Agent Evaluation Engineer for an AI Agent Company
dadaconsultants pte. ltd.- Posted 58 minutes ago
- Be among the first 10 applicants
Job Description
About our client
Our client is a fast-growing, well-funded AI agent company. The product takes on complex work like deep research, data analysis and software development, and carries it through to a finished result. It is a small team building at the frontier, based in Singapore.
About the role
You will develop evaluations grounded in product needs, user tasks and how AI models and agents actually work. Using metrics, experiments and failure analysis, you'll assess whether capability changes are real, investigate gaps in existing evaluations, and inform system development, post-training and model selection.
What you'll do
- Design online and offline metrics that translate user tasks, output quality and practical value into measurable, testable evaluation criteria
- Design evaluation tasks and experiments around agent planning, tool use, context and feedback, comparing performance before and after system changes
- Support post-training evaluations and comparisons of third-party model quality, defining use cases and the conditions results apply to
- Develop new tasks, metrics or experimental methods for issues existing evaluations miss, and test their validity, bias and reproducibility
- Analyze evaluation results and failures, distinguish score changes from real capability changes, and work with product, engineering and model teams to validate improvements
What you bring
- Practical experience evaluating agent products or model post-training, with concrete evidence of metric design and validation
- Deep understanding of AI model and agent mechanisms - task planning, tool use, context management and feedback
- Strong data analysis, engineering and research skills, including experiment design and handling uncertainty in results
- Ability to investigate open-ended problems, develop well-reasoned new evaluation methods, and examine experimental bias
- Independent judgement on AI output quality and real user value, with the ability to explain findings and their limitations clearly
Nice to have
- Familiarity with user research methods (interviews, observation, usability testing), and the ability to translate findings into evaluation criteria
What we offer
- Build at the frontier of AI agents with a fast-paced team
- Unlimited access to our own AI tools
- Competitive salary, reviewed for the right person
If it sounds like your next move, please don't hesitate to apply. Kindly note that only shortlisted candidates will be contacted. Appreciate your understanding. Data provided is for recruitment purposes only.
About Us
Dada Consultants was established in 2017, with the commitment of providing the best recruitment services in Singapore. We are comprised of a dynamic head-hunting team dedicated to sourcing for highly competent professionals in IT industry. We provide enterprises with customized talent solutions, and bring talents to career advancement.
EA Registration Number: R23118712
Business Registration Number: 201735941W. Licence Number: 18S9037
www.dadaconsultants.com
More Info
Key Skills
metric design and validation
AI model and agent mechanisms
experiment design
evaluation methods
user research methods
