Agent Evaluation Engineer
Job Description
Location: Beijing / Singapore
About the Role
Develop evaluations grounded in product needs, user tasks and how AI models and agents work. Use metrics, experiments and failure analysis to assess capability changes and investigate gaps in existing evaluations, informing system development, post-training and model selection.
Responsibilities
- Design online and offline metrics that translate user tasks, output quality and practical value into measurable, testable evaluation criteria.
- Design evaluation tasks and experiments around agent planning, tool use, context and feedback, comparing performance before and after system changes and identifying what affects results.
- Support post-training evaluations and comparisons of third-party model quality, defining use cases, metric definitions and the conditions to which results apply.
- Develop new tasks, metrics or experimental methods for capabilities and user experience issues that existing evaluations miss, and test their validity, bias and reproducibility.
- Analyze evaluation results and failures, distinguish score changes from real capability changes, and validate improvements and refine methods with product, engineering and model teams.
Requirements
- Practical experience evaluating agent products or model post-training, with concrete evidence of metric design and validation skills.
- Deep understanding of AI model and agent mechanisms, including task planning, tool use, context management and feedback, and the ability to use this knowledge to analyze system behavior.
- Data analysis, engineering and research skills, including the ability to design experiments, process evaluation data and analyze uncertainty in results.
- Ability to investigate open-ended problems, develop well-reasoned new evaluation methods, examine experimental bias and how samples support conclusions, and test judgments through experiments.
- Independent judgment on AI output quality and real user value, with the ability to explain findings, their applicability and limitations clearly.
- Familiarity with user research methods such as interviews, observation or usability testing, and the ability to translate findings into evaluation tasks and criteria, are preferred.
Manus excels at various tasks in work and life, getting everything done while you rest at Manus AI.
What we offer:
- Build at the Frontier of AI Agents - Work with a fast-paced team to turn frontier AI breakthroughs into real-world impact.
- Equity & Shared success - Our share incentive plan enables employees to participate in Manus's long-term growth and share in the value we create together.
- Unlimited Manus Tokens - Enjoy unlimited Manus tokens to experiment, build, and supercharge your productivity.
If you're passionate about cutting-edge technology and making a real impact, we'd love to hear from you!
More Info
Key Skills
metric design
validation skills
user research methods
evaluation methods
context management
task planning

