Search Jobs

Search by job, company or skills

Agent Evaluation Engineer

Agent Evaluation Engineer

manus ai
Fresher
  • Posted 10 hours ago
  • Be among the first 10 applicants

Job Description

Location: Beijing / Singapore

About the Role

Develop evaluations grounded in product needs, user tasks and how AI models and agents work. Use metrics, experiments and failure analysis to assess capability changes and investigate gaps in existing evaluations, informing system development, post-training and model selection.

Responsibilities

  • Design online and offline metrics that translate user tasks, output quality and practical value into measurable, testable evaluation criteria.
  • Design evaluation tasks and experiments around agent planning, tool use, context and feedback, comparing performance before and after system changes and identifying what affects results.
  • Support post-training evaluations and comparisons of third-party model quality, defining use cases, metric definitions and the conditions to which results apply.
  • Develop new tasks, metrics or experimental methods for capabilities and user experience issues that existing evaluations miss, and test their validity, bias and reproducibility.
  • Analyze evaluation results and failures, distinguish score changes from real capability changes, and validate improvements and refine methods with product, engineering and model teams.

Requirements

  • Practical experience evaluating agent products or model post-training, with concrete evidence of metric design and validation skills.
  • Deep understanding of AI model and agent mechanisms, including task planning, tool use, context management and feedback, and the ability to use this knowledge to analyze system behavior.
  • Data analysis, engineering and research skills, including the ability to design experiments, process evaluation data and analyze uncertainty in results.
  • Ability to investigate open-ended problems, develop well-reasoned new evaluation methods, examine experimental bias and how samples support conclusions, and test judgments through experiments.
  • Independent judgment on AI output quality and real user value, with the ability to explain findings, their applicability and limitations clearly.
  • Familiarity with user research methods such as interviews, observation or usability testing, and the ability to translate findings into evaluation tasks and criteria, are preferred.

Manus excels at various tasks in work and life, getting everything done while you rest at Manus AI.

What we offer:

  • Build at the Frontier of AI Agents - Work with a fast-paced team to turn frontier AI breakthroughs into real-world impact.
  • Equity & Shared success - Our share incentive plan enables employees to participate in Manus's long-term growth and share in the value we create together.
  • Unlimited Manus Tokens - Enjoy unlimited Manus tokens to experiment, build, and supercharge your productivity.

If you're passionate about cutting-edge technology and making a real impact, we'd love to hear from you!

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

metric design

validation skills

user research methods

evaluation methods

context management

task planning

About Company

Similar Jobs

2-5 yrs
SGD 8,000 - 16,000 per month
Singapore
Skills:
Usability Testing, metric design and validation, Data Analysis, AI model and agent mechanisms, Failure Analysis, experiment design, evaluation methods, user research methods
Singapore
Skills:
Python, Incident management practice, Understanding of LLM agent architectures and failure modes, Clear technical writing, Monitoring alerting and observability, Production operations experience, Evaluation methods and metric design for AI systems, LLM observability and evaluation tooling, Experience with LLM agentic systems, Log trace analysis, Dashboarding tools, Evaluation data-analysis or quality experience, Understanding of telco customer intents and journeys