Search by job, company or skills

ML Evaluation Engineer

3-7 Years
SGD 5,000 - 10,000 per month
Quick Apply
  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

Summary

We are looking for an ML Evaluation Engineer/Data Scientist to bring statistical rigor and measurement excellence to the evaluation of AI and LLM systems. This role will focus on ensuring that evaluation results are reliable, reproducible, and scientifically valid. 

The successful candidate will partner closely with ML Evaluation Engineers and Applied Scientists to answer critical questions such as: Is this performance improvement real, Are our evaluators aligned, and Can we trust these results 

 

Key Responsibilities:

  • Design and execute statistical analyses for AI and LLM evaluation programs. 
  • Measure and improve inter-annotator agreement between human evaluators. 
  • Develop sampling methodologies and evaluation datasets to ensure representative results. 
  • Perform significance testing and statistical validation of model performance changes. 
  • Analyze golden sets and benchmark datasets to assess model quality and consistency. 
  • Build dashboards and reporting frameworks to monitor evaluation outcomes and trends. 
  • Identify sources of evaluation noise, bias, and inconsistency. 
  • Partner with ML Evaluation Engineers to validate automated judges against human raters. 
  • Provide data-driven recommendations regarding evaluation quality and model readiness. 

Must Have Skills:

  • Bachelor's degree in Statistics, Data Science, Computer Science, Mathematics, Machine Learning, or a related field. 
  • Minimum 2+ years of professional experience in Data Science, Evaluation Science, Applied Statistics, or related disciplines. 
  • Strong foundation in statistical analysis, hypothesis testing, experimental design, and measurement methodologies. 
  • Experience working with structured and unstructured data using Python, SQL, and data analysis tools. 
  • Ability to communicate statistical findings clearly to technical and non-technical stakeholders. 
  • Experience building analytical dashboards and reporting solutions. 

 

Nice to haves:

  • Master's degree or PhD in Statistics, Data Science, Computer Science, Quantitative Social Sciences, or a related discipline. 
  • Direct experience supporting LLM, AI, NLP, or ML evaluation initiatives. 
  • Experience measuring agreement between human raters and automated evaluation systems. 
  • Knowledge of annotation quality frameworks, golden-set construction, sampling theory, and evaluation benchmarking. 
  • Familiarity with LLM-as-Judge frameworks and AI evaluation platforms. 

 

We regret to inform that only shortlisted candidates will be notified / contacted.

 

 EA registration number : YAP JIA YI, R25157934

Allegis Group Singapore Pte Ltd, Company Reg No. 200909448N, EA Licence No. 10C4544

More Info

Job Type:
Function:

About Company

Allegis Group is the global leader in talent solutions focused on working harder and caring more than any other provider. We'll go further to understand the needs of our people; our clients, our candidates, and our employees; and to consistently deliver on our promise of an unsurpassed quality experience. That's the Allegis Group difference, and it's consistent across every Allegis Group company. With more than US$11billion in annual revenues and over 500 locations across the globe, our network provides businesses with a comprehensive suite of talent solutions; without sacrificing the niche expertise required to ensure a successful partnership. Our specialised group of companies includes: Aerotek, TEKsystems, Allegis Global Solutions, Aston Carter, Major, Lindsey; Africa, Allegis Partners, MarketSource, and EASi. Visit www.AllegisGroup.com to learn more.

Allegis Group Singapore Pte Ltd,
Company Reg No. 200909448N, EA Licence No. 10C4544

Job ID: 153208065

Beware of Scammers

We don’t charge money for job offers