Serve as the senior technical evaluator within the team, responsible for the design, calibration, and adjudication of evaluation suites and gate thresholds across agent archetypes. Lead gate reviews under the Team Lead's authority, own the regression-pack methodology for AI/ML model changes, and act as the technical custodian of evaluation quality and drift hygiene.
Make an Impact by:
- Design and maintain offline evaluation suites (golden sets, regression packs, adversarial/safety probes) and the continuous-evaluation scoring pipeline across archetypes.
- Conduct operability gate reviews: assess evidence packs, reproduce evaluation results, and recommend go/no-go with documented findings.
- Own the model-update regression pack methodology and adjudicate regression runs against archived baselines with AIML Operations team.
- Calibrate gate thresholds against production reality and maintain evaluation drift hygiene (golden-set rotation, hold-out sets, judge calibration).
- Produce the monthly quality report per agent: eval trends, failure-mode taxonomy, and defect clusters with reproduction traces.
- Mentor members of the team when needed and review their work for quality and consistency.
Skills for Success:
- Bachelor's or Master's degree in Computer Science or a related field
- 6+ years in ML/data/software with strong evaluation or quality focus
- Hands-on experience evaluating LLM or ML systems
- LLM/agent evaluation design and statistical rigour
- Python and evaluation tooling (promptfoo, DeepEval, or custom harnesses)
- Data analysis and metric interpretation
- Tracing/observability tooling
- Analytical rigour and attention to detail
- Clear technical writing for gate findings
- Ability to influence build teams on quality
- Understanding of telco customer intents and journeys