Role Overview
We are looking for a Senior AI Systems Engineer with deep experience in multi-model systems, model routing, AI evaluation, and production AI infrastructure. The role will focus on designing systems that can compare, select, and operate multiple foundation models reliably in production.
.AI AI AI
The ideal candidate should be able to define evaluation standards, translate evaluation results into routing strategies, and build the supporting platform required to optimize model quality, latency, cost, and reliability.
.
Key Responsibilities
Model Routing Strategy.
- Design and implement model-routing strategies across multiple foundation models based on task type, complexity, quality requirements, latency, cost, and operational risk.
. - Develop dynamic model-selection, arbitration, escalation, fallback, and ensemble mechanisms.
- Define routing policies and decision criteria for different task categories and production scenarios.
. - Continuously improve routing performance using evaluation results, production data, and controlled experimentation.
. - Distinguish and optimize for model capability rather than relying solely on availability-based failover or static rule-based routing.
.
Evaluation Strategy, Framework & Benchmarking
- Define task taxonomies, evaluation dimensions, scoring criteria, and acceptance thresholds for different AI use cases.
AI - Design and maintain benchmark suites, golden datasets, annotation standards, and regression test sets.
Golden Dataset - Build automated evaluation pipelines using deterministic checks, LLM-as-a-Judge, rubric-based scoring, pairwise comparison, and human evaluation where appropriate.
LLM-as-a-JudgeRubric Pairwise Comparison - Validate evaluation methods against human judgments and monitor judge consistency, bias, and drift.
Judge - Establish continuous evaluation and regression mechanisms for model, prompt, data, and routing-policy changes.
Prompt. - Ensure evaluation signals are reliable enough to support model comparison and routing decisions.
.
Multi-model and Model-serving Architecture
- Design and build scalable infrastructure for integrating and serving multiple proprietary and open-source models.
- Develop model gateways, unified APIs, version-management mechanisms, and routing infrastructure.
. - Support model lifecycle management, traffic governance, rollout, rollback, and model replacement.
- Improve the reliability, scalability, and operational efficiency of multi-model production systems.
Quality, Cost and Latency Optimization
- Define measurable optimization objectives across model quality, inference cost, latency, reliability, and resource utilization.
- Develop systematic methods to balance competing objectives rather than optimizing a single metric in isolation.
- Compare routing strategies against fixed-model baselines and quantify their operational and performance benefits.
. - Use offline evaluation, online experimentation, and production feedback to improve routing policies continuously.
.
Production Engineering and Observability
- Build production-grade AI services with strong standards for reliability, security, testing, and maintainability.
AI - Implement monitoring, logging, tracing, failure analysis, and performance diagnostics for model calls and routing decisions.
.. - Establish data feedback loops to capture model performance, routing outcomes, failure cases, and human-review results.
. - Support incident investigation, production debugging, regression analysis, and system recovery.
Qualifications
- Hands-on experience building production AI, machine-learning, or distributed systems.
AI - Demonstrated experience in model routing, dynamic model selection, model arbitration, ensemble systems, or multi-model decision systems.
. - Experience designing evaluation frameworks, benchmark datasets, scoring methodologies, regression pipelines, or continuous evaluation systems.
Benchmark Dataset - Strong understanding of the trade-offs among model quality, latency, inference cost, reliability, and operational risk.
- Experience integrating and operating multiple foundation models, including third-party APIs, self-hosted models, or open-source models.
API - Ability to translate ambiguous requirements into measurable evaluation criteria, technical designs, and production systems.
- Strong analytical and problem-solving skills, with the ability to validate technical decisions using data and experiments.
Preferred Qualifications
- Experience with LLM-as-a-Judge, human evaluation, judge calibration, preference evaluation, or benchmark development.
LLM-as-a-JudgeJudge Benchmark - Experience with model serving, inference optimization, model gateways, observability platforms, or LLMOps systems.
LLMOps - Experience with reinforcement learning, contextual bandits, ranking, recommendation systems, search, or information retrieval.
Contextual Bandit - Experience with AI safety, red teaming, model governance, or evaluation of high-risk AI systems.
AI Red Teaming AI
What We Value
- You think in terms of systems, objectives, and measurable outcomes rather than individual prompts or isolated model outputs.
Prompt - You can distinguish true model routing from static rules, failover, or load balancing.
. - You understand that reliable routing depends on reliable evaluation.
. - You are comfortable owning both technical strategy and production implementation.