
Search by job, company or skills
This is a 6-month contract role under Optimum Solutions.
This role supports our client, a leading Southeast Asian superapp that offers a wide range of services, including mobility, food delivery, and digital payments.
Work Schedule and Hours: Monday - Friday, 10am - 7pm (inclusive of a 1-hour lunch break)
Key Responsibilities:
. Take ownership of the ACS embedded-evaluation workstream.
. Establish reliable regression evaluation for at least one additional agent team.
. Make its distributed traces consistently usable for debugging and automated evaluation.
. Ship at least one reusable improvement to EvalsHub or its SDK/tooling.
. Reduce manual evaluation effort through automated datasets, evaluators, and CI integration.
. Document ownership, operational procedures, and onboarding guidance.
. Embed with AI product teams and own measurable agent-quality outcomes.
. Build and maintain golden datasets, regression suites, evaluation pipelines, and quality dashboards.
. Convert product rules and human review procedures into deterministic, numeric, LLM-as-judge, trajectory, tool-selection, SOP-adherence, and multi-turn evaluators.
. Diagnose evaluation and production-trace failures, then convert findings into dataset improvements, evaluator changes, or fixes to prompts, tools, skills, and agent workflows.
. Instrument distributed agent systems using OpenTelemetry, including Temporal workflows and Go/Python services.
. Ensure traces contain consistent inputs, outputs, tool calls, metadata, feedback, and root-agent results.
. Develop reusable capabilities across EvalsHub backend, SDK, CLI/plugins, and supporting frontend interfaces.
. Integrate evaluations into CI/CD and model-release workflows.
. Run model, retrieval, embedding, and agent-architecture experiments balance accuracy, latency, and cost.
. Drive migrations from legacy evaluation and observability systems.
. Write technical designs and documentation, run demonstrations, and enable engineering teams to use evaluation tooling independently.
. Participate in the LLMOps operational rotation and support production-quality integrations.
Qualifications:
. Minimum 2 years of relevant experience in one or more of the following areas:
EvalsHub, LangSmith, LangGraph/LangChain, FastAPI, Temporal, Grafana, Redis, or similar technologies.
React/TypeScript and internal developer-tooling interfaces.
LLM-as-judge, multi-turn evaluation, tool/MCP evaluation, and agent-trajectory analysis.
Text-to-SQL/Text-to-DSL, RAG, semantic retrieval, embedding evaluation, or model benchmarking.
Building platform capabilities that can be reused across several AI products.
. Strong analytical and problem-solving skills related to AI systems, evaluation frameworks, and data-driven quality improvements.
. Excellent project management and cross-department communication abilities.
. A detail-oriented and metrics-driven approach to continuous improvement
Job ID: 151379591
Skills:
Java, Golang, Cassandra, Bash, HBase, Redis, Terraform, Kubernetes, Python, Airflow, Step Functions, GitOps, Cadence, Temporal, ArgoCD
Skills:
React, Typescript, FastAPI, Grafana, Redis, LangGraph, LangChain, EvalsHub, Temporal, LangSmith
Skills:
React, Typescript, Gcp, Nodejs, Azure, AWS
Skills:
Java, Spring Boot, JIRA, Sql, Nosql, Tensorflow, Pytorch, Gcp, FastAPI, Restful Apis, Azure, Python, AWS, Caching technologies, OpenAI
Skills:
Java, Express, Vue, Agile Methodologies, PostgreSQL, Node.js, Spring Boot, Mssql, Sql, Angular, Nosql, React, Typescript, DevSecOps, Javascript, FastAPI, MongoDB, Python, software engineering fundamentals, Next.js, full-stack development