Search by job, company or skills

Senior AI Harness Engineer

5-8 Years
SGD 5,000 - 10,000 per month
  • Posted 20 hours ago
  • Be among the first 10 applicants

Job Description

Avensys is a reputed global IT professional services company headquartered in Singapore. Our service spectrum includes enterprise solution consulting, business intelligence, business process automation and managed services. Given our decade of success we have evolved to become one of the top trusted providers in Singapore and service a client base across banking and financial services, insurance, information technology, healthcare, retail, and supply chain.

We are currently looking to hire Senior AI Harness Engineer.

This is an exciting opportunity to expand your skill set, achieve job satisfaction and work-life balance. More details as below.

Job Role:

Job Description

Senior AI Harness Engineer

Position Overview

We are looking for a Senior AI Harness Engineer to design, build, and maintain the engineering frameworks, evaluation systems, testing infrastructure, and tooling required to develop reliable, scalable, and production-ready AI/LLM applications. The role will focus on building AI harnesses that enable systematic testing, evaluation, benchmarking, observability, and continuous improvement of Generative AI and Agentic AI systems.

The ideal candidate will have strong experience in Python, LLMs, Generative AI, AI agents, evaluation frameworks, RAG, prompt engineering, API integration, cloud platforms, and MLOps/LLMOps, with the ability to build robust engineering solutions around AI models.

Key Responsibilities

. Design and develop AI/LLM evaluation and testing harnesses for Generative AI and Agentic AI applications.

. Build reusable frameworks for model evaluation, prompt testing, regression testing, benchmarking, and performance validation.

. Develop automated test suites to evaluate accuracy, relevance, groundedness, hallucination, toxicity, safety, latency, cost, and response quality.

. Create testing frameworks for LLM-based agents, multi-agent workflows, RAG pipelines, and tool-calling applications.

. Develop and maintain Python-based AI engineering frameworks, utilities, SDKs, and automation tools.

. Implement LLMOps/MLOps pipelines for model, prompt, dataset, and evaluation lifecycle management.

. Build automated CI/CD pipelines for AI applications, including model and prompt regression testing.

. Integrate AI evaluation tools and frameworks such as MLflow, LangSmith, DeepEval, Ragas, Azure AI evaluation, OpenAI evaluation frameworks, or equivalent technologies.

. Design evaluation datasets, test cases, golden datasets, benchmark datasets, and synthetic test data.

. Develop automated mechanisms for prompt/version management and experiment tracking.

. Implement observability and monitoring for AI applications, including LLM traces, token usage, latency, errors, model performance, and cost.

. Develop harnesses for testing RAG systems, including retrieval quality, chunking strategies, embeddings, vector search, reranking, and grounded responses.

. Build evaluation capabilities for function calling, API/tool invocation, MCP-based integrations, and agentic workflows.

. Work with AI/ML engineers and application teams to identify failure patterns and improve model/application performance.

. Establish engineering standards for AI quality, reliability, security, scalability, and responsible AI.

. Troubleshoot complex issues across application code, LLM APIs, vector databases, cloud services, and AI infrastructure.

. Build dashboards and reports to communicate AI evaluation and quality metrics to technical stakeholders.

. Mentor junior engineers and contribute to architecture and technical design decisions.

Required Technical Skills

Programming

. Strong hands-on experience with Python.

. Good knowledge of REST APIs, JSON, asynchronous programming, SDK development, and microservices.

. Experience with Git, GitHub/GitLab/Bitbucket, code reviews, and software engineering best practices.

Generative AI & LLM

. Strong understanding of LLMs and Generative AI application architecture.

. Experience with models such as OpenAI, Azure OpenAI, Anthropic, Gemini, Llama, Mistral, or equivalent.

. Strong knowledge of:

o Prompt Engineering

o RAG

o Embeddings

o Vector databases

o Function/Tool Calling

o Structured outputs

o Agentic AI

o LLM APIs

o Context management

o Model evaluation

AI/Agent Frameworks

Hands-on experience with one or more:

. LangChain

. LangGraph

. LlamaIndex

. Semantic Kernel

. AutoGen

. CrewAI

. OpenAI Agents SDK

. MCP (Model Context Protocol)

AI Evaluation & Testing

Experience with AI/LLM evaluation frameworks such as:

. Ragas

. DeepEval

. LangSmith

. MLflow

. Azure AI Evaluation

. OpenAI evaluation frameworks

. Custom LLM evaluation frameworks

Knowledge of evaluation metrics including:

. Accuracy

. Precision / Recall

. Relevance

. Faithfulness

. Groundedness

. Context relevance

. Hallucination detection

. Toxicity and safety

. Robustness

. Latency

. Token consumption

. Cost per request

RAG & Vector Technologies

. Experience designing and testing RAG pipelines.

. Knowledge of vector databases such as:

o Azure AI Search

o Pinecone

o Weaviate

o Milvus

o Qdrant

o FAISS

o Elasticsearch/OpenSearch

. Experience evaluating retrieval quality, embeddings, reranking, chunking, and search strategies.

Cloud & DevOps

Experience with one or more cloud platforms:

. Microsoft Azure

. AWS

. Google Cloud Platform

Good understanding of:

. Docker

. Kubernetes

. CI/CD

. GitHub Actions / Azure DevOps / Jenkins

. Infrastructure automation

. Cloud monitoring and logging

. Secrets and configuration management

MLOps / LLMOps

. Experience implementing MLOps/LLMOps practices.

. Model and prompt versioning.

. Experiment tracking.

. Dataset management.

. Automated evaluation pipelines.

. Continuous monitoring.

. Model/application performance tracking.

. Production deployment and rollback strategies.

Preferred Skills

. Experience building AI testing platforms or internal AI developer tools.

. Experience with agent evaluation and multi-agent systems.

. Experience with MCP servers and MCP-based tool integrations.

. Knowledge of AI security, prompt injection, jailbreak testing, data leakage, and adversarial evaluation.

. Experience with synthetic data generation and automated test-case generation.

. Knowledge of responsible AI and AI governance.

. Experience with SQL and NoSQL databases.

. Familiarity with Kafka or event-driven architectures.

. Experience with TypeScript/Node.js is an advantage.

. Experience working in Agile/Scrum environments.

Key Deliverables

The Senior AI Harness Engineer will be expected to deliver:

. Reusable AI testing and evaluation frameworks.

. Automated LLM and agent regression test suites.

. AI quality and performance dashboards.

. Automated evaluation pipelines integrated with CI/CD.

. Benchmark and golden datasets.

. RAG and agent evaluation capabilities.

. Prompt/model comparison frameworks.

. AI observability and monitoring solutions.

. Production-ready engineering standards for GenAI applications.

Experience

. 8+ years of overall software/engineering experience.

. 4+ years of experience in AI/ML or Generative AI engineering.

. Strong hands-on experience with Python and AI/LLM technologies.

. Proven experience developing AI evaluation, testing, automation, or LLMOps solutions.

. Experience working with enterprise-scale AI applications and production deployments.

Soft Skills

. Strong analytical and problem-solving skills.

. Ability to translate AI quality requirements into measurable engineering solutions.

. Strong communication and collaboration skills.

. Ability to work with AI researchers, ML engineers, software developers, DevOps engineers, and business stakeholders.

. Strong ownership mindset with a focus on building reliable and maintainable AI systems.

. Ability to mentor engineers and provide technical leadership.

Education

Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Engineering, or a related field.

CONSULTANT DETAILS

Consultant Name: Abinaya R
Reg No: R1765546
Avensys Consulting Pte Ltd
EA Licence 12C5759

Privacy Statement: Data collected will be used for recruitment purposes only. Personal data provided will be used strictly in accordance with the relevant data protection law and Avensys privacy policy.

More Info

Job Type:
Industry:
Employment Type:

Job ID: 153335327

Beware of Scammers

We don’t charge money for job offers