Software Quality Assurance Engineer GenAI / LLM
d l resources pte ltd- Posted 2 hours ago
- Be among the first 10 applicants
Job Description
Primary Focus: Software Quality Assurance / Software Testing / Test Automation - GenAI, LLM & Agentic AI
Secondary Exposure: Solution Analysis / Technology Solution Design / Enterprise Integration
Domain / Project: Global Markets, Capital Markets Banking Technology & Market Risk Technology
Role Overview
We are looking for a Senior GenAI Quality Engineer / Solution Analyst to design, analyse, test and validate production-grade Generative AI (GenAI), Large Language Model (LLM), RAG and Agentic AI applications within a complex enterprise environment.
This is not a traditional manual QA or software testing role.
The role combines:
Software Quality Engineering
GenAI / LLM Testing & Evaluation
Agentic AI / AI Agent Testing
UI & API Testing
Test Automation
Solution Analysis
Enterprise Integration Testing
Observability & Troubleshooting
You will work across discovery, solution design, development, testing and release, translating business requirements into clear application behaviours and validating end-to-end application quality across user interfaces, APIs, data flows, LLMs, RAG components, AI agents and enterprise integrations.
Key Responsibilities
GenAI / LLM Quality Engineering
Define and execute end-to-end quality engineering and test strategies covering:
Web / UI workflows
REST APIs
Backend services
Enterprise integrations
GenAI applications
LLM workflows
RAG pipelines
Agentic AI / AI Agent interfaces
Perform GenAI / LLM testing and evaluation covering:
Response quality
Task completion
Grounding
Faithfulness
Relevance
Consistency
Citation accuracy
Hallucination risk
Safe failure behaviour
Test non-deterministic / probabilistic AI systems using:
Evaluation datasets
Repeat testing
Quality thresholds
Acceptance criteria
Regression evaluation
Validate RAG / Retrieval-Augmented Generation solutions, including retrieval quality, grounding and response accuracy.
Agentic AI / AI Agent Testing
Test end-to-end Agentic AI and AI Agent workflows, including:
Multi-turn conversations
Context handling
Agent planning
Tool selection
Tool calling / function calling
Tool inputs and outputs
State transitions
Memory and state
Human-in-the-loop approvals
Handoffs
Retries
Timeouts
Fallback behaviour
Error recovery
Termination conditions
Partial failures
Validate that AI agents behave correctly across both successful and failure scenarios.
Software & API Quality Engineering
Perform:
Functional Testing
Integration Testing
API Testing
Regression Testing
Exploratory Testing
Negative Testing
Resilience Testing
Basic Performance Testing
End-to-End Testing
Design comprehensive REST API tests covering:
API contracts
Authentication
Authorisation
Input validation
Error handling
Idempotency
Rate limits
Downstream system failures
Test web application behaviour across browsers and realistic end-user journeys, including:
Loading states
Interrupted sessions
Error messages
Feedback capture
Accessibility fundamentals
Test Automation
Develop and maintain risk-based test automation that reduces:
Regression testing time
Manual testing effort
Release cycle time
Production risk
Use automation frameworks and tools such as:
Playwright
Cypress
Selenium
pytest
REST Assured
Postman
Equivalent UI / API automation frameworks
Apply pragmatic automation principles by prioritising stable, high-value and frequently executed test scenarios.
GenAI Evaluation & AI Safety Testing
Validate LLM and GenAI applications for:
Grounded responses
Hallucinations
Retrieval quality
Citation accuracy
Prompt behaviour
Prompt injection
Unsupported requests
Restricted content handling
Safe failure behaviour
Adversarial scenarios
Support AI evaluation / LLM evaluation using appropriate evaluation datasets, quality metrics and repeatable evaluation approaches.
Exposure to AI Red Teaming / Adversarial Testing would be advantageous.
Observability & Troubleshooting
Use application and GenAI observability to identify the source of defects across:
Application
LLM / Model
RAG / Retrieval
Data
API / Integration
Platform
Analyse:
Logs
Distributed traces
API requests / responses
Payloads
Network calls
Database records
Agent execution traces
Exposure to observability and LLM evaluation tools such as:
Langfuse
LangSmith
OpenTelemetry
Elastic / Elasticsearch
Splunk
is advantageous.
Solution Analysis & Design
The role also acts as a hands-on Solution Analyst for GenAI applications.
Responsibilities include:
Partner with product owners, business users, architects, engineers and GenAI specialists during discovery and solution design.
Analyse proposed GenAI use cases and determine whether the requirement should use:
Conventional application logic
Deterministic business rules
Search / retrieval
RAG
Workflow automation
Agentic AI
Human approval
Translate business requirements into:
Functional requirements
End-to-end solution flows
User journeys
Acceptance criteria
Interface behaviour
Decision rules
Non-functional requirements
Map interactions across:
User Interfaces
APIs
LLMs / Models
Prompts
RAG / Retrieval components
Enterprise data sources
AI Agent tools
Downstream enterprise systems
Analyse solution design trade-offs involving:
Quality
Complexity
Cost
Latency
Security
Data access
Maintainability
Operational risk
Identify missing controls, integration assumptions, ownership gaps, failure scenarios and operational risks before development begins.
Support the design of:
Human-in-the-loop approval
Fallback flows
Escalation
Exception handling
Solution Documentation
Produce practical technical and functional artefacts including:
Process Flows
Sequence Diagrams
Context Diagrams
Interface Specifications
Decision Tables
User Stories
Acceptance Criteria
Test Scenarios
Traceability Documentation
Maintain traceability across:
Business Requirement → Solution Design → Implementation → Test / Evaluation Scenario → Release Evidence
Release Quality & Governance
Create and maintain:
Test scenarios
Test datasets
Reusable regression scenarios
Test evidence
Defect reports
Quality metrics
Release quality reports
Provide evidence-based release recommendations identifying:
Known defects
Known limitations
Residual risks
Quality concerns
Areas requiring production monitoring
Core Requirements
Experience
5-8 years of experience in Software Quality Engineering, Test Engineering, Test Automation, SDET or similar hands-on software testing roles.
Strong experience testing complex enterprise applications.
Strong experience testing:
Web applications
REST APIs
Backend services
Enterprise integrations
Test Automation / Programming
Hands-on experience with one or more of:
Playwright
Cypress
Selenium
pytest
REST Assured
Postman
Equivalent automation frameworks
Working programming knowledge of:
Python
Java
JavaScript
TypeScript
Candidates should be capable of developing, reviewing and troubleshooting test automation.
Software Engineering / DevOps
Experience with:
Git
Pull Requests
CI/CD
Automated Testing
Test Reporting
Defect Management
Experience validating distributed systems including:
Asynchronous Processing
Queues
Batch Processing
APIs
Downstream Dependencies
Enterprise Integrations
GenAI / LLM Requirements
Practical understanding of:
Generative AI / GenAI
Large Language Models / LLM
LLM Evaluation
LLM Testing
Retrieval-Augmented Generation / RAG
RAG Evaluation
Agentic AI
AI Agents
Multi-Agent Workflows
Prompts / Prompt Engineering
Context Windows
Embeddings
Tool Calling
Agent Memory & State
LLM Observability
Candidates should understand how GenAI applications differ from conventional deterministic software and how to validate probabilistic AI behaviour.
Security & Risk Testing
Understanding of software and GenAI security fundamentals including:
Access Control
Authentication / Authorisation
Sensitive Data Handling
Input Validation
Auditability
Prompt Injection
AI Safety Testing
Adversarial Testing
Nice to Have
Experience with:
Banking / Financial Services
Regulated enterprise environments
Contract Testing
Service Virtualisation
Synthetic Monitoring
Performance Testing
AI Red Teaming
Accessibility Testing / WCAG
Kubernetes
OpenShift
AWS
Containerised Application Deployment
Key Domain / Technical Skills
1. Software Quality Engineering, API Testing & Test Automation
2. GenAI / LLM Evaluation, RAG & Agentic AI Testing
3. Solution Analysis, Observability & Enterprise Integration
Key Search Keywords
GenAI Quality Engineer,AI Quality Engineer,LLM Quality Engineer,Generative AI Testing,GenAI Testing,LLM Testing,LLM Evaluation,AI Evaluation,Agentic AI Testing,AI Agent Testing,RAG Testing,RAG Evaluation,Retrieval-Augmented Generation,Software Quality Engineering,Quality Engineering,Software QA,Test Automation,SDET,Automation Testing,API Testing,REST API Testing,UI Testing,Integration Testing,Regression Testing,End-to-End Testing,Playwright,Cypress,Selenium,pytest,REST Assured,Postman,Python,Java,JavaScript,TypeScript,CI/CD,Git,Prompt Testing,Prompt Injection,Hallucination Testing,Grounding,Faithfulness,AI Safety Testing,Adversarial Testing,AI Red Teaming,Langfuse,LangSmith,OpenTelemetry,Elastic,Splunk,Observability,Distributed Systems,Kubernetes,OpenShift,Solution Analysis
More Info
Key Skills
LLMs
Acceptance Criteria
Meaningful Use
Operational Risk Assessment
