T
LLM Engineer
T
LLM Engineer
techknowledgey pte. ltd.- Posted 3 hours ago
- Be among the first 10 applicants
Job Description
We are seeking a Deep Learning Architect with 3-4 years of experience to help design and build production-grade GenAI systems. In this role, you will contribute architecture coverage across the team-reviewing system designs, identifying gaps, and guiding technical decisions at the solution level. You will work on end-to-end LLM system design, Retrieval-Augmented Generation (RAG) pipelines, and multi-agent architectures, with a strong focus on production readiness. Strong coding depth is non-negotiable.
Sounds great - what will I do
- Design and contribute to end-to-end LLM system architecture for real-world enterprise use cases (from requirements to production).
- Pre-train, fine-tune LLMs and domain-specific models using techniques such as CPT, SFT, LoRA, and QLoRA for client-specific use cases.
- Design and run model evaluation pipelines to benchmark performance, accuracy, and cost across different fine-tuning approaches.
- Optimise models for latency, throughput, token efficiency, and inference cost in production environments.
- Work alongside agent orchestration and architecture teams to integrate fine-tuned models into multi-agent pipelines.
- Implement prompt versioning, rollback strategies, and model monitoring to ensure reliability post-deployment.
- Translate business requirements from client engagements into model adaptation strategies with clear success criteria.
- Contribute to internal knowledge sharing on fine-tuning best practices, tooling, and emerging techniques.
- Define enterprise integration patterns for GenAI systems (identity/access controls, auditability, data boundaries, governance, and compliance alignment).
- Improve production reliability: latency/throughput optimization, token efficiency, cost control, and robust failure handling.
- Collaborate with cross-functional stakeholders (engineering, data, product, client teams) to deliver high-impact solutions on tight timelines.
- Contribute hands-on code, perform code reviews, and raise the engineering bar through strong software fundamentals.
Ideal Experience
- 3-6 years of experience in ML engineering, LLMs, or model development roles.
- Hands-on experience with Continual pre-training (CPT), supervised fine-tuning (SFT), LoRA, or QLoRA on LLMs.
- Strong Python programming skills. Ability to write clean, testable, production-ready code.
- Experience running model evaluation and benchmarking pipelines in a structured way.
- Solid understanding of transformer architectures and how fine-tuning affects model behaviour.
- Experience deploying fine-tuned and pre-trained models to cloud environments with attention to cost and latency.
- Strong problem-solving skills with the ability to work independently on client-facing projects.
- Solid software engineering fundamentals: APIs, data structures, testing, debugging, and performance optimization.
- Ability to review designs, communicate trade-offs clearly, and collaborate effectively in a fast-paced environment.
- Experience with AWS-native GenAI building blocks (e.g., Bedrock, OpenSearch, Lambda, ECS/EKS) and secure enterprise deployments.
What's in it for me
- Impact-first culture where your work ships to production and drives measurable business outcomes.
- Professional growth through mentorship, continuous learning, and exposure to diverse enterprise environments.
- A high-performing team that values curiosity, integrity, ownership, and engineering excellence.
- Opportunities for regional and international exposure as projects require.
More Info
Job Type:
Industry:
Employment Type:





