Key Responsibilities
- Own the design, development, maintenance, and evolution of the in-house AIOps / ML / LLM platform, including related cloud and on-premise Kubernetes solutions.
- Translate client, security, compliance, and internal requirements into practical platform designs with cross-functional teams.
- Build and operate production ML / LLM workflows, including retraining, deployment, inference serving, monitoring, rollback, and optimisation.
- Troubleshoot production issues across application, infrastructure, networking, Linux, Kubernetes, and ML serving layers.
Qualifications / Requirements
- Strong software/platform engineering fundamentals, including system design, API design, distributed systems, scalability, reliability, observability, authentication/authorization, testing, and maintainable code design.
- Practical understanding of the ML / LLM lifecycle, including data pipelines, model training/retraining, evaluation, experiment tracking, deployment, monitoring, and production feedback loops.
- Strong development experience in Python, with working proficiency in Go and C++ for reading, debugging, maintaining, and extending existing production codebases.
- Strong Linux, networking, and Kubernetes fundamentals, including production troubleshooting, service connectivity, ingress, resource limits, workload debugging, and deployment operations.
- Experience designing, deploying, and operating production platforms on AWS, Azure, GCP, or on-premise environments.
- Experience building CI/CD, automation, and MLOps / LLMOps workflows for production ML / LLM systems.
- Strong communication skills and ability to work with AI, deployment, infrastructure, and security teams.
Good to Have
- Deep experience operating Kubernetes in bare-metal, air-gapped, or restricted on-premise environments.
- Experience with MLflow, Kubeflow, vLLM, TensorRT, TGI, or similar ML / LLM platform tools.
- Exposure to TypeScript / React or Java-based services.