
Search by job, company or skills
AI Singapore (AISG) is a national AI programme launched by the National Research Foundation (NRF), Singapore, to build and anchor deep national capabilities in AI.AISG is supported through a government-wide partnership including the NRF, Ministry of Digital Development and Information (MDDI), Infocomm Media Development Authority (IMDA), Economic Development Board (EDB) and Enterprise Singapore (ESG). We bring together research institutions and the vibrant ecosystem of AI start-ups and companies to support impactful research, develop talent, and power Singapore's AI efforts.
This position will be hosted at the Nanyang Technological University (NTU) under VP (Artificial Intelligence & Digital Economy)s office and we welcome you to join our community.
We're looking for an AI Engineer to join the Platform team within AI Products at AISG. In this role, you will be collaborating with different internal teams to design and implement optimized inference workflows and support the team to build customized LLM-based solutions.
Your work will directly contribute to the deployment, optimisation and management of large language models (LLMs) in production, retrieval-augmented generation (RAG) services, AI agent orchestration platforms and GPU-enabled AI infrastructure.
Responsibilities:
Platform operations and reliability
Own day-to-day operations of SEA-LION API Farm, our multi-cloud LLM inference platform - monitoring environment health, GPU capacity, performance, cost, and security posture.
Optimise LLM inference across various modalities to drive business value and support production goals.
Diagnose and troubleshoot performance and reliability issues on API Farm.
Build new API services such as batch API services, MCP services.
Infrastructure, CI/CD, and automation
Manage high performance AI clusters and storage systems using infrastructure-as-code (e.g. Terraform) across different cloud providers.
Develop and maintain CI/CD pipelines, container build/registry workflows, and deployment automation so teams can ship safely and frequently.
Strengthen observability across the stack including logs, metrics, traces, and dashboards, and reduce toil by automating repetitive operational tasks.
AI-assisted ops and continuous improvement
Use AI tools (e.g. Claude, Copilot, Cursor) appropriately in your daily work responsibilities.
Build internal tools leveraging AI to reduce manual effort in day-to-day operations.
Requirements:
You should be a hands-on engineer who is comfortable operating cloud and GPU infrastructure end-to-end, who understands how to deploy and run large language models reliably in production, and who actively uses AI tools to make platform work faster and more reliable.
A degree in Computer Science, Information Technology, or equivalent.
1-3 years of DevOps, SRE, or platform engineering experience, with a track record of operating production systems at scale.
Hands-on experience operating workloads on different cloud providers including IaC (e.g. Terraform), containers and orchestration (e.g. Docker, Kubernetes), and managed services for compute, storage, and networking.
Strong knowledge on Inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers).
Hands-on experience deploying and serving LLMs in production - model serving, GPU scheduling, autoscaling, latency/throughput optimisation, and inference cost management.
REST API design, model context protocol (MCP), Internet authentication patterns (e.g. OAuth).
Strong fundamentals in CI/CD, observability (logs/metrics/traces), and incident response.
Demonstrated use of AI tools (e.g. Claude, Copilot, Cursor) in your day-to-day engineering - for code generation, review, debugging, and documentation - with a clear sense of where they help and where they don't.
Solid scripting/programming skills (e.g. Python, Bash) and comfortable reading other people's code across the stack.
Strong communication skills with the ability to explain technical concepts.
Good to Have:
Experience with multimodal AI models (e.g. vision language models, audio language models).
C/C++/Rust/Go or other relevant programming languages.
Contributions to open-source AI/ML projects.
We regret that only shortlisted candidates will be notified.
Hiring Institution: NTUJob ID: 151734835
Skills:
ticketing systems , workflow engines , Logging, Authentication, Python, Apis, Containers, tool calling, CI CD, Memory Design, agent architecture, enterprise APIs, multi-agent workflows, identity systems, orchestration frameworks, production observability, data platforms, cloud platforms, model routing, operational tools, LLM agentic AI, RAG, ML automation, backend services, code repositories
Skills:
Power Automate, Sql, Clustering, Azure Machine Learning, Python, Power Apps, retrieval-augmented generation, anomaly detection, time-series forecasting, Classification, prompt engineering, Regression, Microsoft Power Platform, large language models, Azure AI services, Copilot Studio, machine learning models
Skills:
MLops, Pytorch, Docker, Kubernetes, Python, KV caching, batching strategies, quantization, vLLM
Skills:
Java, Golang, C, Python, LangChain, RAGAS, RAG evaluation frameworks, Pinecone, AutoGen, Vector Databases, Vibe Coding, Milvus, ChromaDB, LlamaIndex, Prompt Engineering, RAG architectures, ETL pipelines
Skills:
Machine Learning, Data Science, Git, MLops, Docker, Databricks, Azure Machine Learning, Kubernetes, Python, LangChain, GenAI, LLMs, vector databases, LangGraph, AWS SageMaker, RAG, cloud-based AI platforms, AI agent frameworks, model monitoring