

Search by job, company or skills
About the Company
We are working with A1, a company incubated and backed by BJAK, whose mission is to build the next generation of AI-native applications that fundamentally change how people communicate and get things done. A1's first application, AI Email Triage, reimagines email by moving users from reading and writing emails to learning from and approving AI-completed work-- making it efficient, smart and delightful.
A1's core capabilities are Agentic AI - can reason through multi-step workflows and use external tools to complete tasks; Permission-Based Actions - always asks for approval before taking actions such as sending emails or updating your calendar, keeping users in control; Context & Memory - remembers user preferences and past context to deliver increasingly personalised, accurate, and consistent assistance over time.
BJAK is the largest insurance platform in Southeast Asia with presence in Japan, United Kingdom and growing.
Role Summary
As an ML Platform Engineer, you will build the infrastructure and systems that power A1's AI capabilities. You will design and operate the systems behind the AI stack, from model training and evaluation to deployment, inference, observability, and continuous improvement. You will work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, and cost-efficient production systems. You will build the platforms, tooling, and infrastructure that enable the team to experiment quickly and bring AI capabilities to production with confidence.
Responsibilities
Key Performance Indicators
Our Ideal Candidate
Job ID: 152963943
Skills:
Git, Agile Scrum, Apache Spark, Python, Generative AI, AWS Bedrock, RAG architectures
Skills:
AWS, Python, Azure, Gcp, MLops, Spark, Ray, Dask
Skills:
triton , multi-threading, Tensorflow, memory pools, template programming, vLLM, TensorRT-LLM, XLA, CUDA programming, kernel scheduling, lock optimization, GDB debugging, operator fusion, memory access optimization, thread pools, graph optimization, GPU optimization, Warp execution models, RPC frameworks, performance profiling, TensorRT, VRAM scheduling