Role : AI Ops Engineer
Company : R Systems Direct payroll with leading reputed MNC
Permanent Opportunity
Singapore Business Hours
We are seeking an experienced AI Ops Engineer to support the deployment, monitoring, reliability, and optimization of enterprise AI platforms. The ideal candidate will have strong cloud, DevOps, and MLOps experience with expertise in maintaining production AI workloads.
Be a Part of Something BIG!
- Apply AI, ML, and LLM-based tools to solve real-world operational and reliability challenges
- Identify, prototype, and validate AI-driven improvements to day-to-day operations using clear, measurable metrics
- Operationalise successful AI use cases into reliable, scalable, and maintainable production capabilities
- Improve system observability, incident detection, root-cause analysis, and operational efficiency through automation and intelligence
- Support the end-to-end lifecycle of AIOps, MLOps, and LLMOps solutions, from experimentation to production
- Ensure operational reliability, robustness, and maintainability of AI-enabled systems
- Collaborate closely with senior engineers and operations teams to align solutions with business and operational needs
- Continuously evaluate emerging AI technologies and tools for practical applicability and impact
- Embed best practices for monitoring, governance, and responsible use of AI in operational environments
Make an Impact by:
- Explore and evaluate AI, ML, and LLM-based tools that can improve operational efficiency and visibility
- Build small proof-of-concepts or prototypes to validate ideas in operational contexts
- Measure outcomes using simple, clearly defined metrics
- Analyze logs, metrics, traces, and events to identify anomalies and automation opportunities
- Assist in deploying, monitoring, and maintaining ML or LLM-enabled solutions
- Automate repetitive operational tasks using scripts, workflows, or lightweight services
- Work with senior engineers to turn validated ideas into stable, scalable capabilities
- Document solutions, trade-offs, and lessons learned for reuse by other teams
- Stay informed on relevant AI and open-source tooling, and suggest ideas worth testing
Skills for Success:
- Bachelor's or Master's degree in Computer Science or a related field
- 1–2 year experience:
- Python, SQL, Apache Spark, or similar languages
- Data Wrangling, Analysis and Visualisation
- Containers and deployment fundamentals (Docker; Kubernetes)
- Version Control (Git)
- Testing (pytest)
- MLOps/GenAI Ops tools
Internships, graduate projects, hackathons, open-source contributions, or relevant work experience in:
- Software development
- Automation
- AI / ML / GenAI-related project
Familiarity with
- Software engineering best practices, including OO/Functional design, reproducibility, and testing
- Logs, metrics, and monitoring concepts
- ML/GenAI tooling (LLM, RAG, LlamaGuard, Presidio, MLFlow, LangFuse etc)
- Monitoring and observability stacks (Prometheus, Grafana, OpenTelmetry)
- Hands-on experience with cloud ML services (Databricks, Azure ML, etc.)
- Hands-on, curious, and proactive
- Pragmatic and focused on operational impact
- Disciplined about scaling what works and discarding what doesn't
- Collaborative and able to communicate ideas clearly
- Self-driven and proactive, comfortable working in a fast-paced environment
- Familiarity with AI and data development process in telco environment