Data & AI Operations Engineer #AIDA
Be a Part of Something BIG!
- Apply AI, ML, and LLM-based tools to solve real-world operational and reliability challenges
- Identify, prototype, and validate AI-driven improvements to day-to-day operations using clear, measurable metrics
- Operationalise successful AI use cases into reliable, scalable, and maintainable production capabilities
- Improve system observability, incident detection, root-cause analysis, and operational efficiency through automation and intelligence
- Support the end-to-end lifecycle of AIOps, MLOps, and LLMOps solutions, from experimentation to production
- Ensure operational reliability, robustness, and maintainability of AI-enabled systems
- Collaborate closely with senior engineers and operations teams to align solutions with business and operational needs
- Continuously evaluate emerging AI technologies and tools for practical applicability and impact
- Embed best practices for monitoring, governance, and responsible use of AI in operational environments
Make an Impact by:
- Explore and evaluate AI, ML, and LLM-based tools that can improve operational efficiency and visibility
- Build small proof-of-concepts or prototypes to validate ideas in operational contexts
- Measure outcomes using simple, clearly defined metrics
- Analyze logs, metrics, traces, and events to identify anomalies and automation opportunities
- Assist in deploying, monitoring, and maintaining ML or LLM-enabled solutions
- Automate repetitive operational tasks using scripts, workflows, or lightweight services
- Work with senior engineers to turn validated ideas into stable, scalable capabilities
- Document solutions, trade-offs, and lessons learned for reuse by other teams
- Stay informed on relevant AI and open-source tooling, and suggest ideas worth testing
Skills for Success:
- Bachelor's or Master's degree in Computer Science or a related field
- 1-2 year experience:
- Python, SQL, Apache Spark, or similar languages
- Data Wrangling, Analysis and Visualisation
- Containers and deployment fundamentals (Docker Kubernetes)
- Version Control (Git)
- Testing (pytest)
- MLOps/GenAI Ops tools
- Internships, graduate projects, hackathons, open-source contributions, or relevant work experience in:
- Software development
- Automation
- AI / ML / GenAI-related project
- Familiarity with
- Software engineering best practices, including OO/Functional design, reproducibility, and testing
- Logs, metrics, and monitoring concepts
- ML/GenAI tooling (LLM, RAG, LlamaGuard, Presidio, MLFlow, LangFuse etc)
- Monitoring and observability stacks (Prometheus, Grafana, OpenTelmetry)
- Hands-on experience with cloud ML services (Databricks, Azure ML, etc.)
- Hands-on, curious, and proactive
- Pragmatic and focused on operational impact
- Disciplined about scaling what works and discarding what doesn't
- Collaborative and able to communicate ideas clearly
- Self-driven and proactive, comfortable working in a fast-paced environment
- Familiarity with AI and data development process in telco environment