Date: 28 Jul 2026
Company: Singapore Pools (Pte) Ltd
Work that powers communities.
Who We Are
Singapore Pools was established by the Singapore government on 23 May 1968 to provide safe and trusted betting to counter illegal gambling. As a not-for-profit organisation, it makes contributions to the Tote Board to fund a wide range of causes in social service, community development, sports, arts, education and health sectors.
Since 2004, over $5 billion have been channelled to the Tote Board. In addition, Singapore Pools also contributes about $2 billion annually to the Government in the form of taxes and duties. Its responsible gaming practices have been awarded the highest level of certification (Level 4) by the World Lottery Association's Responsible Gaming Framework since 2012.
Since inception, Singapore Pools staff have a long-standing commitment to doing good and giving back to those in need. Staff volunteers support activities held all year round, from helping disadvantaged children, youth-at-risk, underprivileged families, and elderly, to conserving the environment.
Job Purpose
The Site Reliability Engineer (SRE) drives enterprise operational resilience by architecting scalable cloud infrastructure, managing the enterprise DevOps toolchain, and ensuring centralized system observability across hybrid cloud environments. Leveraging a strong software engineering background, the SRE implements GitOps methodologies and develops custom automation applications and API-based services. By heavily utilizing Infrastructure as Code (IaC) and GenAI tools, this role eliminates manual operations, drives efficiency, optimizes cloud expenditures, and ensures maximum system uptime against stringent SLAs.
What You'll Do
- Software Engineering & Automation: Write code and develop automated, API-based applications (including GenAI integrations) to streamline operational reporting and eliminate recurring manual tasks.
- Reliability & Operations: Implement SRE best practices for observability, availability, performance, and incident response. Define, measure, and govern Service Level Objectives (SLOs) and Error Budgets in collaboration with product engineering teams. Participate in on-call rotations, execute postmortem/RCAs, and identify/fix production bottlenecks.
- Hybrid Cloud Infrastructure: Support product teams in building fault-tolerant applications by enforcing infrastructure deployment via rigorous IaC code reviews.
- CI/CD & Toolchain: Maintain DevOps systems and enforce GitOps workflows and integrate automated security/vulnerability scanning into CI/CD pipelines for all application, container, and IaC deployment approvals.
- Observability: Create and maintain consolidated operations dashboards by integrating telemetry from disparate monitoring tools and native AWS/Azure metrics.
- FinOps: Generate actionable FinOps reporting to track and optimize hybrid cloud spending.
- Resilience & Documentation: Execute disaster recovery (DR), backup, redundancy, and capacity planning strategies while maintaining high-quality runbooks and operational documentation.
Who You Are
- Degree-qualified in Computer Science, Engineering, Information Science or related IT Disciplin, with 4 to 7 years of proven, hands-on experience in software engineering, cloud architecture, DevOps, or SRE roles.
- Hold professional certifications such as ITIL, FinOps Certified Practitioner, AWS Certified Solutions Architect (Associate or Professional), AWS Certified DevOps Engineer (Professional), AWS Certified CloudOps Engineer (Associate), Microsoft Certified Azure Administrator Associate (AZ-104), Azure Solutions Architect Expert (AZ-305), or DevOps Engineer Expert (AZ-400).
- Cloud Architecture & Engineering: Deep hands-on experience building scalable hybrid cloud infrastructures (AWS and Azure) and containerization. Strong understanding of modern hosting, networking design patterns, and applying the Six Pillars of operational excellence across environments.
- Software Engineering: Strong background building production-level software in Python, Golang, or Java. Experience developing and deploying API-based services and serverless applications on AWS Lambda and Azure Functions.
- Automation & Infrastructure as Code (IaC): High proficiency in IaC and configuration management tools (e.g., Terraform, Ansible, CloudFormation) and container orchestration systems (e.g., Kubernetes). Proven capability in enforcing deployment pipelines through rigorous code review processes and working within Agile methodologies.
- Version Control: Proficient in Git, including advanced branching strategies and GitOps paradigms.
- Systems Knowledge: Deep expertise in architecting and managing centralized Observability platforms, utilizing GenAI, time-series databases, and diverse monitoring tools (e.g., Prometheus, Grafana, Dynatrace, Splunk, InfluxDB) to trace distributed cloud applications. Solid understanding of Database Administration and Networking architecture.
- Professional & Interpersonal Skills: Ability to make sound, logical, data-based decisions on complex issues while considering risks. Strong communication and interpersonal skills to collaborate and build relationships with internal and external stakeholders.
- Strong interest in technological trends and disruptions impacting Cloud Engineering and SRE.
What We Offer
- Comprehensive total rewards package
- Health & wellness benefits
- Continuous learning and upskilling opportunities
- Volunteerism and community initiatives
Only shortlisted candidates will be contacted for further career conversations.