Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)
Argyll Scott5-7 Years
- Posted 12 hours ago
- Be among the first 10 applicants
Job Description
1-year contract on our client's payroll
We are partnering with a leading organisation in Singapore to hire a Site Reliability Engineer to provide SRE and production support for their internal developer platform. You will help keep the platform available, stable and performant, while driving automation to reduce manual operational effort.
Key Responsibilities
- Monitor platform health and respond to alerts across production services
- Troubleshoot and resolve incidents, coordinating timely resolution and clear updates to technical and business stakeholders
- Lead root-cause analysis and post-incident follow-up
- Improve service reliability through SLIs/SLOs, alert tuning and observability enhancements
- Automate repetitive operational tasks and maintain runbooks
- Optimise developer workflows and enhance developer experience
- Participate in on-call rotation for critical production services where required
Requirements
- Degree in Computer Science, Information Technology or a related field
- 5 to 7 years of software engineering experience in at least one of JavaScript, Java, Python or .NET
- 2 to 4 years of hands-on SRE or production support experience
- 3+ years of AWS experience (certifications advantageous)
- 3+ years with Docker, Kubernetes, EKS and Helm (certifications advantageous)
- Hands-on experience with Terraform and/or CloudFormation
- Proficiency in CI/CD workflows and GitHub Actions
- Knowledge of artifact repositories such as JFrog
- Strong Linux administration and Shell scripting skills
- Experience with CloudWatch, Splunk and Datadog
- Ability to diagnose complex issues across multiple technology layers
- Strong communication, collaboration and problem-solving skills, with an agile, fast-learning mindset

