N
Site Reliability Engineer (SRE) - Contract -
N
Site Reliability Engineer (SRE) - Contract -
nicoll curtin technology pte. ltd.Early Applicant
- Posted a day ago
- Be among the first 10 applicants
Job Description
Site Reliability Engineer (SRE)
Location: Singapore, Onsite
Employment Type: Contract (12 months renewable based project demand / performance)
Key Responsibilities
- Provide Site Reliability Engineering and production support for a developer platform, ensuring high availability, stability, and performance.
- Monitor platform health, investigate incidents, and coordinate timely resolutions with technical and business stakeholders.
- Perform root cause analysis and implement preventive actions to improve reliability and reduce recurring issues.
- Define and manage monitoring, alerting, SLIs, and SLOs for critical production services.
- Automate operational processes, enhance runbooks, and minimise manual support effort.
- Diagnose and resolve complex issues across applications, infrastructure, and cloud environments.
- Improve developer productivity and optimise engineering workflows.
- Participate in on-call support rotation for critical production services when required.
- Promote SRE, DevOps, and software engineering best practices across teams.
Required Skills & Experience
- 5-7 years of software engineering experience with JavaScript, Java, Python, or .NET.
- 2-4 years of experience supporting production systems within an SRE or production support environment.
- Minimum 3 years of AWS experience, including cloud infrastructure and platform operations.
- Hands-on expertise with Docker, Kubernetes, EKS, and Helm.
- Experience with Infrastructure as Code tools such as Terraform and CloudFormation.
- Strong knowledge of CI/CD pipelines, GitHub Actions, and artifact repositories such as JFrog.
- Proficiency in Linux administration and Shell scripting.
- Experience with observability and logging tools including CloudWatch, Splunk, and Datadog.
- Good understanding of incident management, service monitoring, alerting, SLIs/SLOs, and post-incident reviews.
- Strong analytical, troubleshooting, communication, and stakeholder management skills.
- Education: Degree in Computer Science, Information Technology, or a related discipline.
