Search by job, company or skills

Site Reliability Engineer

5-8 Years
  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

We are hiring for the Site Reliability Engineer (SRE) role on behalf of our global MNC client in Singapore. This is a 12-month extendable contract with a hybrid work arrangement. If you have a strong background in platform engineering, DevOps, and production operations, this is an exciting opportunity to work on enterprise-scale platforms and ensure their reliability, scalability, and operational excellence.

As a Site Reliability Engineer, you will be responsible for deploying and operationalizing enterprise platform solutions across development, testing, and production environments. You will establish production-grade reliability by implementing logging, observability, monitoring, alerting, and operational resilience while integrating platforms with enterprise SIEM and Incident Response workflows. The role also involves managing environment configurations, release processes, deployment automation, and ensuring platforms meet enterprise standards for availability, scalability, security, and recoverability. You will support production readiness, troubleshoot platform and infrastructure issues, and work closely with cybersecurity, infrastructure, application development, and operations teams to deliver secure and highly available services.

The ideal candidate has 5–8 years of experience as a Site Reliability Engineer, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer, with hands-on experience managing production systems across Dev, QA, UAT, and Production environments. You have strong expertise in Linux, shell scripting, Git, CI/CD, Docker, Kubernetes, observability platforms, logging, monitoring, tracing, and alerting. You are experienced in integrating enterprise monitoring, SIEM, and Incident Response workflows, and have a solid understanding of production reliability practices including SLIs, SLOs, incident management, capacity planning, and post-incident reviews. Experience with cloud platforms (AWS, Azure, or Google Cloud), Infrastructure as Code tools such as Terraform, Helm, or Argo CD, identity and access management, secrets management, and automation using Python or Go will be highly valued. Candidates with experience operating enterprise-scale, high-availability platforms in regulated environments will be at an advantage.

If you have the relevant experience please apply with your updated resume. Only shortlisted candidates will be contacted.

More Info

Job Type:
Industry:
Employment Type:

Job ID: 152459625

Similar Jobs

Singapore

Skills:

GcpDockerCloud InfrastructureApache KafkaKubernetesAWSerror budgetsSLIsAI modelsdistributed tracinginfrastructure-as-codeAI-assisted engineering toolsSLOsstructured loggingapplication instrumentation

Singapore

Skills:

.NETJavaWiresharkPrometheusSpring BootGrafanaDatadogJenkinsTerraformDockerAnsibleECSDynatraceGitlabSplunkKubernetesPythonCorvil

Singapore

Skills:

BashDnsSSLGcpDockerAnsibleNetworking ProtocolsLuaLoad BalancersLinux InternalsMongoDBPuppetOracleTlsKubernetesPythonGolangAWSAliCloudDevOps practicesdistributed system componentsSpinnakerGit-based source control

Singapore, Beach Road

Skills:

LoggingAutomated TestingIncident ResponseDistributed Systemsbusiness continuity strategiesalerting systemsDisaster Recoveryreliability standardsMonitoringoperational runbookson-prem environmentsdeployment workflows

Singapore

Skills:

JavaGcpDistributed SystemsAzurePythonKubernetesAWSGo

Beware of Scammers

We don’t charge money for job offers