Search by job, company or skills

Site Reliability Engineer

5-8 Years
SGD 8,000 - 10,000 per month
Early Applicant
  • Posted 6 hours ago
  • Be among the first 10 applicants

Job Description

Strong experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.

Hands-on experience deploying and operating production systems across multiple environments such as Dev, QA, UAT, and Prod.

Strong knowledge of Linux, shell scripting, Git, CI/CD, Docker, and Kubernetes.

Experience with observability platforms, logging, metrics, tracing, alerting, and dashboarding.

Experience integrating systems with enterprise monitoring, alerting, SIEM, or Incident Response workflows.

Experience defining and implementing runbooks, operational procedures, escalation paths, and production support models.

Strong troubleshooting skills across application, infrastructure, networking, container, and cloud layers.

Familiarity with production reliability practices such as SLIs, SLOs, incident management, capacity planning, failure handling, and post-incident review.

Experience working with APIs, gateways, microservices, service accounts, secrets management, and enterprise integration patterns.

Ability to collaborate with cybersecurity, AI engineering, infrastructure, application development, and operations teams.

More Info

Job Type:
Industry:
Function:
Employment Type:

Job ID: 151688251

Similar Jobs

Singapore

Skills:

Kubernetescloud networkingmicroservice architectureedge computingLoad Balance

Singapore

Skills:

JavaGolangCloudformationPrometheusGrafanaGitTerraformAnsibleInfluxdbDynatraceSplunkAzureKubernetesPythonAWSGenAI

Singapore

Skills:

LinuxautomationKubernetesScriptingmicroservices architectureobservability monitoring toolscloud-hosted applicationscontainer technologies

Singapore

Skills:

GithubSentryPostgreSQLPrometheusArtifactoryGrafanaJiraDatadogNew RelicJenkinsDockerLinuxAnsibleMySQLDynatraceMongoDBSplunkPythonKubernetes

Singapore

Skills:

LoggingIdentity And Access ManagementDockershell scriptingTerraformSiemAWSGoogle CloudGitLinuxIncident ManagementHelmAzureKubernetespost-incident reviewsSLIsInfrastructure as Code toolsautomation using PythonalertingCapacity PlanningIncident Response workflowsenterprise monitoringtracingsecrets managementcloud platformsMonitoringproduction reliability practicesArgo CDobservability platformsSLOs