Search Jobs

Search by job, company or skills

Kubernetes & Site Reliability Engineer (SRE)

Kubernetes & Site Reliability Engineer (SRE)

opensource pte. ltd.
5-8 Years
SGD 8,000 - 11,000 per month
  • Posted 6 hours ago
  • Be among the first 10 applicants

Job Description

Role Overview

  • We are looking for experienced Kubernetes & Site Reliability Engineers to support highly scalable, business-critical production platforms for a global technology customer in Singapore.
  • The role requires strong hands-on expertise in Kubernetes, Linux, production reliability, automation, observability, incident management and troubleshooting of distributed systems.
  • Candidates should be comfortable operating large-scale production environments where availability, performance, automation and operational excellence are critical.

Key Responsibilities

  • Operate, maintain and troubleshoot large-scale Kubernetes-based production environments.
  • Ensure reliability, scalability, availability and performance of critical services.
  • Investigate complex production issues and perform detailed root-cause analysis.
  • Participate in incident response and drive permanent corrective actions.
  • Automate repetitive operational activities and improve platform reliability.
  • Build and improve monitoring, alerting, logging and observability frameworks.
  • Define and track SLIs, SLOs and operational reliability metrics.
  • Support Kubernetes upgrades, configuration changes, patching and platform improvements.
  • Work closely with application engineering, infrastructure, platform, security and DevOps teams.
  • Perform capacity planning, performance tuning and reliability improvements.
  • Develop and maintain operational runbooks, automation scripts and troubleshooting documentation.
  • Participate in production readiness reviews and ensure applications meet operational standards.

Mandatory Skills

  • Strong hands-on experience with Kubernetes administration and troubleshooting.
  • Strong understanding of Kubernetes architecture, including:
  • Pods
  • Deployments
  • StatefulSets
  • Services
  • Ingress
  • ConfigMaps / Secrets
  • RBAC
  • Storage
  • Networking
  • Strong Linux systems administration and troubleshooting skills.
  • Good understanding of networking concepts such as DNS, TCP/IP, load balancing and service connectivity.
  • Strong understanding of Site Reliability Engineering principles.
  • Experience supporting large-scale, high-availability production systems.
  • Strong incident management and RCA experience.
  • Hands-on scripting/automation experience using Python, Bash/Shell or similar.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, ELK/OpenSearch, Datadog or equivalent.
  • Experience working with CI/CD and automated deployment environments.
  • Strong debugging and problem-solving capabilities.

Preferred Skills

  • Infrastructure-as-Code experience using Terraform, Ansible or equivalent.
  • Helm or similar Kubernetes package/deployment management tools.
  • GitOps experience using tools such as Argo CD or Flux.
  • Knowledge of service mesh concepts.
  • Experience with container security and Kubernetes security practices.
  • Experience with cloud or private-cloud infrastructure.
  • Familiarity with distributed systems and microservices architectures.
  • Exposure to performance engineering and capacity management.
  • Experience working in globally distributed engineering environments.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Similar Jobs

5-8 yrs
SGD 8,000 - 11,000 per month
Singapore
Skills:
Capacity Management, Incident Management, Performance Engineering, Distributed Systems, Helm, Linux systems administration, Microservices architectures, Root-cause analysis, Kubernetes security practices, Service mesh concepts, Container security, Cloud or private-cloud infrastructure, Kubernetes administration