O
Kubernetes & Site Reliability Engineer (SRE)
O
Kubernetes & Site Reliability Engineer (SRE)
opensource pte. ltd.- Posted 6 hours ago
- Be among the first 10 applicants
Job Description
Role Overview
- We are looking for experienced Kubernetes & Site Reliability Engineers to support highly scalable, business-critical production platforms for a global technology customer in Singapore.
- The role requires strong hands-on expertise in Kubernetes, Linux, production reliability, automation, observability, incident management and troubleshooting of distributed systems.
- Candidates should be comfortable operating large-scale production environments where availability, performance, automation and operational excellence are critical.
Key Responsibilities
- Operate, maintain and troubleshoot large-scale Kubernetes-based production environments.
- Ensure reliability, scalability, availability and performance of critical services.
- Investigate complex production issues and perform detailed root-cause analysis.
- Participate in incident response and drive permanent corrective actions.
- Automate repetitive operational activities and improve platform reliability.
- Build and improve monitoring, alerting, logging and observability frameworks.
- Define and track SLIs, SLOs and operational reliability metrics.
- Support Kubernetes upgrades, configuration changes, patching and platform improvements.
- Work closely with application engineering, infrastructure, platform, security and DevOps teams.
- Perform capacity planning, performance tuning and reliability improvements.
- Develop and maintain operational runbooks, automation scripts and troubleshooting documentation.
- Participate in production readiness reviews and ensure applications meet operational standards.
Mandatory Skills
- Strong hands-on experience with Kubernetes administration and troubleshooting.
- Strong understanding of Kubernetes architecture, including:
- Pods
- Deployments
- StatefulSets
- Services
- Ingress
- ConfigMaps / Secrets
- RBAC
- Storage
- Networking
- Strong Linux systems administration and troubleshooting skills.
- Good understanding of networking concepts such as DNS, TCP/IP, load balancing and service connectivity.
- Strong understanding of Site Reliability Engineering principles.
- Experience supporting large-scale, high-availability production systems.
- Strong incident management and RCA experience.
- Hands-on scripting/automation experience using Python, Bash/Shell or similar.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, ELK/OpenSearch, Datadog or equivalent.
- Experience working with CI/CD and automated deployment environments.
- Strong debugging and problem-solving capabilities.
Preferred Skills
- Infrastructure-as-Code experience using Terraform, Ansible or equivalent.
- Helm or similar Kubernetes package/deployment management tools.
- GitOps experience using tools such as Argo CD or Flux.
- Knowledge of service mesh concepts.
- Experience with container security and Kubernetes security practices.
- Experience with cloud or private-cloud infrastructure.
- Familiarity with distributed systems and microservices architectures.
- Exposure to performance engineering and capacity management.
- Experience working in globally distributed engineering environments.
More Info
Key Skills
Reliability Requirements
Automated Operation Monitoring
systems reliability
Configuration Changes
