O
Kubernetes & Site Reliability Engineer (SRE)
O
Kubernetes & Site Reliability Engineer (SRE)
opensource technologies pte. ltd.- Posted 5 hours ago
- Be among the first 10 applicants
Job Description
Role Overview
- We are looking for experienced Kubernetes & Site Reliability Engineers to support highly scalable, business-critical production platforms for a global technology customer in Singapore.
- The role requires strong hands-on expertise in Kubernetes, Linux, production reliability, automation, observability, incident management and troubleshooting of distributed systems
- Candidates should be comfortable operating large-scale production environments where availability, performance, automation and operational excellence are critical.
Key Responsibilities
- Operate, maintain and troubleshoot large-scale Kubernetes-based production environments
- Ensure reliability, scalability, availability and performance of critical services.
- Investigate complex production issues and perform detailed root-cause analysis.
- Participate in incident response and drive permanent corrective actions.
- Automate repetitive operational activities and improve platform reliability.
- Build and improve monitoring, alerting, logging and observability frameworks.
- Define and track SLIs, SLOs and operational reliability metrics
- Support Kubernetes upgrades, configuration changes, patching and platform improvements.
- Work closely with application engineering, infrastructure, platform, security and DevOps teams.
- Perform capacity planning, performance tuning and reliability improvements.
- Develop and maintain operational runbooks, automation scripts and troubleshooting documentation.
- Participate in production readiness reviews and ensure applications meet operational standards.
Mandatory Skills
- Strong hands-on experience with
- Kubernetes administration and troubleshooting
- Strong understanding of Kubernetes architecture, including:
- Pods
- Deployments
- StatefulSets
- Services
- Ingress
- ConfigMaps / Secrets
- RBAC
- Storage
- Networking
- Strong
- Linux systems administration and troubleshooting skills.
- Good understanding of networking concepts such as DNS, TCP/IP, load balancing and service connectivity.
- Strong understanding of
- Site Reliability Engineering principles
- Experience supporting large-scale, high-availability production systems.
- Strong incident management and RCA experience.
- Hands-on scripting/automation experience using
- Python, Bash/Shell or similar
- Experience with monitoring and observability tools such as
- Prometheus, Grafana, Splunk, ELK/OpenSearch, Datadog or equivalent
- Helm or similar Kubernetes package/deployment management tools.
- GitOps experience using tools such as Argo CD or Flux.
- Knowledge of service mesh concepts.
- Experience with container security and Kubernetes security practices.
- Experience with cloud or private-cloud infrastructure.
- Familiarity with distributed systems and microservices architectures.
- Exposure to performance engineering and capacity management.
- Experience working in globally distributed engineering environments.
More Info
Key Skills
Linux systems administration
Microservices architectures
Root-cause analysis
Kubernetes security practices
Service mesh concepts
Container security
Cloud or private-cloud infrastructure
Kubernetes administration
