Search by job, company or skills

Arient Solutions

Site Reliability Engineer

Early Applicant
Quick Apply
  • Posted 13 days ago
  • Be among the first 10 applicants
5-8 Years
SGD 5,500 - 7,000 per month

Information Technology,

IT Infrastructure,

GovTech

Job Description

As a Site Reliability Engineer, you will be responsible for designing and operatingGitLab,AWSandKubernetes-based infrastructure and solutions that power our platform,to ensure the stability, scalability, and performance of our runtime platform.

Responsibilities:

As a Site Reliability Engineer, you will be responsible for:

Toil Reduction & Automation

  • Identify repetitive tasks and develop automation via CI/CD pipelines, ensuring integration with cross-functional teams to reduce manual intervention and improve operational efficiency.

Observability & System Health

  • Implement comprehensive observability solutions (logs, metrics, traces, alerts) around the four Golden Signals (latency, traffic, errors, saturation), and build automation for proactive system health assessments and self-remediation.

Production Support & Incident Management

  • Participate in on-call rotations, promptly respond to incidents to minimize MTTR, and conduct thorough post-incident reviews to implement preventive measures and improve system resilience.

Security & Compliance

  • Design and implement solutions that are secure and compliant by collaborating with dedicated security teams, conducting regular audits, and integrating advanced vulnerability scanning tools.

Maintenance, Optimisation & Performance

  • Identify and resolve performance bottlenecks and operational issues, define and track KPIs (e.g., MTTR, system uptime, cost efficiency), and drive ongoing optimisation efforts.

Strategic Customer Engagement

  • Act as a technical advisor for tenants, guiding them on containerization, and best practices for cloud-native deployments, and participating in strategic initiatives to enhance platform scalability and performance.

Knowledge Sharing & Documentation

  • Develop and maintain detailed playbooks, runbooks, and documentation to facilitate team-wide knowledge sharing, streamline incident response, and ensure that critical processes are well understood across the team.

Continuous Learning & Innovation

  • Stay current with the latest AWS, Kubernetes, and industry developments, and proactively recommend improvements and innovative solutions to maintain a competitive and reliable platform.

Requirements:

  • Proven experience as a Site Reliability Engineer or similar role, with a strong background in containerization, orchestration, and cloud-native technologies.
  • Proven ability to troubleshoot and resolve complex technical issues in containerized applications.
  • Demonstrated experience with incident management, including post-incident reviews and continuous improvement.
  • Strong documentation skills and experience in knowledge sharing across teams.
  • Deep understanding of AWS, Kubernetes (including AWS EKS), and operational best practices, with familiarity in multi-cloud or hybrid environments.
  • Solid grasp of networking, security, and storage in both AWS and Kubernetes contexts.
  • Experience integrating Kubernetes with AWS cloud technologies (e.g., Secrets Manager, Load Balancers) and using infrastructure-as-code (Terraform or similar).
  • Hands-on experience with containerization tools (Kubernetes, Kustomize, Helm) and automation scripting (Go, Python, Bash, or equivalent).
  • Ability to write and maintain automated tests or conduct thorough manual testing for automation scripts, ensuring the reliability and effectiveness of automated solutions.
  • Familiarity with CI/CD tools (GitLab CI/CD, ArgoCD) and version control systems (Git).
  • Experience with observability/monitoring tools (Prometheus, Grafana, ELK Stack) and defining SLOs and Error Budgets.
  • Certifications such as Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) are a plus.
  • Experience with developing Kubernetes operators using Go, service mesh technologies, and Chaos Engineering is a plus.

Date Posted: 18/09/2025

Job ID: 126099275

Report Job

About Company

Founded in 2013, Arient Solutions is an independent specialized recruiting & staffing firm headquartered in Tirunelveli, Tamil Nadu. We chip in as your HR partner in providing an array of HR related services. Our success is forged upon our personalized, long-term relationships with both our clients and candidates together with an underlying knowledge of the sectors we operate in. We are now a leading Human Resource Employment Services Company with proven track record in recruiting candidates for a wide range of industries and job roles.

User Avatar
0 Active Jobs
View More
Last Updated: 28-09-2025 02:56:12 PM
Home Jobs in Singapore Site Reliability Engineer

Similar Jobs