Site Reliability Engineer
- Posted a day ago
- Be among the first 10 applicants
Job Description
About the Role
Join an international backend team to own the stability, observability, and automation of the region's production environment. Our client's smart hardware and mobile applications run across multiple AWS regions. You will hold operational access to the regional production environment, handle incidents alongside the regional tech lead, and work with the HQ infrastructure team on containerization, unified gateway, and middleware governance initiatives — giving the region independent release, monitoring, and emergency-response capability.
Responsibilities
- Regional infrastructure operations: run day-to-day operations, capacity planning, and cost optimization for the region's AWS resources and Kubernetes clusters; keep environments consistent and configuration traceable
- Observability: build and maintain dashboards, alert rules, and log platforms for applications, middleware (MySQL / Redis / Kafka / RocketMQ), and infrastructure; ensure alerts are accurate and actionable
- Incident response and drills: take part in regional incident handling; execute and automate mitigation actions (rollback, scaling, degradation switches); run degradation and failure drills and maintain incident runbooks
- Release and change management: build and maintain CI/CD pipelines supporting canary releases and fast rollback; enforce change review and checklists; advance infrastructure-as-code
- Middleware and database operations: own backup, recovery, tuning, and capacity assessment for regional databases and middleware; work with developers on slow queries and performance issues
- Security and access: manage regional IAM accounts and access under company policy; support rollout of WAF, gateway, and other security capabilities; enforce least privilege and audit trails for data access
Requirements
- Bachelor's degree or above; Computer Science / Software Engineering preferred
- 5+ years in operations / SRE / platform engineering
- Proficient with AWS (EC2, ALB, EKS, RDS, ElastiCache, S3, IAM, VPC, CloudWatch), with multi-region operations experience
- Strong Kubernetes skills: cluster operations, troubleshooting, and resource governance
- Experience with Terraform or similar infrastructure-as-code tools, and CI/CD tools such as Jenkins / GitLab CI
- Solid Linux and networking fundamentals; scripting proficiency in at least one of Python / Shell / Go
- Experience deploying, monitoring, and troubleshooting MySQL, Redis, and Kafka / RocketMQ
- Familiar with Prometheus / Grafana / ELK or comparable observability stacks
- Experience handling production incidents and postmortems; able to execute mitigation calmly and methodically under pressure
- Rigorous and detail-oriented; committed to documented, auditable changes
- Must be legally authorized to work in the country of employment




