Job Description
We are seeking an experienced Platform Engineer / Site Reliability Engineer (SRE) to support the deployment, operations, and reliability of enterprise platform services. The ideal candidate will have strong hands-on experience with Kubernetes, observability platforms, CI/CD pipelines, and production support in large-scale environments.
Key Responsibilities
- Manage and support Kubernetes-based applications and platform services in production environments.
- Troubleshoot and resolve issues related to pods, deployments, services, ingress, scaling, and platform performance.
- Design, implement, and maintain monitoring, alerting, and observability solutions using Grafana, Prometheus, Splunk, or similar tools.
- Support CI/CD pipelines and deployment processes across development, testing, and production environments.
- Perform production incident management, root cause analysis (RCA), and service reliability improvements.
- Work closely with development, infrastructure, and security teams to ensure platform stability and availability.
- Automate operational tasks using scripting and infrastructure-as-code tools where applicable.
- Contribute to system capacity planning, performance optimization, and operational readiness.
Required Skills & Experience
- 5+ years of experience in Platform Engineering, SRE, DevOps, Infrastructure Engineering, or related roles.
- Strong hands-on experience with Kubernetes and containerized workloads.
- Experience with Helm, Kubernetes deployments, services, ingress, scaling, and troubleshooting.
- Good experience with Grafana, Prometheus, Datadog, Splunk, or equivalent monitoring and observability tools.
- Experience working with CI/CD tools such as Jenkins, GitLab CI/CD, GitHub Actions, ArgoCD, or similar.
- Strong Linux administration and troubleshooting skills.
- Experience with production support, incident management, and RCA.
- Exposure to cloud platforms such as AWS, Azure, or GCP.
- Working knowledge of scripting and automation using Python, Shell, or similar tools.
Preferred Skills
- Experience in banking, financial services, or other large-scale enterprise environments.
- Exposure to Infrastructure as Code tools such as Terraform or Ansible.
- Knowledge of SRE concepts including SLI, SLO, SLA, and Error Budgets.
- Experience with high-availability systems, observability platforms, and distributed systems.
Employment Type
- Full-time / Contract (as applicable)
Location
Experience
This version is concise enough for MCF posting while still attracting the exact profile type that the client has been responding positively to: Kubernetes + SRE + Monitoring + Production Support engineers rather than pure software developers.