
Search by job, company or skills
We are looking for a Site Reliability Engineer (SRE) with experience in platform engineering, DevOps, and production operations. The role involves building reliable systems, automating infrastructure, and ensuring observability across mission-critical applications. You will work closely with development, infrastructure, and operations teams to design scalable solutions, improve system resilience, and define operational best practices.
System reliability: Ensure high availability, performance, and resilience of production systems.
Infrastructure automation: Build and maintain CI/CD pipelines, automate deployments, and manage containerized workloads.
Observability & monitoring: Implement logging, metrics, tracing, alerting, and dashboards using modern monitoring tools.
Incident response: Integrate systems with enterprise monitoring, SIEM, and incident management workflows participate in on-call rotations.
Operational excellence: Define and implement runbooks, escalation paths, and production support models.
Collaboration: Work with cross-functional teams in Agile environments to deliver reliable and secure systems.
Continuous improvement: Conduct root cause analysis, implement permanent fixes, and drive automation initiatives to reduce manual effort.
5+ years of experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.
Solid knowledge of Linux administration, Shell scripting, Git, CI/CD, Docker, and Kubernetes.
Hands-on experience with observability platforms (Grafana, Prometheus, ELK, CloudWatch, Azure Monitor, etc.).
Experience integrating systems with enterprise monitoring, alerting, SIEM, or incident response workflows.
Proven ability to define and implement runbooks, operational procedures, escalation paths, and production support models.
Familiarity with cloud platforms (AWS, Azure, GCP) and infrastructure-as-code tools (Terraform, Ansible).
Excellent problem-solving skills, ability to troubleshoot complex systems, and experience in Agile/DevOps practices.
Excellent communication skills and ability to collaborate with cross-functional stakeholders.
EA License : 02C3423
EA Personnel : R22108699
Job ID: 153341231
Skills:
Kibana, BigQuery, Elk, Apis, Prometheus, Grafana, Devops, Terraform, Itil Processes, DataFlow, Scripting, CI CD automation, monitoring and observability tools, Cloud Composer, IT risk and security concepts, SRE practices, Google Cloud technologies, cloud-native services
Skills:
Google Cloud, Jenkins, Docker, Azure, Kubernetes, AWS, cloud services products, application operations, automation operations, GitLab CI, cloud platform deployment and management, container technologies
Skills:
Java, Nginx, Hadoop, OpenStack, Shell, Gcp, Docker, Linux, Ansible, Spark, Python, Kubernetes, AWS, Volcano Engine, Go, Flink, Aliyun
Skills:
Java, Golang, Cloudformation, Prometheus, Grafana, Git, Terraform, Ansible, Influxdb, Dynatrace, Splunk, Azure, Python, Kubernetes, AWS, GenAI
Skills:
AWS (CloudWatch, CloudTrail), Azure (Monitor, App Insights), Terraform, Ansible, Python, Datadog, Dynatrace, Splunk, Elk, CI/CD, X-Ray, Log Analytics