
Search by job, company or skills
Strong experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.
Hands-on experience deploying and operating production systems across multiple environments such as Dev, QA, UAT, and Prod.
Strong knowledge of Linux, shell scripting, Git, CI/CD, Docker, and Kubernetes.
Experience with observability platforms, logging, metrics, tracing, alerting, and dashboarding.
Experience integrating systems with enterprise monitoring, alerting, SIEM, or Incident Response workflows.
Experience defining and implementing runbooks, operational procedures, escalation paths, and production support models.
Strong troubleshooting skills across application, infrastructure, networking, container, and cloud layers.
Familiarity with production reliability practices such as SLIs, SLOs, incident management, capacity planning, failure handling, and post-incident review.
Experience working with APIs, gateways, microservices, service accounts, secrets management, and enterprise integration patterns.
Ability to collaborate with cybersecurity, AI engineering, infrastructure, application development, and operations teams.
Job ID: 151688251
Skills:
Kubernetes, cloud networking, microservice architecture, edge computing, Load Balance
Skills:
Java, Golang, Cloudformation, Prometheus, Grafana, Git, Terraform, Ansible, Influxdb, Dynatrace, Splunk, Azure, Kubernetes, Python, AWS, GenAI
Skills:
Linux, automation, Kubernetes, Scripting, microservices architecture, observability monitoring tools, cloud-hosted applications, container technologies
Skills:
Github, Sentry, PostgreSQL, Prometheus, Artifactory, Grafana, Jira, Datadog, New Relic, Jenkins, Docker, Linux, Ansible, MySQL, Dynatrace, MongoDB, Splunk, Python, Kubernetes
Skills:
Logging, Identity And Access Management, Docker, shell scripting, Terraform, Siem, AWS, Google Cloud, Git, Linux, Incident Management, Helm, Azure, Kubernetes, post-incident reviews, SLIs, Infrastructure as Code tools, automation using Python, alerting, Capacity Planning, Incident Response workflows, enterprise monitoring, tracing, secrets management, cloud platforms, Monitoring, production reliability practices, Argo CD, observability platforms, SLOs