Search Jobs

Search by job, company or skills

Observability Engineer

Observability Engineer

u3 infotech pte. ltd.
5-7 Years
SGD 7,000 - 12,000 per month
Early Applicant
  • Posted 21 days ago
  • Be among the first 10 applicants

Job Description

Key Responsibilities:

Observability Strategy and Governance

  • Define and own Enterprise Observability Architecture aligned with operational resilience mandates (MAS TRM, DORA, APRA CPS 230).
  • Deploy and optimize observability platforms (Datadog, Dynatrace, Splunk) for full-stack visibility across infra, application, network, and user experience.
  • Establish governance standards for telemetry data (metrics, logs, traces), ensuring consistency, retention compliance, and security controls.
  • Integrate observability platforms with incident management, ITSM, and AIOps systems for predictive alerting and anomaly detection.

Reliability Engineering and Automation

  • Implement SRE frameworks for infrastructure and business-critical applications.
  • Automate runbooks, alerts, self-healing actions, and auto-remediation workflows via Python, Ansible, and Terraform.
  • Partner with Application, Infrastructure, and Cyber teams to codify operational reliability into the delivery lifecycle.
  • Conduct resilience testing, chaos engineering, and capacity validation.
  • Develop error budget policies and reliability scorecards for key production services.

Cloud Observability and Platform Engineering

  • Architect and manage observability for Cloud-native workloads in AWS and Azure.
  • Integrate cloud observability into landing zones and CI/CD pipelines for continuous compliance.
  • Implement IaC models using Terraform and Ansible for consistent, auditable provisioning.
  • Collaborate with Cloud, DevOps, and Security teams on real-time telemetry aligned to audit requirements.

Operational Excellence and Stakeholder Management

  • Drive reduction in incident recurrence, MTTR, and manual intervention through observability-led automation.
  • Deliver executive dashboards highlighting availability, reliability KPIs, and operational risk indicators.
  • Act as technical advisor to senior management during major incidents, post-incident reviews, and audits.

Skillset Requirements:

  • At least 5 years of experience in Infrastructure, Cloud, or Site Reliability Engineering (SRE) related roles, with minimum 3 years in an SRE SME capacity, ideally within financial institutions or regulated environments.
  • Hands-on expertise with Observability Platforms: Datadog, Dynatrace, Splunk, ELK.
  • Hands-on expertise with Automation/IaC: Terraform, Ansible, Python, CI/CD tools.
  • Hands-on expertise with Cloud Platforms: AWS (CloudWatch, X-Ray, CloudTrail), Azure (Monitor, Log Analytics, App Insights).
  • Deep understanding of SRE principles, service health modelling, error budgets, and auto-remediation design.
  • Familiarity with financial sector operational resilience frameworks, regulatory compliance, and incident governance.
  • A good team player with excellent written and verbal communication skills, able to coordinate across diverse stakeholders.
  • Certification in at least one of the following required:
  • Datadog Certified Observability Professional / Dynatrace Certified Associate
  • Terraform/Ansible/Python Certified Expert
  • Certification in the following will be advantageous:
  • AWS Certified DevOps Engineer / Azure DevOps Expert
  • SRE Foundation/Practitioner (DevOps Institute)
  • ITIL v4 Managing Professional

    U3 Privacy Notice:
    Please refer to U3's Privacy Notice for Job Applicants/Seekers at https://u3infotech.com/privacy-notice-job-applicants/. When you apply, you voluntarily consent to the collection, use and disclosure of your personal data for recruitment/employment and related purposes.

    Key Skills

    Similar Jobs

    3-5 yrs
    SGD 7,000 - 14,000 per month
    Singapore
    Skills:
    Dashboards, Mq, Kafka, metrics, Docker, Terraform, Ansible, ECS, Azure, Logging, AWS, Alerting, Incident Troubleshooting, Observability Engineering, OpenTelemetry, Infrastructure Engineering, tracing, Platform Engineering
    3-5 yrs
    SGD 7,000 - 9,000 per month
    Singapore
    Skills:
    Mq, Prometheus, Kafka, Grafana, Cloud, Docker, Terraform, Ansible, ECS, Dynatrace, AWS-native observability and monitoring capabilities, SHIP-HATS, Instrumentation, CI CD, Elastic, Cloud-native storage and telemetry data services, OpenTofu, event-driven telemetry patterns, OpenTelemetry, Event Telemetry, Data Storage, Observability, Infrastructure
    3-5 yrs
    Singapore
    Skills:
    Dashboards, Mq, Kafka, Terraform, Docker, metrics, Ansible, ECS, Azure, AWS, Logging, Alerting, Incident Troubleshooting, Observability Engineering, OpenTelemetry, Infrastructure Engineering, tracing, Platform Engineering
    3-5 yrs
    SGD 8,500 - 9,500 per month
    Singapore
    Skills:
    Docker, Terraform, Ansible, Prometheus, Dynatrace, Grafana, AWS, Elastic, OpenTofu, OpenTelemetry, CI CD
    3-5 yrs
    Singapore
    Skills:
    Python Scripting, Kibana, Linux, Elasticsearch, Logstash, Elk Stack, Rest Apis, Sdlc, AWS monitoring services