Search Jobs

Search by job, company or skills

Senior Observability Engineer (AI GPU Cloud) - Telemetry/Prometheus/Cutting-edge technology

Senior Observability Engineer (AI GPU Cloud) - Telemetry/Prometheus/Cutting-edge technology

dada consultants
Fresher
  • Posted 18 hours ago
  • Be among the first 10 applicants

Job Description

About the Role

Our client is a large-scale AI cloud infrastructure provider that offers GPU compute capacity, high-performance training and inference infrastructure. As a Senior Observability Engineer, you will own the design and scaling of the company's entire telemetry stack — from hardware-level metrics collection through to multi-tenant billing pipelines.

Key Responsibilities

  • Architect and scale a high-throughput observability platform capable of ingesting and querying extremely high-cardinality metrics reliably across a large, distributed compute environment.
  • Integrate low-level hardware telemetry (accelerator health, network fabric statistics, and out-of-band system management data) into a unified cluster-wide monitoring layer.
  • Build custom kernel-level diagnostic tooling to trace network congestion, I/O latency, and distributed workload bottlenecks across large compute clusters.
  • Develop automated dashboards and alerting pipelines that proactively identify and isolate degraded hardware before it impacts running workloads.
  • Design metering pipelines that support accurate, multi-tenant usage-based billing derived from real-time compute and network utilization data.
  • Partner with infrastructure and scheduling teams to define observability standards for large-scale AI/ML workloads, ensuring visibility into job-level efficiency and resource usage.
  • Lead architecture and design reviews for the observability stack, mentoring engineers on best practices for high-performance telemetry collection and analysis.

Requirements

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related discipline
  • Software or site reliability engineering experience, with expertise in Prometheus/OpenTelemetry
  • Advanced proficiency in Go, with substantial experience in Kubernetes metric exporters and operators
  • Familiarity with kernel-level tracing tools (eBPF, BCC) and performance tuning of Linux systems at scale
  • Strong familiarity with AI hardware performance metrics (GPU power states, compute utilisation, memory bandwidth) and high-performance network telemetry

If you are passionate about technology and meet the above requirements, please don't hesitate to apply. Please note that only shortlisted candidates will be contacted. Appreciate your understanding. Data provided is for recruitment purposes only.

Dada Consultants Pte Ltd

Website: www.dadaconsultants.com

EA License No.: 18S9037

Business Registration Number: 201735941W

Key Skills

BCC

OpenTelemetry

Multi-tenant billing pipelines

Kernel-level tracing tools

eBPF

Telemetry collection

High-performance network telemetry

Kubernetes metric exporters

About Company