Senior Observability Engineer (AI GPU Cloud) - Telemetry/Prometheus/Cutting-edge technology
dada consultants- Posted 18 hours ago
- Be among the first 10 applicants
Job Description
About the Role
Our client is a large-scale AI cloud infrastructure provider that offers GPU compute capacity, high-performance training and inference infrastructure. As a Senior Observability Engineer, you will own the design and scaling of the company's entire telemetry stack — from hardware-level metrics collection through to multi-tenant billing pipelines.
Key Responsibilities
- Architect and scale a high-throughput observability platform capable of ingesting and querying extremely high-cardinality metrics reliably across a large, distributed compute environment.
- Integrate low-level hardware telemetry (accelerator health, network fabric statistics, and out-of-band system management data) into a unified cluster-wide monitoring layer.
- Build custom kernel-level diagnostic tooling to trace network congestion, I/O latency, and distributed workload bottlenecks across large compute clusters.
- Develop automated dashboards and alerting pipelines that proactively identify and isolate degraded hardware before it impacts running workloads.
- Design metering pipelines that support accurate, multi-tenant usage-based billing derived from real-time compute and network utilization data.
- Partner with infrastructure and scheduling teams to define observability standards for large-scale AI/ML workloads, ensuring visibility into job-level efficiency and resource usage.
- Lead architecture and design reviews for the observability stack, mentoring engineers on best practices for high-performance telemetry collection and analysis.
Requirements
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related discipline
- Software or site reliability engineering experience, with expertise in Prometheus/OpenTelemetry
- Advanced proficiency in Go, with substantial experience in Kubernetes metric exporters and operators
- Familiarity with kernel-level tracing tools (eBPF, BCC) and performance tuning of Linux systems at scale
- Strong familiarity with AI hardware performance metrics (GPU power states, compute utilisation, memory bandwidth) and high-performance network telemetry
If you are passionate about technology and meet the above requirements, please don't hesitate to apply. Please note that only shortlisted candidates will be contacted. Appreciate your understanding. Data provided is for recruitment purposes only.
Dada Consultants Pte Ltd
Website: www.dadaconsultants.com
EA License No.: 18S9037
Business Registration Number: 201735941W
More Info
Key Skills
BCC
OpenTelemetry
Multi-tenant billing pipelines
Kernel-level tracing tools
eBPF
Telemetry collection
High-performance network telemetry
Kubernetes metric exporters
