Search by job, company or skills

Senior Platform & SRE Engineer (80 Platform / 20 SRE)

6-8 Years
Early Applicant
  • Posted 16 hours ago
  • Be among the first 10 applicants

Job Description


Senior Platform & SRE Engineer (80% Platform / 20% SRE)

What's on offer:

  • Good Salary + Equity
  • Health Insurance
  • Great team and Technology
  • Visa on offer for the right candidate

ABOUT THE CLIENT & COMPANY

Our client is a fast-growing, high-impact technology enterprise building secure AI infrastructure for scientific and engineering R&D teams. Their platform enables organizations to seamlessly connect complex data, machine learning models, and expert domain workflows in environments where reliability, strict data control, and total traceability matter.

With a highly ambitious engineering team distributed across Singapore and the United States, they are turning cutting-edge material science research and advanced technical computing into scalable, production-grade software.

ABOUT THE PLATFORM & TECH STACK

The platform is designed to deploy seamlessly across public clouds, customer-managed Virtual Private Clouds (VPCs), and custom infrastructure. Key highlights of the technical stack and architectural vision include:

  • Container Orchestration & Packaging: Built heavily on Docker and Kubernetes, utilizing custom Helm Charts for modular, reproducible deployments across diverse environments.
  • Infrastructure as Code (IaC): Core platform environments and automated provisioning managed systematically via Terraform.
  • Cloud & Infrastructure Strategy: Primary cloud footprint on AWS, with ongoing expansion into multi-cloud support (GCP, Azure) and an active roadmap to architect and build their own private cloud infrastructure.
  • AI & ML Integration: Microservices architecture supporting machine learning data pipelines, model registries, automated evaluation harnesses, and secure data-provenance tracking.
  • Security & Operational Control: Enterprise-grade security automation featuring Zero Trust IAM, secrets management, dynamic TLS certificate workflows, and complete telemetry for restricted or customer-managed environments.

ABOUT THE ROLE

This is a deeply hands-on senior engineering role with a primary focus on Platform Engineering (80%) complemented by Site Reliability Engineering (20%). You will own the release, deployment, and operational resilience path for the AI platform across secure enterprise environments.

You will build the core deployment tooling and automation that make every release fully versioned, reproducible, observable, and safe to operate-taking installations from initial proof-of-concept through to full-scale production.

WHAT YOU WILL DO

  • Platform Engineering (80% Focus):
  • Design, build, and maintain robust infrastructure and deployment automation using Terraform, Kubernetes, Docker, and Helm.
  • Architect deployment paths across AWS, hybrid cloud, and upcoming private cloud environments.
  • Own platform release engineering, including CI/CD pipelines, artifact versioning, and automated production readiness gates.
  • Build operational controls for customer-managed installations, covering configuration validation, credential handling, certificate workflows, and documented recovery paths.
  • Site Reliability Engineering (20% Focus):
  • Design health checks, validation suites, and automated recovery mechanisms for production deployments.
  • Build operational visibility into customer environments using structured logging, metrics, alerting, and distributed tracing.
  • Debug complex platform incidents, conduct root-cause analyses (RCA), and translate lessons learned into reusable platform automation.

WHAT WE ARE LOOKING FOR

  • Background: Bachelor's or Master's degree in Computer Science, Software Engineering, or a related technical discipline, with 6+ years of production software engineering experience.
  • Core Tech Mastery: Deep, practical expertise with Kubernetes, Docker, Helm Charts, and Terraform.
  • Cloud & Multi-Environment: Strong experience with AWS, with a solid understanding of cloud-native architecture. Comfort or interest in multi-cloud environments or bare-metal/private cloud infrastructure is a big plus.
  • Release & CI/CD: Proven track record of owning build, release, and CI/CD pipelines for multi-component distributed systems.
  • Security Controls: Familiarity with secrets management (e.g., Vault), dynamic certificate handling, IAM permissions, and securing enterprise software deployments.
  • Operational SRE Mindset: Strong debugging and incident-response skills using standard telemetry stacks (metrics, logs, traces).

NICE TO HAVE

  • Experience with private cloud architecture or hybrid cloud deployments.
  • Exposure to GitOps paradigms (e.g., ArgoCD, Flux) and progressive delivery rollout patterns.
  • Familiarity with packaging, versioning, or monitoring machine learning models in production (MLOps).
  • Experience building or scaling early-stage engineering functions in high-growth companies.
  • Material Science background, experience or interest.

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 151496425