You are the engineering backbone that enables Centience's teams to build fast and ship confidently. Your primary mission is to eliminate the gap between experimentation and production — designing and operating CI/CD pipelines, MLOps infrastructure, and automated deployment workflows with speed, reliability, and full audit traceability. Day-to-day, you work closely with the Head of Data Science & Analytics to productionise the machine learning models and AI products the business delivers. Your reporting will be to the MD & the Head of Data Science & Analytics.
Key Responsibilities
1. CI/CD & MLOps Pipeline Engineering
- Design, build, and maintain CI/CD pipelines for ML models, data products, and analytical applications using GitHub Actions, Jenkins, or AWS CodePipeline.
- Implement end-to-end MLOps workflows (MLflow, SageMaker Pipelines, or Kubeflow) covering experiment tracking, model versioning, registry, validation gating, and automated deployment.
- Automate model evaluation, shadow and A/B testing, and canary/blue-green deployments for zero-downtime releases with full rollback capability.
2. Infrastructure as Code & Environment Automation
- Own infrastructure-as-code (Terraform and/or AWS CDK) across dev, staging, and production with environment parity.
- Automate provisioning of ML compute environments (SageMaker Studio, JupyterHub, GPU clusters) so teams can spin up and tear down on demand.
- Manage containerisation and orchestration (Docker, ECS/EKS), environment isolation, and secrets management (AWS Secrets Manager, HashiCorp Vault).
3. Monitoring, Observability & Reliability
- Build and maintain observability stacks (CloudWatch, Grafana, Prometheus) for all pipelines, model endpoints, and platform services.
- Implement model drift detection, production data quality monitoring, and automated alerting for SLA breaches.
- Define and enforce SLOs/SLAs; lead incident response and post-mortems to eliminate repeat failures.
4. DevSecOps & Compliance Integration
- Embed security into every pipeline stage — SAST/DAST scanning, dependency and container image checks, and compliance policy enforcement.
- Ensure deployment artifacts are immutable, signed, and version-controlled, and collaborate with the Cloud Security Engineer to meet PDPA/GDPR and internal standards.
5. Collaboration & Enablement
- Act as the embedded platform engineer — translating model and pipeline requirements into production-grade infrastructure decisions.
- Reduce time-to-production via self-service deployment templates and runbooks, and train engineers on deployment best practice and production readiness.
Qualifications
- Bachelor's degree in Computer Science, Software Engineering, or related field.
- 5+ years of DevOps/MLOps engineering experience; at least 2 years supporting ML or data science workloads in production.
- Deep proficiency in IaC (Terraform, AWS CDK, CloudFormation) and CI/CD platforms (GitHub Actions, Jenkins, AWS CodePipeline).
- Hands-on experience with ML platforms (SageMaker Pipelines, MLflow, Kubeflow) and containerisation/orchestration (Docker, Kubernetes/EKS, ECS).
- Proficiency in Python and Bash, and experience with monitoring tools (CloudWatch, Prometheus, Grafana, Datadog).
- AWS DevOps Engineer Professional or ML Specialty preferred; familiarity with data engineering stacks (Airflow, dbt, Spark) a strong plus.