THE ROLE
The Client is engaging a senior data scientist on a professional-services basis - from AWS Professional Services or an AWS Partner - to work full-time, fully on-site, as part of the Client team. The role carries two sequential mandates, which are also its two headline KPIs:
- KPI 1 - Delivery assurance during the build. Ensure the appointed vendor delivers to industry best practice - best-in-class algorithm design, training pipelines, deployment architecture and documentation - challenging and validating every material technical decision on the Client's behalf while contributing hands-on to the build.
- KPI 2 - Production ownership after go-live. Take over the system at go-live and own 24/7 production support - operations, monitoring, retraining, bug fixes and enhancements - while enabling the Client's internal team toward self-sufficiency.
The role demands equal parts modeling depth (contextual bandits in production), AWS engineering and MLOps depth (SageMaker AI, deployment, operations), and the professional spine to challenge a top-tier delivery vendor constructively and hold its output to standard.
KEY RESPONSIBILITIES
Phase 1 - Delivery assurance alongside the appointed vendor (build phase)
- Algorithm assurance: challenge and validate the bandit formulation - action space, reward design, exploration/exploitation policy, feature set, cold-start handling and off-policy evaluation (IPS, doubly robust) - and require quantitative model-validation evidence before any model is promoted.
- Training-pipeline assurance: hold the vendor to reproducible, fully versioned pipelines (data, features, code, models) with automated evaluation gates and Model Registry promotion discipline; no manual, unrepeatable steps anywhere in the path to production.
- Deployment assurance: review the serving architecture against the latency budget at peak traffic (load testing), auto-scaling behavior, canary/shadow deployment, one-step rollback, and verified fail-safe behavior of the rules-based fallback.
- Engineering standards: enforce code review, automated testing, infrastructure-as-code, least-privilege security (IAM, KMS, VPC) and cost-awareness across everything delivered into the Client's AWS accounts and repositories.
- Hands-on co-build: contribute directly to feature engineering, SageMaker AI training and inference workloads, data pipelines and Action API integration - this role is a builder and reviewer, not an observer.
- Experimentation: co-design the A/B tests (baselines, power analysis, guardrail metrics, measurement windows) and independently validate the measured conversion and AOV uplift.
- Knowledge absorption: require complete documentation as work is produced - not only at the end; participate in all sprint ceremonies; verify every handover deliverable (code, models, pipelines, runbooks) against jointly agreed acceptance criteria before takeover.
Phase 2 - 24/7 production ownership (post go-live)
- 24/7 support: own round-the-clock production support under agreed SLAs - incident detection and response, root-cause analysis, and bug fixes through to verified resolution.
- Operations: run monitoring and alerting (SageMaker Model Monitor, CloudWatch), drift detection, retraining pipelines, model refresh cadence and champion-challenger releases; manage availability, latency and cloud cost.
- Enhancements: deliver a continuous stream of improvements - new behavioral features and signals, reward refinement, per-surface and per-market policy tuning - each validated through experimentation before full rollout.
- Capability enablement: train and mentor the Client's internal engineering and data science team hands-on; keep documentation and runbooks current; progress to an agreed milestone at which the Client operates independently.
REQUIRED QUALIFICATIONS
- Contextual bandit expertise (core). Demonstrable, shipped experience designing and operating contextual multi-armed bandit systems in production - e.g., LinUCB, Thompson Sampling, neural contextual bandits - including reward design, exploration policies, delayed/sparse rewards, cold-start strategies and off-policy evaluation. Recommender-system experience alone is not sufficient.
- Amazon SageMaker AI mastery (core). Expert, current, end-to-end command of the SageMaker AI platform: Studio, training and processing jobs, real-time inference endpoints with auto-scaling, SageMaker Pipelines, Model Registry, Feature Store (online/offline), Model Monitor and Clarify, and Experiments.
- Deployment & MLOps (core). Proven production ML operations on AWS: CI/CD for ML, infrastructure-as-code (CloudFormation/CDK or Terraform), containerization (ECR/ECS/EKS), blue-green, canary and shadow deployments, observability, incident management and cost optimization - including hands-on production on-call experience.
- Retail domain experience (required). Working experience in fashion retail or broader retail/e-commerce - in-house or through delivered client engagements - with fluency in retail e-commerce concepts and metrics (conversion funnel, AOV, merchandising, seasonality).
- Broader AWS data & serving stack. S3, Glue, Athena/EMR, Kinesis, Lambda, API Gateway, Step Functions/EventBridge; security and governance with IAM, KMS and VPC design, including data-residency controls.
- Real-time decisioning. Experience engineering low-latency (sub-100 ms budget) decision services with caching, asynchronous patterns and graceful fallback in consumer-scale web environments.
- Experimentation rigor. Strong applied A/B testing practice: experiment design, power analysis, guardrail metrics, sequential-testing pitfalls, and revenue-metric measurement (conversion rate, AOV).
- Engineering foundation. Expert Python and SQL; PyTorch or TensorFlow; familiarity with bandit/RL tooling (e.g., Vowpal Wabbit, RLlib).
- Experience profile. Typically 7+ years of applied machine learning, including 3+ years on production personalization/decisioning systems at consumer scale, with at least one engagement carrying a system through handover or takeover into supported operation.
- Consulting maturity. Track record in client-facing professional services: architecture reviews, executive communication, documentation quality, and the ability to challenge a delivery vendor constructively while keeping the joint team productive.
- Eligibility & location. Currently engaged by AWS Professional Services, or by an AWS Partner (Services) - Advanced or Premier tier preferred, ideally holding the AWS Machine Learning Competency; able to work full-time on-site in Singapore for the duration of the engagement and to sustain 24/7 on-call ownership after go-live.
PREFERRED QUALIFICATIONS (GOOD TO HAVE)
- AWS certifications: AWS Certified Machine Learning - Specialty and/or AWS Certified Machine Learning Engineer - Associate; AWS Certified Solutions Architect a plus.
- Platform familiarity: working knowledge of clickstream/event analytics platforms (e.g., Amplitude) and e-commerce or marketing platforms (e.g., Salesforce Commerce Cloud) - good to have, not required; deep expertise is not expected.
- Cross-cloud data integration: experience ingesting and reconciling data into AWS from non-AWS cloud environments.
- Takeover experience: taking over, operating and improving systems built by a third party, including acceptance testing and playbook-driven operations.
- Regional experience: multi-market Asia-Pacific consumer e-commerce.