Focus: Production ASR, GPU Infrastructure, Reliability, and Semantic Data Systems
Location: Singapore (Remote-friendly)
Type: Full-time
Company: VALSEA
About VALSEA
VALSEA is building Southeast Asia's speech intelligence infrastructure.
We turn real-world regional speech (accent-heavy, mixed-language, messy audio) into reliable, workflow-ready outputs using a combination of production-grade ASR and a proprietary semantic understanding layer.
We build systems that must work every day, under real customer usage, with clear failure modes and fast recovery.
The Role
This is a senior systems role focused on keeping mission-critical speech systems alive, predictable, and scalable.
You will own:
- GPU-based ASR inference running in production
- Reliability of APIs and downstream workflow endpoints
- Real-time debugging when things break (because they will)
- Coordination and technical oversight of interns building country-specific semantic datasets
What You Will Own
1. Production ASR & GPU Infrastructure
Deploy and operate GPU-backed ASR inference services on AWS
- Handle model migration, versioning, rollouts, and rollbacks
- Manage latency, batching, concurrency, and cost tradeoffs
- Ensure inference servers are always online and recover cleanly from failures
- Set up health checks, autoscaling logic, and failure detection
You are expected to know what breaks in GPU inference systems and how to fix it without guessing.
2. Reliability & Real-Time Debugging
- Be accountable for system uptime and API correctness
- Debug production issues in real time, not tomorrow
- Diagnose failures across:
- GPU inference
- Audio pipelines
- APIs
- Schema mismatches
- Partial or corrupted inputs
- Reduce customer-facing downtime to minutes, not hours
If something fails, you own the fix end-to-end.
3. Semantic Core & Dataset Oversight
- Technically manage interns buildingcountry-specific semantic datasets
- Define:
- Data formats
- Validation rules
- Quality thresholds
- Review outputs to ensure they integrate cleanly into the semantic core
- Prevent bad data from silently corrupting production behaviour
You are not labelling data yourself full-time, but you are responsible for making sure the system remains coherent.
4. Stable Interfaces & Workflow Outputs
- Design and maintain durable JSON schemas
- Ensure outputs are consistent and safe for downstream workflows
- Handle ambiguity, low confidence, and missing data explicitly
- Treat graceful degradation as a first-class system behaviour
How You'll Work
- You are comfortable making decisions with incomplete specs
- You document assumptions and tradeoffs
- You prioritise predictability over cleverness
- You stay calm when production breaks
- You do not wait for permission to fix critical issues
This is a role for someone who keeps systems alive, not someone who waits to be told what to do.
Required Experience
You must have:
- 5+ years in backend, systems, platform, or infrastructure engineering
- Hands-on production experience with GPU-based inference or ML serving
- Experience deploying and operating systems on AWS
- Built and owned real-time or near-real-time APIs
- Debugged live production systems under user traffic
- Strong Python and Linux fundamentals
- Experience with Docker and CI/CD pipelines
- Comfort being accountable for system behaviour, not just code
Strong Signals
We are especially interested if you have:
- Operated GPU inference under cost and latency constraints
- Worked on speech, audio, or media pipelines
- Designed or enforced data schemas in production systems
- Used infrastructure-as-code (Terraform, AWS CDK)
- Built queue-based or async systems (SQS, Kafka, Redis)
- Participated in on-call rotations or owned reliability metrics
Who This Role Is NOT For
Do not apply if:
- You prefer research over production
- You dislike being on-call or handling outages
- You avoid debugging under pressure
- You need highly detailed specs before moving forward
- You want a narrow, siloed role
- You are optimising for title, visibility, or resume optics
This role requires ownership, humility, and operational toughness.
Compensation & Growth
- Cash + meaningful ESOP
- Growth is measured by system ownership and trust, not titles
- As a founding engineer, your decisions will shape VALSEA's core infrastructure for years.
If you take pride in building systems that:
- survive real users
- fail clearly
- recover fast
- and improve steadily over time
You will feel at home here.