Requirements:
Data Ingestion
- Build and maintain reliable pipelines for ingesting data from diverse external sources
- Develop robust, schedulable, production-grade data ingestion workflows
Document Parsing & Structured Extraction
- Extract structured data from PDF, HTML, XBRL, and other document formats
- Handle scanned PDFs, complex tables, multilingual documents (English & Chinese), and other real-world document challenges
Data Warehouse Design
- Design and maintain well-structured schemas on modern cloud data warehouses such as Snowflake, BigQuery, or PostgreSQL
- Build scalable, maintainable data models for downstream analytics
Pipeline Orchestration & Monitoring
- Develop production ETL workflows using Airflow, Prefect, or Dagster
- Implement data quality validation, monitoring, alerting, and data lineage
LLM Engineering
- Build production applications using OpenAI, Anthropic, and open-source LLMs
- Develop structured output pipelines and Retrieval-Augmented Generation (RAG) workflows
- Design lightweight evaluation frameworks and quality monitoring for LLM-powered systems
Backend Services
- Build query APIs using FastAPI (or similar frameworks) to serve downstream analyst tools
Cloud Infrastructure
- Deploy and operate systems on AWS or GCP
- Leverage managed cloud services to minimize operational overhead
Requirements
- 5–10 years of experience building and operating production software systems
- Strong proficiency in Python and SQL
- Solid experience with data modeling, testing, and performance optimization
- Hands-on experience parsing messy real-world documents (PDF, HTML, XBRL, etc.)
- 2–3 years of production experience building and deploying LLM-powered applications
- Experience with at least one workflow orchestration framework (Airflow, Prefect, or Dagster)
- Experience designing schemas for at least one modern cloud data warehouse
- Excellent communication skills, with the ability to collaborate closely with research teams and company leadership
- Professional working proficiency in English
Strongly Preferred
- Previous experience as a Founding Engineer or early engineering hire at an AI startup
- Built end-to-end data platforms in an early-stage startup environment
Nice to Have
- Experience building data infrastructure that supports NLP or LLM workloads
- Hands-on experience with dbt
- Experience designing multi-tenant systems and access control
- Familiarity with graph databases (e.g. Neo4j)
Application:
- Apply to this job posting, and send your CV with the job title as the subject line to: [Confidential Information] & https://www.linkedin.com/in/treasa-wong/