Search Jobs

Search by job, company or skills

Senior Data Engineer/Data Engineer

Senior Data Engineer/Data Engineer

apba tg human resource pte. ltd.
10-13 Years
SGD 12,000 - 20,000 per month
Early Applicant
  • Posted 11 hours ago
  • Be among the first 10 applicants

Job Description

Job Description

This role sits within NCS AI Central's (AIC) Forward Deployed Engineering (FDE) model - the combined capability that takes AI solutions from proof-of-concept through to hardened production systems. You will operate across both fast-moving FDE engagements (POC/POV, pilot deployments for strategic and lighthouse clients) and steady-state system development and maintenance work - bringing the same rigor and a reusable, asset-fed approach to both.

What will you do:

1. Data Pipeline Engineering & AI-Readiness

. Design and build ingestion, cleaning, and transformation pipelines that turn messy, real-world client data into AI-ready datasets.

. Build batch and streaming pipelines (Airflow/Prefect/Kafka) that keep data flowing reliably into AI systems without manual intervention.

. Own data quality - deduplication, schema validation, completeness checks - upstream of any model or RAG pipeline.

. Proactively flag data gaps or quality issues that would degrade model/RAG performance downstream, before they surface as an AI Engineer's problem in testing.

2. RAG & Vector Store Architecture

. Architect document/data ingestion and indexing pipelines for Retrieval-Augmented Generation (RAG) systems - chunking strategy, embeddings, hybrid/vector search.

. Design and operate vector database and search infrastructure (pgvector/Pinecone/OpenSearch) at production scale and query volume.

3. Data Governance & Compliance

. Implement PII redaction, data residency, and access-control patterns aligned to PDPA and sector-specific requirements (Healthcare, Government, Transport).

. Maintain clear data lineage and metadata governance so engagement teams and auditors can trace how client data flows into AI outputs.

4. FDE & Development/Maintenance Coverage

. During FDE engagements: rapidly assess and prepare a client's data landscape during Discover/POC, identifying data-readiness gaps early.

. During system development & maintenance engagements: build and operate production-scale data pipelines handling the full volume and complexity of live client systems (e.g., Healthcare or Transport data at scale).

. Contribute reusable ingestion/indexing patterns back into the shared internal asset library to accelerate future engagements.

5. Collaboration & Leadership

. Partner closely and continuously with AI Engineers and AI Architects - understanding what a given model, RAG pipeline, or agent actually needs from the data layer, and translating that into concrete pipeline and schema design decisions.

. Own the definition of AI-ready data for each engagement jointly with AI Engineers - agreeing on chunking strategy, metadata, freshness, and quality thresholds before pipelines are built, not after retrieval quality suffers.

. Sit in solution design conversations alongside AI Engineers and AI Architects, so data architecture and model/RAG architecture are designed together rather than data being treated as a downstream dependency.

. Mentor junior data engineers and set data engineering standards across engagements.

Qualifications

The ideal candidate should possess:

. 10+ years in data engineering, including production-scale pipeline design (not just analytics/reporting pipelines).

. Strong SQL and at least one systems language (Python/Scala/Java) hands-on with batch and streaming frameworks (Airflow, Spark, Kafka).

. Experience building data pipelines for AI/ML or RAG use cases - embeddings, vector indexing, hybrid search.

. Solid understanding of data governance, PII handling, and access-control patterns in regulated environments.

. Comfortable moving between fast, exploratory data assessment (FDE/POC) and disciplined, high-volume production pipeline engineering (system development & maintenance).

. Working understanding of core AI/LLM concepts - tokenization, embeddings, chunking strategy, context windows, RAG, and agentic workflows - sufficient to hold a real technical conversation with AI Engineers and AI Architects about what AI-ready data means for a given use case, not just how to move and clean it.

Preferred Qualifications

. Experience with vector databases (pgvector, Pinecone, Weaviate) and search platforms (OpenSearch/Azure AI Search).

. Exposure to Singapore Government data environments (GCC/HCC) and compliance regimes (IM8, PDPA).

. Experience with sector-specific data complexity - Healthcare (clinical data governance) or Transport/Aviation systems.

. Familiarity with data cataloguing and lineage tooling.

. Prior experience embedded within an AI/ML delivery team (not just a data platform team) - i.e., has sat alongside AI Engineers day-to-day and adjusted pipeline/schema design based on model or RAG performance feedback.

Tech Stack (Illustrative)

. Languages: Python, SQL (Scala/Java a plus)

. Pipelines: Airflow/Prefect, Spark, Kafka/Debezium

. Storage/Search: Postgres, S3/Blob, pgvector/Pinecone/Weaviate, OpenSearch/Azure AI Search

. Governance: Presidio (PII redaction), data catalogue/lineage tooling

. Cloud: AWS/Azure/GCP GCC/HCC exposure a plus

More Info

Job Type:
Industry:
Employment Type:

Key Skills

pgvector

Pinecone

Prefect

OpenSearch

PII handling

Data Pipeline Engineering

Metadata governance

Similar Jobs

10-13 yrs
SGD 10,000 - 20,000 per month
Anson, Singapore
Skills:
Java, Hadoop, Scala, Spark, Clickhouse, Go, Flink
12-14 yrs
Singapore
Skills:
Java, AWS Glue, REST, Gcp, Docker, Spark, Azure, Python, AWS, Apache Iceberg, Parquet, gRPC APIs, Trino
7-10 yrs
SGD 5,000 - 8,300 per month
Singapore
Skills:
snowflake , BigQuery, Power Bi, Apache Spark, Tableau, Redshift, Sql, Apache Airflow, Databricks, Python, AWS, dbt
8-10 yrs
Singapore
Skills:
Data Modelling, Databricks, Automation, Data Integration, Sql, Etl, ELT, DataOps, Data Quality Management
5-10 yrs
Singapore
Skills:
data engineering , Pyspark, Data Architecture, Databricks, Python, Sql