About the Company
We are seeking highly skilled Data Engineers with strong expertise in Python, Hadoop, AI, and Big Data technologies to join a leading banking client's enterprise Data & AI transformation programme.
About the Role
The ideal candidate will have extensive experience designing and building enterprise-scale Data Lakes, Data Warehouses, Big Data platforms, and AI-ready data pipelines that support Machine Learning, Generative AI, Advanced Analytics, and Regulatory Reporting. If you enjoy working on large-scale data platforms and AI-driven solutions in the banking industry, we'd love to hear from you.
Responsibilities
- Design, develop, and maintain scalable enterprise Data Pipelines and ETL/ELT processes.
- Build high-performance Big Data solutions using Python, Hadoop, Spark, and distributed computing technologies.
- Develop and optimise Enterprise Data Warehouse, Data Lake, and Lakehouse platforms.
- Build AI-ready data platforms to support Machine Learning, Generative AI, LLM, and Advanced Analytics use cases.
- Process and optimise structured, semi-structured, and unstructured datasets.
- Design reusable data frameworks, APIs, and data integration components.
- Collaborate closely with Data Scientists, AI Engineers, Data Architects, Business Analysts, and Product Owners.
- Implement data governance, metadata management, data quality, lineage, and security controls.
- Optimise data processing performance and troubleshoot production issues.
- Participate in DevOps, CI/CD, automation, and code reviews.
Qualifications
- Bachelor's Degree in Computer Science, Engineering, Information Technology, or a related discipline.
- Banking or Financial Services experience is mandatory.
- Proven experience delivering enterprise Data Engineering, AI, Big Data, and Data Warehouse solutions.
- Strong understanding of enterprise data architecture, data governance, and security.
- Experience working in Agile/Scrum delivery environments.
- Excellent analytical, problem-solving, and communication skills.
Required Skills
- 4–10 years of Data Engineering experience
- Strong hands-on experience with Python
- Strong experience with the Hadoop ecosystem (HDFS, Hive, YARN, MapReduce)
- Apache Spark / PySpark
- SQL and advanced database development
- ETL/ELT pipeline development
- Enterprise Data Warehouse
- Data Lake / Lakehouse architecture
- Big Data technologies
- Linux / Unix
- Git
Preferred Skills
- Apache Kafka
- Apache Airflow
- Databricks
- Snowflake
- Delta Lake
- Docker
- Kubernetes
- REST APIs
- CI/CD pipelines
- AWS, Azure, or Google Cloud Platform (GCP)
- MLflow
- LangChain or LlamaIndex
- Vector Databases (Pinecone, Chroma, Weaviate, FAISS)
- MLOps
EA License No. 01C4394 • EA Registration No. R1113321 (Jacob Tijo)