Job Summary
The Data Engineer will design, develop, and maintain scalable, reliable data pipelines on Databricks and cloud platforms, integrating diverse data sources to support analytics, reporting, and machine learning. Collaborate with cross-functional teams to enhance data platform governance, monitoring, and reliability.
Responsibilities
- Design, develop, and maintain ETL pipelines for centralized data storage systems such as Delta Lake to ensure efficient data ingestion and transformation
- Integrate data from databases, APIs, log files, streaming platforms, and external providers to support diverse analytics needs
- Develop data transformation routines to clean, normalize, and aggregate complex or inconsistent datasets for accurate analysis
- Apply data processing techniques to handle large-scale and varied data sources effectively
- Contribute to the development and enforcement of frameworks and best practices for code development and deployment to ensure quality and consistency
- Implement data governance policies aligned with company standards to maintain data integrity and compliance
- Collaborate with analytics and product teams to design and operationalize data pipelines that meet business requirements
- Work with infrastructure teams to advance cloud-based data platforms leveraging Azure and Databricks technologies
- Explore and evaluate new tools and techniques to enhance data platform capabilities on Azure, Databricks, or related platforms
- Monitor data pipelines continuously to detect, troubleshoot, and resolve issues promptly, ensuring high availability
- Develop monitoring tools, alerts, and automated error-handling mechanisms to maintain pipeline reliability
- Analyze business requirements and translate them into data extraction and pipeline development tasks
- Participate in requirement grooming and refinement sessions with users to clarify data needs
- Optimize pipeline performance and batch scheduling to maximize efficiency and minimize latency
- Develop dashboards, reports, scorecards, and data visualizations to support data-driven decision-making
- Perform system integration testing (SIT), data validation, and profiling to confirm data accuracy and completeness
- Validate completeness and consistency of ETL loads to ensure reliable data delivery
- Support user acceptance testing (UAT) and production implementation activities to ensure smooth deployment
Required competencies and certifications
- Databricks Certified Data Engineer Associate (strongly preferred)
- Databricks Certified Data Engineer Professional (strongly preferred)
Preferred competencies and qualifications
- 3 or more years of experience in data engineering with scalable pipelines
- Strong experience designing data solutions including data modelling and distributed computing architectures
- Hands-on experience with data processing jobs using PySpark, Spark SQL, and Databricks notebooks/jobs
- Experience orchestrating data pipelines with Azure Data Factory (ADF), Airflow, or similar tools
- Experience with both real-time and batch data processing
- Experience building pipelines on Azure AWS experience is beneficial
- Proficiency in SQL including window functions and performance optimization
- Understanding of DevOps tools, Git workflows, and CI/CD pipelines
- Familiarity with Scrum methodology and experience working in Scrum teams
- Ability to apply Scrum practices in a practical project context
- Strong problem-solving and collaborative mindset
- Experience with streaming technologies such as Apache Kafka, Apache Flink, or AWS Kinesis
- Ability to design and implement real-time data processing pipelines