Data Engineer
epergne solutions- Posted 7 hours ago
- Be among the first 10 applicants
Job Description
Position Summary / Project Description
This Data Engineer role is designed to augment our existing data engineering team to accelerate the implementation and operationalization of our enterprise data platform built on Databricks. The successful candidate will be responsible for leading the migration of legacy data pipelines to Databricks, establishing operational excellence frameworks, and ensuring seamless integration with existing healthcare information systems and clinical data workflows.
The role involves working on critical projects including the Healthcare Data Analytics Platform modernization and advanced analytics infrastructure supporting machine learning operations. The position holder will be expected to mentor junior team members.
Key projects include migrating existing ETL processes from traditional data warehouse systems to Databricks Delta Lake architecture, implementing real-time streaming analytics for clinical monitoring systems. The role requires close collaboration with infrastructure teams, compliance officers.
Role and Responsibilities
Operational Responsibilities:
Monitor and maintain production data pipelines to ensure 99.9% uptime and optimal performance
Implement comprehensive logging, alerting, and monitoring systems using Application monitoring tools
Perform regular health checks performance, job execution times, and resource utilization to identify and resolve bottlenecks proactively
Manage incident response procedures for pipeline failures, including root cause analysis, resolution, and post-incident reviews
Establish and maintain disaster recovery procedures and backup strategies for critical data assets within the Databricks environment
Conduct regular performance tuning of Spark jobs and Databricks cluster configurations to optimize cost and execution efficiency
Maintain comprehensive documentation for operational procedures, runbooks, and troubleshooting guides
Coordinate scheduled maintenance windows and system upgrades with minimal business impact
Manage user access controls, workspace configurations, and security policies within Application environments
Requirements / Qualifications
Education & Experience:
Degree in Computer Science or Computer Engineering
Minimum 5 years working experience in system operations compliance and management areas
Project hands-on experience specifically with AWS platform (primary requirement)
project experience in cloud operations or cloud architecture
Must be cloud certified (AWS)
Core Technical Skills:
proficiency in Databricks platform, including workspace management, cluster configuration, and job orchestration
Strong expertise in Apache Spark within Databricks environment, including Spark SQL, DataFrames, and RDDs
Good in-depth understanding of data warehouse concepts, data profiling, data verification and advanced analytics techniques
Strong knowledge of monitoring, incident management, and cloud cost control
Technology Stack Experience:
Databricks
AWS cloud services and architecture
IDMC (Informatica Data Management Cloud)
Tableau for data visualization
Oracle Database management
ML Ops practices within Databricks environment
STATA for statistical analysis is advantage
Amazon SageMaker integration with Databricks
DataRobot platform integration
More Info
Key Skills
Amazon SageMaker
Cloud cost control
ML Ops practices
DataRobot
ETL processes
IDMC Informatica Data Management Cloud
Advanced analytics techniques


