Search by job, company or skills

Data Engineer-Spark,Scala

Early Applicant
  • Posted 12 hours ago
  • Be among the first 10 applicants

Job Description

Data Engineer (Spark/Scala)

About The Role

We are seeking an experienced Data Engineer to design, build, and optimize complex data workflows across on-premises and cloud environments. This role requires deep hands-on expertise in Apache Spark, Databricks, and Scala/PySpark, along with strong SQL and Python skills, to build robust, high-performance data pipelines. You will work extensively on complex on-prem workflows, integrating data across multiple file systems and formats, migrating and modernizing legacy processes, and ensuring efficient, reliable data movement across heterogeneous environments.

Key Responsibilities

  • Design, develop, and maintain large-scale data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark
  • Build and support complex on-premises data workflows, including migration/hybrid on-prem-to-cloud integration patterns
  • Integrate data across diverse file systems (on-prem file shares, NAS, HDFS, S3) and formats JSON, Parquet, Fixed-Length, CSV, Excel, Avro
  • Write efficient, optimized SQL for data extraction, transformation, and loading across relational databases
  • Connect to and extract data efficiently from various source databases, tuning queries and pipelines for performance at scale
  • Develop and maintain workflow orchestration using Airflow (or similar schedulers) for reliable, monitored pipeline execution
  • Write clean, production-grade Python code for data processing, automation, and tooling
  • Build and maintain unit/integration tests for data pipelines to ensure data quality and reliability
  • Create and maintain clear technical documentation for pipelines, data flows, and system architecture
  • Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data quality issues in complex, multi-system workflows
  • Collaborate with cross-functional teams (data science, analytics, application engineering) to support downstream data consumption
  • Support cloud integration efforts, particularly with Azure, as workloads evolve from on-prem to hybrid/cloud architectures

Required Qualifications

Primary Skills:

  • Strong hands-on experience with Apache Spark and Databricks for large-scale data processing
  • Proficiency with Amazon S3 for data storage and pipeline integration
  • Strong SQL skills - query optimization, complex joins, performance tuning
  • Proven experience integrating data across various file systems and formats: JSON, Parquet, Fixed-Length, CSV, Excel, Avro, etc.
  • Strong knowledge of Scala Spark and PySpark for distributed data processing
  • Strong Python programming skills for scripting, automation, and data engineering tasks
  • Strong experience connecting to and efficiently extracting data from databases (relational/other), including performance-conscious extraction strategies
  • Demonstrated experience working on complex on-prem data workflows (multi-system integration, legacy system data extraction, hybrid on-prem/cloud pipelines)
  • Experience leveraging coding assistant tools and implementing AI agents to enhance development productivity and task execution.

Secondary Skills:

  • Experience with Azure cloud services (storage, compute, data services)
  • Experience with Apache Airflow for workflow orchestration and scheduling
  • Experience writing automated tests for data pipelines (unit, integration, data quality checks)
  • Strong documentation skills able to clearly document pipelines, data lineage, and technical designs

Good To Have:

  • Working knowledge of Java
  • Familiarity with React for building internal tooling/dashboards
  • Experience with Prefect for workflow orchestration
  • PBM (Pharmacy Benefit Management) / Healthcare domain knowledge

Data Engineer (Spark/Scala)

Good To Have:

  • Working knowledge of Java
  • Familiarity with React for building internal tooling/dashboards
  • Experience with Prefect for workflow orchestration
  • PBM (Pharmacy Benefit Management) / Healthcare domain knowledge.

Data Engineer (Spark/Scala)

Good To Have:

  • Working knowledge of Java
  • Familiarity with React for building internal tooling/dashboards
  • Experience with Prefect for workflow orchestration
  • PBM (Pharmacy Benefit Management) / Healthcare domain knowledge

Skills: scala,pipelines,data,spark

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152706011

Similar Jobs

Bengaluru, India

Skills:

Big Data - Data ProcessingJavaSparkApacheScalaPysparkSparksqlTechnology

Bengaluru, India

Skills:

JavaAgile MethodologiesScalaSparkBig DataApacheSdlcSoftware Configuration Management SystemsFunctional ProgrammingData Processing

Bengaluru, India

Skills:

ScalaElastic SearchGitlabAkka LEGOM frameworkPostgres databaseGitlab PipelinesMicroservice architecturesContainerized applicationsSlick Connector for Database integrationApache Pulsar

Bengaluru

Skills:

ScalaAPI designSqlKubernetesCI/CD

Beware of Scammers

We don’t charge money for job offers