Search by job, company or skills

GCP Data Engineer Postgres Python Pyspark Developer

Early Applicant
  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

Information Details

1 Role** GCP Data Engineer Postgres Python Pyspark Developer (Big Query, Cloud Storage, Dataproc, Airflow)

2 Required Technical Skill Set** GCP Data Engineer to design, build, and optimize scalable data pipelines and analytics solutions using Big Query, Cloud Storage, Dataproc, and Airflow.

3 No of Requirements** 5

4 Desired Experience Range** 7+ Years

5 Location of Requirement HYDERABAD

6 Immediate Joiners Needed YES

Desired Competencies (Technical/Behavioral Competency)

Must-Have** (Ideally should not be more than 3-5)

· GCP Services: Big Query, Cloud Storage, Dataproc, Cloud Composer (managed Airflow) or self-managed Airflow.

· Airflow: Strong experience in DAG creation, operators/hooks, scheduling, backfilling, retry strategies, and CI/CD for DAG deployments.

· Programming: Proficiency in Postgres Programming including Tables, Triggers, Views, Stored procedures, Python and Pyspark (PySpark, Airflow DAGs), SQL (advanced BigQuery SQL).

· Data Modeling: Dimensional modeling (Star/Snowflake), data vault basics, and schema design for analytics. ·

Performance Tuning: BigQuery partitioning/clustering, predicate pushdown, job stats review, Dataproc executor tuning. ·

Version Control & CI/CD: Git, branching strategies, pipelines for deploying Airflow DAGs and config.

· Operational Excellence: Monitoring with Stackdriver/Cloud Logging, debugging pipeline failures, and root-cause analysis. · involves end-to-end ownership of data ingestion, transformation, orchestration, and performance tuning for batch and near real-time workflows.

Good-to-Have

· Streaming: Pub/Sub, Dataflow (Apache Beam) for near real-time pipelines.

· Orchestration Patterns: Event-driven pipelines, dependency management, and cross-environment promotion.

· Data Governance: Catalog/lineage tools (e.g., Data Catalog), PII handling, row-level security, column-level encryption.

· Containers & Infra: Docker, Terraform for IaC on GCP; Kubernetes concepts.

· BI Integration: Experience integrating with Looker, Tableau, or Power BI.

· Certifications: Google Professional Data Engineer / Cloud Architect.

Responsibility of / Expectations from the Role

1 Data Pipeline Development: Build robust ETL/ELT pipelines using Apache Airflow (DAG creation, scheduling, monitoring) to orchestrate data workflows across GCP services.

2 Data Warehousing: Design and optimize BigQuery schemas, partitioning/clustering strategies, materialized views, and query performance tuning.

3 Data Processing: Implement scalable data processing using Dataproc (Spark/Hive), including job configuration, optimization, and cost control.

4 Data Ingestion & Storage: Manage ingestion from diverse sources (APIs, files, streaming) and design storage strategies using Cloud Storage (lifecycle policies, tiers, security).

5 Quality & Observability: Implement data validation (e.g., Great Expectations or custom checks), logging/alerting, and SLA monitoring for pipelines.

6 Security & Governance: Apply IAM, service accounts, VPC SC, CMEK, and access policies across GCP resources; ensure compliance with data governance standards.

7 Cost & Performance: Optimize queries, cluster usage, and storage to balance cost/performance; leverage reservation/flex slots and job-level optimizations.

8 Collaboration: Work closely with analytics, product, and business teams to translate requirements into scalable data solutions; create documentation and handover materials.

More Info

Job Type:
Industry:
Employment Type:

Job ID: 152252595

Beware of Scammers

We don’t charge money for job offers