GCP Data Engineer
equinix
Job Description
-
Design, develop, and maintain scalable ETL/ELT pipelines using Cloud Dataflow, Cloud Composer (Apache Airflow), Dataform/dbt and Cloud Functions
-
Build real-time streaming data pipelines using Cloud Pub/Sub, Kafka and Dataflow
-
Implement automated data quality checks and monitoring across all data workflows
-
Optimize pipeline performance and cost efficiency through proper resource allocation and scheduling
-
Architect and implement data lake and data warehouse solutions using Dataproc, BigQuery, Cloud Storage, and Cloud SQL
-
Design optimal data models, partitioning strategies, and clustering for analytical workloads
-
Manage data lifecycle policies and implement automated archival and retention strategies
-
Ensure data security, encryption, and access control across all storage layers
-
Build and optimize BigQuery datasets for analytics and reporting use cases
-
Create and maintain dimensional models and fact tables for business intelligence
-
Implement data marts and aggregation layers for improved query performance
-
Support self-service analytics through proper data cataloging and documentation
-
Having good knowledge of Dataplex and Analytics hub
-
Integrate data from various sources including databases, APIs, SaaS applications, and file systems
-
Implement change data capture (CDC) solutions for real-time data synchronization
-
Work with third-party data providers and external data feeds
-
Implement comprehensive monitoring and alerting using Cloud Monitoring and Cloud Logging
-
Troubleshoot data pipeline issues and implement robust error handling mechanisms
-
Maintain data lineage documentation and impact analysis capabilities