Data Engineer
ford
Job Description
- GCP Pipeline Development: Design, build, and maintain highly scalable ETL/ELT data pipelines using Python and GCP-native data processing tools (e.g., Cloud Run, Cloud Functions).
- AI/ML Infrastructure Support: Engineer feature stores, robust data feeds specifically optimized for machine learning training and inference. Work closely with ML Engineers to operationalize models using Vertex AI.
- Data Integration & Ingestion: Write clean, modular Python code to ingest data from diverse sources (APIs, streaming platforms, on-prem databases) into BigQuery and Google Cloud Storage (GCS).
- System Optimization: Optimize BigQuery architecture, partition/cluster tables, and tune complex SQL queries to ensure performance and cost-efficiency at a massive scale.
- Software Engineering Best Practices: Champion best practices in Python development, including version control (Git), CI/CD pipelines (Cloud Build / GitHub Actions), code reviews, and comprehensive unit/integration testing.
- Data Quality & Governance: Implement robust data quality checks, alerting, and monitoring to ensure the data feeding our AI models is accurate and reliable.