Data platform Engineer
ripplehire
Job Description
-
Design, build, and maintain robust, scalable, and efficient data pipelines for batch and streaming data processing. Develop and optimize ETL/ELT workflows using Python and PySpark to process large-scale datasets.
-
Architect and manage cloud infrastructure on AWS (S3, EMR, Glue, Lambda, Redshift, EC2, IAM, etc.).
-
Use Terraform to provision, manage, and version-control infrastructure as code (IaC). Collaborate with data scientists, analysts, and business stakeholders to understand data requirements and deliver reliable data solutions.
-
Ensure data quality, integrity, and security across pipelines and storage systems.
-
Monitor and troubleshoot production data pipelines, ensuring high availability and performance.
-
Implement CI/CD pipelines for data engineering workflows.
-
Document technical designs, data flows, and architecture for internal reference.
-
Mentor junior data engineers and contribute to best practices within the team.
-