Senior Spark / PySpark Data Engineer
webspiders
Job Description
- Design, develop, and optimize large-scale ETL/ELT pipelines using Apache Spark and PySpark.
- Develop scalable data transformation and processing solutions using PySpark and Python.
- Build distributed data-processing applications capable of handling large volumes of data.
- Develop reusable and maintainable Spark/PySpark frameworks and data-processing components.
- Optimize Spark jobs for performance, scalability, memory utilization, and execution efficiency.
- Work with complex transformations, joins, aggregations, partitioning, and large datasets.
- Implement data validation, quality checks, error handling, and monitoring within data pipelines.
- Work with AWS data services including EMR, Glue, S3, and Redshift.
- Develop data pipelines supporting data lakes, warehouses, analytics, and downstream applications.
- Troubleshoot production data pipeline and Spark processing issues.
- Identify and resolve performance bottlenecks in Spark/PySpark workloads.
- Collaborate with Data Engineering, Cloud, AI/ML, and Product teams to deliver reliable data solutions.