Snr Data Engineer
alight
Job Description
- Build and maintain high volume ETL/ELT pipelines across Hadoop (HDFS, Hive, Spark, Kafka) and AWS (Glue, EMR, Lambda, Step Functions, Redshift).
- Develop distributed data processing solutions using PySpark, Spark SQL, and scalable cloud serverless patterns.
- Implement reusable data ingestion frameworks for batch, ability to design & implement Orchestration process and Leverage AI
- Optimize data workflows using partitioning, bucketing, compression, file formats (Parquet/ORC).
- Understanding hybrid data lake architectures using S3 + HDFS, ensuring data governance and best practices are adheres
- Experience to deliver complex projects in an Agile environment
- Assist in Design and build the robust, scalable and secure software solutions across the having no/least adoption
- Define clear technical specifications and make architecture decisions that align with business goals and long-term scalability.
- Implement best practices (including secure code guidelines) through the implementation of unit tests, automation, leverage and code reviews. Drive continuous improvement in code quality and maintainability.
- Troubleshooting issues and proactively solving problems as they arise, ensuring the smooth operation of full stack applications
- Ability to understand the data flow diagram, data modelling and Lineages
- Job orchestration using Airflow, Control M, Step Functions, or event-driven triggers.
- Ensure data is protected and compliant with regulatory standards.
- Work closely with business stakeholders to enable high quality datasets.
- Work on best practice adoption and provide guidance to peers/juniors in team.
- Ability to respond on incidents, and troubleshooting Spark performance issues, job failures, and cluster bottlenecks.
- Collaborate closely with team members, QA and cross product teams to streamline release processes.
- Collaborate with business stakeholders to gather, analyse, and translate data into technical solutions