Databricks with AWS & Pyspark
cognizant
Job Description
- Develop and implement scalable data processing solutions using Spark in Scala to enhance data-driven decision-making.
- Manage and administer Databricks Unity Catalog to ensure data governance and security compliance.
- Utilize Delta Sharing to facilitate secure and efficient data sharing across various platforms.
- Configure and maintain Databricks CLI for seamless integration and automation of workflows.
- Design and execute Delta Live Pipelines to streamline data ingestion and transformation processes.
- Implement Structured Streaming solutions to handle real-time data processing and analytics.
- Collaborate with cross-functional teams to integrate risk management strategies into data solutions.
- Leverage Apache Airflow for orchestrating complex data workflows and ensuring timely execution.
- Optimize data storage and retrieval using Amazon S3 and Amazon Redshift to improve performance.
- Develop Python scripts to automate data processing tasks and enhance operational efficiency.
- Utilize Databricks SQL for querying and analyzing large datasets to derive actionable insights.
- Implement Databricks Delta Lake to ensure data reliability and consistency across the platform.
- Manage Databricks Workflows to automate and streamline data engineering processes.