Databricks with AWS & Pyspark

cognizant

Hyderabad 7 Years Exp Posted 4h ago

Job Description

  • Develop and implement scalable data processing solutions using Spark in Scala to enhance data-driven decision-making.
  • Manage and administer Databricks Unity Catalog to ensure data governance and security compliance.
  • Utilize Delta Sharing to facilitate secure and efficient data sharing across various platforms.
  • Configure and maintain Databricks CLI for seamless integration and automation of workflows.
  • Design and execute Delta Live Pipelines to streamline data ingestion and transformation processes.
  • Implement Structured Streaming solutions to handle real-time data processing and analytics.
  • Collaborate with cross-functional teams to integrate risk management strategies into data solutions.
  • Leverage Apache Airflow for orchestrating complex data workflows and ensuring timely execution.
  • Optimize data storage and retrieval using Amazon S3 and Amazon Redshift to improve performance.
  • Develop Python scripts to automate data processing tasks and enhance operational efficiency.
  • Utilize Databricks SQL for querying and analyzing large datasets to derive actionable insights.
  • Implement Databricks Delta Lake to ensure data reliability and consistency across the platform.
    • Manage Databricks Workflows to automate and streamline data engineering processes.