Data Lake SME
hpe `
Job Description
- Data Lake Architecture & Development
- Design and implement data ingestion pipelines for structured, semi-structured, and unstructured data.
- Develop and manage ETL/ELT processes for large-scale data processing.
- Optimize storage and retrieval strategies across on-prem and cloud-based data lakes.
- Data Integration & Processing
- Integrate data from multiple sources (databases, APIs, streaming platforms).
- Implement real-time and batch processing using Apache Spark, Kafka, or Flink.
- Support metadata management, data lineage, and cataloging.
- Performance & Optimization
- Tune queries and pipelines for high performance and cost efficiency.
- Implement partitioning, indexing, and caching strategies for large datasets.
- Automate routine ETL/ELT workflows for reliability and speed.
- Security & Governance
- Ensure compliance with data governance, privacy, and regulatory standards (GDPR, HIPAA, etc.).
- Implement encryption, masking, and role-based access control (RBAC).
- Collaborate with cybersecurity teams to align with Zero Trust and IAM policies.
- Collaboration & Support
- Partner with data scientists, analysts, and application teams for analytics enablement.
- Provide L2/L3 support for production pipelines and troubleshoot failures.
- Mentor junior engineers and contribute to best practices documentation.