Data Engineer

oraclecloud

Mumbai 7 Years Exp Posted 1h ago

Job Description

  • Build and maintain Data Factory pipelines for API-based data ingestion (REST, OData).

  • Develop Python ETL scripts for custom extraction, API polling, and file ingestion.

  • Develop and maintain Spark/Python notebooks (Python/PySpark) and SQL-based transformation logic.

  • Create Dataflows Gen2 (Power Query) for handling Excel/CSV-based messy data.

  • Implement incremental loads, watermarking, and CDC patterns.

  • Integrate data from Dataverse and other enterprise sources into Fabric.

  • Own Bronze → Silver transformations (cleansing, deduplication, standardisation)

  • Support Silver → Gold transformations (star schema, fact/dimension modelling)

  • Prepare feature datasets for AI/ML workloads and support notebook-based model development

  • Implement data quality frameworks (validation rules, anomaly detection)

  • Build and optimise Power BI semantic models (Direct Lake)

  • Prepare datasets for reporting and basic analytics use cases.

  • Monitor pipeline performance, platform health, and Fabric capacity (F64+), including cost optimisation.

  • Maintain documentation (data dictionaries, pipeline inventory, architecture decisions)

    • Manage Fabric capacity — pause/resume scheduling, scaling, and cost optimisation

Similar Openings for You