Data Engineer
oraclecloud
Job Description
-
Build and maintain Data Factory pipelines for API-based data ingestion (REST, OData).
-
Develop Python ETL scripts for custom extraction, API polling, and file ingestion.
-
Develop and maintain Spark/Python notebooks (Python/PySpark) and SQL-based transformation logic.
-
Create Dataflows Gen2 (Power Query) for handling Excel/CSV-based messy data.
-
Implement incremental loads, watermarking, and CDC patterns.
-
Integrate data from Dataverse and other enterprise sources into Fabric.
-
Own Bronze → Silver transformations (cleansing, deduplication, standardisation)
-
Support Silver → Gold transformations (star schema, fact/dimension modelling)
-
Prepare feature datasets for AI/ML workloads and support notebook-based model development
-
Implement data quality frameworks (validation rules, anomaly detection)
-
Build and optimise Power BI semantic models (Direct Lake)
-
Prepare datasets for reporting and basic analytics use cases.
-
Monitor pipeline performance, platform health, and Fabric capacity (F64+), including cost optimisation.
-
Maintain documentation (data dictionaries, pipeline inventory, architecture decisions)
-
Manage Fabric capacity — pause/resume scheduling, scaling, and cost optimisation
-