Data Engineer

ice

Hyderabad NM Years Exp Posted 2h ago

Job Description

  • Own Apache Airflow end-to-end. You will write DAGs that handle complex multi-step workflows across ingestion, transformation, validation, and delivery. That means getting the fundamentals right: idempotency, backfill safety, SLA alerting, dynamic task mapping, and thorough testing. You will also review other engineers' DAGs and raise the standard across the team.
  • Keep Airflow running well on Kubernetes. You will manage the deployment using KubernetesExecutor or CeleryKubernetesExecutor, handle Helm upgrades, tune pod templates and resource limits, and deal with scheduler performance issues before they become incidents.
  • Build ETL and ELT pipelines that hold up in production. Sources will vary: APIs, message queues, databases, object storage. The expectation is that pipelines are fault-tolerant, incremental where it makes sense, and straightforward to debug when something goes wrong.
  • Treat data quality as part of the job, not an afterthought. Schema validation, row-count reconciliation, and anomaly detection should be baked into pipelines from the start. Freshness and accuracy are part of your definition of done.
  • Design and maintain a lakehouse architecture using Apache Iceberg as the primary table format, following a Bronze, Silver, and Gold medallion structure. Schema evolution, partitioning strategies, and time-travel queries will be regular concerns, not edge cases.
  • Use Databricks to manage cluster compute, Delta Lake workflows, Unity Catalog, and scheduled jobs. You will keep things cost-efficient, well-organised, and easy for others to navigate.
  • Build and maintain dbt projects that sit alongside the Airflow orchestration layer. Models, tests, snapshots, and macros should be well-structured and documented. Test failures and freshness checks should surface where people can act on them.
  • Contribute reusable components to the shared platform: pipeline templates, custom Airflow operators, and utility libraries that make the whole team faster and more consistent.
  • Partner with data scientists and ML engineers to deliver feature pipelines and training datasets with the freshness and reproducibility their models need.
  • Instrument your work. Pipelines should have metrics, dashboards, and alerts that surface real problems early. You will own runbooks and participate in on-call cover.
    • Play an active part in architecture discussions, tooling evaluations, and mentoring. Your experience should benefit the team, not just your own workstream.

Similar Openings for You