Data Platform Engineer III
stancorpglobal
Job Description
• Build and operate batch and streaming ingestion pipelines on Databricks (PySpark, Lakeflow/Delta Live Tables).
• Integrate event streams from Confluent Kafka into the Lakehouse and manage schema evolution.
• Build and operate Lakebase-backed operational data patterns alongside the lakehouse.
• Build ingestion and parsing pipelines for structured and unstructured content (documents, text) feeding the lakehouse, graph and vector stores.
• Operate and maintain Neo4j environments under established runbooks: backups and restores, monitoring, health checks and routine upgrades.
• Support vector index operations in Databricks Vector Search and Azure AI Search: run index refresh and sync pipelines and monitor retrieval data freshness.
• Execute disaster recovery procedures for platform data stores: backup/restore testing, failover drills and multi-region runbook execution.
• Implement lineage capture and freshness monitoring against published data contracts and SLAs.
• Automate data-quality and contract testing in CI/CD so violations fail fast, and respond to pipeline incidents.
• Maintain catalog and metadata hygiene in Unity Catalog, and contribute runbooks and operational documentation for the pipelines you own.