Senior Data Engineer - Databricks & Streaming - Healthcare AI (Onsite, Evening Shift, Lahore, PKR Salary)
HR POD - Hiring Talent Globally- Location
- Lahore, Pakistan
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted yesterday
Requirements
- 4+ years of experience in data engineering, with substantial production experience in Databricks.
- Strong experience with Spark SQL, PySpark, Delta Lake, Medallion Architecture, and Delta Live Tables (DLT).
- Hands-on experience with Structured Streaming or equivalent production-grade streaming ingestion using Azure Event Hubs, Kafka, or Kinesis.
- Strong understanding of checkpoint recovery, watermarking, and deduplication strategies.
- Demonstrated experience debugging source-to-warehouse data discrepancies.
- Ability to walk through a real-world incident involving mismatched record counts and explain how the root cause was identified and resolved.
- Proven experience integrating third-party REST APIs in production.
- Experience handling pagination edge cases, rate and row limits, retries, and schema drift.
- Experience with entity resolution or data matching involving messy, real-world text data.
- Experience with a metrics or semantic layer such as Holistics AML/AQL, dbt Metrics, or LookML.
- Working understanding of why non-additive measures cannot be reliably calculated from pre-aggregated rollups.
- Strong SQL and Python skills, with the ability to own data pipelines end-to-end with minimal oversight.
- Strong written and spoken English, with the ability to collaborate effectively with a US-based team asynchronously.
- Healthcare data experience, including referrals, payer taxonomy, claims/eligibility, or other PHI-adjacent datasets.
- Familiarity with HIPAA handling expectations.
- Experience with voice-agent, call-center, or telephony/conversation data.
- Familiarity with call transcripts and containment or outcome metrics.
- Hands-on experience with Holistics, specifically AML/AQL modeling.
- Experience with the broader Azure ecosystem beyond Event Hubs, including ADLS, ADF, and Key Vault.
Responsibilities
- Build and maintain resilient ingestion pipelines for third-party vendor REST APIs.
- Work primarily with voice-AI observability and telephony platforms.
- Handle different pagination schemes, including offset/limit and page/cursor models.
- Manage row and rate limits through time-windowing and adaptive bisection.
- Implement robust schema-drift handling through contract and column-presence checks.
- Ensure alerts are triggered when fields are renamed, moved, or removed rather than silently propagating null values.
- Own streaming ingestion from Azure Event Hubs into Databricks using Structured Streaming and/or Auto Loader.
- Manage checkpoints and offsets, watermarking, and at-least-once deduplication.
- Perform source-parity reconciliation across ingested and production data.
- Investigate row counts, dropped or duplicated events, late-arriving data, and schema mismatches.
- Identify and resolve the root cause when ingested data does not match production sources.
- Develop and maintain Delta Lake pipelines using a Medallion Architecture (Bronze Silver Gold).
- Use Spark SQL and PySpark to build and maintain production data pipelines.
- Implement idempotent MERGE upserts.
- Work with Delta Live Tables and materialized-view constraints, including CREATE OR REFRESH and LIVE references.
- Understand and manage differences between DLT and job execution contexts.
- Build entity-resolution pipelines for dirty, free-text data.
- Normalize practice, provider, and payer names using regex, canonical dictionaries, fuzzy matching, confidence-scored crosswalks, and override tables.
- Maintain the semantic and metrics layer with rigorous metric definitions.
- Define and maintain accurate denominators, data grain, and cohort boundaries.
- Ensure the correct handling of non-additive aggregates, including medians and percentiles that cannot be reliably supported through aggregate-aware pre-aggregation.
- Ensure every metric remains accurate and reproducible.
- Instrument data quality across the entire pipeline.
- Monitor data freshness, source parity, data contracts, and other critical quality checks.
- Build alerting mechanisms that identify data issues before they reach dashboards.
Skills
- Databricks
- Spark
- SQL
- PySpark
- Delta Lake
- Delta Live Tables
- Azure Event Hubs
- Kafka
- AWS Kinesis
- dbt
- Python
- HIPAA
- Azure
- Azure Key Vault
- Cursor
More jobs at HR POD - Hiring Talent Globally
All 12Engineering Manager (Onsite, Lahore, USD Salary)
HR POD - Hiring Talent Globally · Lahore, Pakistan · 3d ago
Full Stack Developer - Chauffeur Booking Platform (Onsite, Night Shift, Lahore, PKR Salary)
HR POD - Hiring Talent Globally · Lahore, Pakistan · 11d ago
Senior AI/ML Engineer (Remote, EST, Anywhere in Pakistan, USD Salary)
HR POD - Hiring Talent Globally · Remote (EST) · Pakistan · 13d ago
Senior Integration Engineer (Remote, Anywhere in Pakistan, EUR Salary)
HR POD - Hiring Talent Globally · Pakistan · 13d ago
CRO Specialist (Remote, Anywhere in Pakistan, USD Salary)
HR POD - Hiring Talent Globally · Lahore, Pakistan · 15d ago
Similar roles
Senior AI Data Engineer
Strategic Systems International · Mexico · Argentina · Lahore, Pakistan · today
Senior Data Analyst
Strategic Systems International · Lahore, Pakistan · 3d ago
Data Engineer
Talentmanagementsolution · PER - Karachi, PK · PER - Lahore, PK · PER - Islamabad, PK · 3d ago
Senior AI/ML & Data Engineer - Islamabad
Devs · Lahore, Punjab, Pakistan · 9d ago
Senior Data Scientist
inDrive · Islamabad, Pakistan · 11d ago
Senior Data Engineer (Pyspark, Databricks)
Strategic Systems International · Lahore, Pakistan · 14d ago