Start Date : 01-Jun-2026
Primary Skill : ML,CONTRACTS,VERSIONING,SNOWFLAKE,PERFORMANCE TUNING,TECHNICAL DESIGN,DATA ACCESS,DATA INFRASTRUCTURE,APACHE,SCALA,ENTERPRISE DATA,PRAGMATIC,DISTRIBUTED DATA,TOOLING,ARCHITECTURE,UNITY,DISTRIBUTED,FRAMEWORKS,COST EFFICIENCY,ORCHESTRATION,PIPELINES,DATA QUALITY,GOVERNANCE,INFLUENCE,PIPELINE,ENGINEERING,RELIABILITY,CONTRACT,INFRASTRUCTURE,DATA PROCESSING,TESTING,DESIGN,PROCESSING,DEPLOYMENT,PRODUCTION,GITHUB,REAL-TIME,TRACK RECORD
ROLE OVERVIEW
We are looking for a Data Platform Lead to own the design, build, and governance of our enterprise data platform. This is a role within the Data COE, responsible for driving platform maturity across a federated, multi-cloud environment. You will set engineering standards, eliminate data duplication and drift across domains, and enable reliable, governed data access — without creating centralisation bottlenecks.
KEY RESPONSIBILITIES
- Design and evolve the enterprise data platform spanning Snowflake and Databricks across Azure and AWS
- Build and maintain production-grade ELT/ETL pipelines using Spark, dbt, Airflow, and Kafka
- Define and enforce platform standards for ingestion, transformation, storage, and access across federated domain teams
- Drive data quality, data contract adoption, and lineage visibility across the estate
- Own cloud storage architecture across S3 and ADLS; lead CI/CD practices for all platform assets
- Lead containerized workload deployment using Kubernetes for data platform services
- Coordinate with domain teams to align on standards; represent platform engineering in architecture forums
- Mentor mid-level engineers and conduct technical design reviews for platform-impacting changes
MUST-HAVE SKILLS
- Snowflake — expert-level: data modelling, performance tuning, role-based access, zero-copy sharing, dynamic data masking
- Databricks — production experience: Delta Lake, Unity Catalog, cluster management, Databricks Workflows
- Python and SQL — strong proficiency; Scala advantageous
- Apache Spark — distributed data processing at scale
- Apache Kafka — real-time event streaming, pipeline design, topic management
- Apache Airflow — pipeline orchestration, DAG design, dependency management
- dbt — transformation layer modelling, testing, documentation
- Cloud storage — S3 and ADLS; familiarity with Parquet, Delta, and partitioning strategies
- Kubernetes — container orchestration for data platform services
- CI/CD — GitHub Actions or equivalent for data platform assets (pipelines, schemas, infra)
- 8+ years of data engineering experience, with 3+ years in a lead or staff IC capacity
GOOD-TO-HAVE SKILLS
- Data quality frameworks — Great Expectations or equivalent, embedded in pipeline execution
- Feature Store platforms — Feast, Tecton, or equivalent; ML data infrastructure design
- Data catalogue and lineage tooling — Amundsen, DataHub, or equivalent
- Data contracts as an engineering practice — schema registries, versioning, producer/consumer agreements
- Open table formats — Apache Iceberg; lakehouse architecture patterns
- Cross-cloud experience — Snowflake (Azure), Databricks, and BigQuery integrations
WHAT WE'RE LOOKING FOR
- Systems thinker — able to see how pipeline, storage, and governance decisions interact at scale
- Able to influence without authority across autonomous, federated domain teams
- Opinionated on standards and quality, but pragmatic in how they are adopted across an existing estate
- Strong communicator — able to translate platform complexity to both engineering peers and business stakeholders
- Ownership-driven, with a track record of improving platform reliability, cost efficiency, and developer experience
Job Snapshot
Employment Type: Full Time
Minimun Education:
Bachelors
Location : Multiple Locations
Experience: At least 13 year(s)
Date Posted:
20-May-2026
Category:
Data Science
Remote:
No
Similar Open Jobs
About Us