This position blends hands-on development, solution design, and team leadership, playing a key role in both the initial platform build and its ongoing evolution.
Key Responsibilities:
Architecture & Design
- Drive the implementation of a medallion architecture (Bronze, Silver, Gold)
- Design scalable data models and reusable data products
- Optimize the Gold layer for reporting, semantic models, and analytical workloads
Data Engineering (Hands-on)
- Develop data transformation and enrichment logic using PySpark and SparkSQL
- Build and maintain end-to-end data pipelines
- Implement efficient incremental loading strategies (e.g. Delta merge, watermarking)
- Refactor and modernize legacy ETL processes
Orchestration & Platform
- Design and manage Fabric pipelines for scheduling and orchestration
- Replace legacy orchestration tools such as SQL Agent or SSIS
- Enable and support event-driven pipeline execution
Data Quality & Governance
- Establish data validation, monitoring, and quality assurance processes
- Ensure consistency, accuracy, and reliability across all data layers
DevOps & Best Practices
- Implement CI/CD pipelines for data platform components
- Promote Git-based version control and collaborative development workflows
Collaboration & Mentorship
- Partner with stakeholders to translate business requirements into data solutions
- Provide guidance and mentorship to junior engineers
- Support testing, migration activities, and documentation efforts
Requirements:
- 5-12 years of experience in data engineering or data platform development
- Hands-on experience with Microsoft Fabric, Databricks, or Azure Synapse
- Strong programming skills in Python / PySpark
- Advanced SQL capabilities, including complex transformations and optimisation
- Experience with Delta Lake (merge, partitioning, optimization)
- Proven experience building data pipelines and orchestration frameworks
- Solid understanding of dimensional modelling (e.g. star schema, SCD)
Nice to Have:
- Experience with Fabric-specific tools (Lakehouses, Dataflows Gen2)
- Knowledge of Power BI, semantic models, or DirectLake
- Familiarity with Azure DevOps and CI/CD pipelines
- Experience in data migration or modernisation initiatives
- Exposure to data product or data mesh concepts