Databricks & Lakehouse
Production hands-on: PySpark and Spark SQL, Delta Lake medallion architecture (Bronze / Silver / Gold), Databricks Jobs, notebooks and cluster sizing; partitioning, join strategy, caching and broadcast joins tuned against real workloads for both runtime and cost. Microsoft Fabric (OneLake, Lakehouse, Warehouse, Data Factory, Direct Lake) equally in production.
Data Migration
Source-system migration and cutover as delivered work: scoping what a legacy model actually computes, target model design, transfer and transformation of business logic, historical backfill and restatement, and reconciliation batteries (control totals, record-level diffs) that prove source-to-target parity before sign-off.
SAP BW & Source Systems
Familiar with the SAP product line and BW concepts: InfoObjects, InfoProviders / DSO, star-schema extracts and transformation logic, but no hands-on BW development; stated plainly. What transfers is the harder half: reading an undocumented legacy warehouse, mapping its data structures and embedded business rules, and reproducing them provably in a lakehouse. Enterprise sources worked with directly: Oracle, PostgreSQL, MSSQL, HRIS and CRM mirrors, REST APIs and file feeds.
Data Modelling & BI
Dimensional modelling: star and galaxy schemas with conformed customer, product and transaction dimensions, SCD handling, standardised metric definitions. Reusable marts and analytics-ready datasets. Power BI (semantic models, advanced DAX, Direct Lake, RLS), the layer that consumes the migrated model, so target design accounts for how it will be queried.
Cloud
Azure in production: ADLS Gen2, Azure SQL, Key Vault, Data Factory. AWS in working use; S3 and Athena for operational revenue data. Storage, compute separation, partitioning and cost-aware scanning patterns are the same on either side.
LLM-based Tooling
Daily practice: LLM-assisted translation and refactoring of legacy SQL and stored procedure logic into PySpark, transformation scaffolding, reverse-engineering of undocumented models, and generation of test and reconciliation harnesses, every output validated against source data rather than trusted.
SQL & Python SQL
15+ yrs, advanced — complex multi-source joins, CTEs, window functions, incremental and idempotent MERGE logic, execution plans and query optimisation (Spark SQL, PostgreSQL, Athena, T-SQL, Oracle). Python 7+ yrs: PySpark daily, plus pandas, API clients and automation. Airflow DAG design, dbt, Git-based Dev / Test / Prod promotion.
Data Quality
Validation and integrity rules inside the pipeline, reconciliation batteries, coverage and freshness guards, diff checks, failure and SLA alerting; data dictionaries, lineage documentation and data contracts between layers; root-cause investigation with the fix applied at source, not patched downstream.
Delivery & Domains
Requirements discovery and gap analysis directly with business, risk, compliance and finance owners; user stories and acceptance criteria; documentation and handover. Financial services, payments, AML, regulatory reporting, US retirement and wealth, digital lending; enterprise manufacturing (BMW).