SystemsSheet 06 of 08
Data lakes and lakehouses with quality built in
A lake keeps every file you receive, in open formats, at low cost. A lakehouse adds tables, transactions and governance on top. We build them with the Medallion pattern (Bronze, Silver and Gold), so raw data is never lost and business data is always tested.
- Pattern
- Medallion: Bronze, Silver, Gold
- Formats
- Delta Lake, Apache Iceberg
- Built before
- Healthcare lakehouse on Databricks and AWS
Medallion lakehouse with a quality gate
Switch the environment to see the services it runs on. Play the walkthrough, or pick a step.
Delivered on Databricks and AWS: Auto Loader, Delta Lake, Unity Catalog, and Snowflake publishing.
See it run
A nightly load from landing to Gold. The quality gate passes with warnings, quarantines 318 rows with reasons, and the change summariser explains the batch.
Change summariser: "Claims volume is 3% above the 8-week average, driven by one partner's late February resubmission. No schema changes."
Component view
The same system as an exploded 3D drawing. Each plate is one step, and each block is one component in one of four materials.
Drag to rotate. Select a layer to inspect it.
This drawing needs WebGL, which is turned off in this browser. The parts list describes every layer.
Parts list
Numbered bottom to topFiles land in your own cloud storage. Arrival events start the pipeline, so nothing waits for a schedule unless it should.
- Object storage Your existing system
- Arrival events Deterministic code
- Landing checks Deterministic code
Typical toolingAmazon S3, Azure Data Lake Storage, Google Cloud Storage, EventBridge, Event Grid
Data exactly as received, stored as tables with ingestion time, source and file hash. Nothing is transformed, so any later step can be re-run from here.
- Raw tables Deterministic code
- Audit columns Deterministic code
- Schema registry Deterministic code
Typical toolingDelta Lake or Apache Iceberg, Auto Loader, Glue crawlers
IDs, codes, dates and names are standardised and history is tracked. When a batch looks unusual, an LLM summarises what changed for the engineer on call.
- Standardisation Deterministic code
- SCD Type 2 history Deterministic code
- Change summariser AI step
Typical toolingPySpark, SQL MERGE, Delta change data feed
Data moves from Silver to Gold only after it passes its checks. Failing rows are quarantined, and data owners decide on exceptions.
- Expectation suites Deterministic code
- Quarantine Deterministic code
- Owner sign-off Human review
Typical toolingGreat Expectations, Lakeflow (DLT) expectations, Soda, dbt tests
Business aggregates such as member months, claims roll-ups and risk scores, stored and clustered for the filters people use most.
- Business aggregates Deterministic code
- Clustered tables Deterministic code
- Feature tables Deterministic code
Typical toolingDelta Lake, Z-ordering or liquid clustering, feature tables
One catalog controls who can see what, down to rows and columns, and records lineage from raw file to report. An LLM drafts table and column descriptions for owners to approve.
- Catalog and permissions Deterministic code
- Lineage Deterministic code
- Description drafter AI step
Typical toolingUnity Catalog, AWS Lake Formation, Microsoft Purview, Dataplex
Gold tables are shared with the warehouse, partners, BI tools and machine learning without copying data more than needed.
- Warehouse sync Deterministic code
- Delta Sharing Deterministic code
- BI and ML consumers Your existing system
Typical toolingDelta Sharing, Snowflake COPY INTO, Iceberg REST catalog, BI connectors
Open formats, in your own storage
Your data stays in your cloud account in open table formats, so any engine can read it and you are not tied to one vendor.
| Choice | Options | How we decide |
|---|---|---|
| Table format | Delta Lake, Apache Iceberg | Which engines must read and write the tables |
| Engine | Databricks, Spark, Athena or Trino, Snowflake, BigQuery | Workload size, team skills and existing contracts |
| Catalog | Unity Catalog, AWS Glue and Lake Formation, Polaris, Dataplex | Where permissions and lineage need to live |
| Quality | Great Expectations, Lakeflow expectations, Soda, dbt tests | Where checks run and who acts on failures |
Where it runs
The drawing stays the same. The services change with the environment you already run, and the feasibility analysis picks the fit.
Databricks on AWS, Azure or Google Cloud
Built beforeAWS native
Azure
Google Cloud
Snowflake
Feasibility first
Before anything is built, we check whether this system is worth building for you, and where it should run.
- What data you receive, its formats and volumes, and how long it must be kept.
- Who consumes it: BI, data science, partners, regulators.
- Cloud, table format and catalog choices, and what they cost at your volume.
- Security, residency and retention rules.
- The migration path from any existing lake or warehouse.
Built before
Work delivered by Ashish Adhikari, who leads engineering at YoursSherpa.
- Medallion lakehouse on Databricks and AWS for US healthcare data: Bronze, Silver and Gold Delta tables, SCD Type 2 for members and providers, Great Expectations gates, a schema registry, Z-ordered Gold tables and Snowflake publishing.Abacus Insights
- Databricks Certified Data Engineer Associate, covering Spark, Delta Lake, Delta Live Tables, Workflows and Unity Catalog.Databricks, 2024
Start with a feasibility call
Tell us about one process or data problem. In the first call we will say which parts we would automate with code, which need an agent, and which we would leave alone.
Send a short note through the contact form and we will set up the call.
What helps us prepare
- The process or system you have in mind, and who works on it today.
- Where the data lives: cloud, platform and main tools.
- Security or hosting rules we need to work within.