YoursSherpa

SystemsSheet 06 of 08

Data lakes and lakehouses with quality built in

A lake keeps every file you receive, in open formats, at low cost. A lakehouse adds tables, transactions and governance on top. We build them with the Medallion pattern (Bronze, Silver and Gold), so raw data is never lost and business data is always tested.

Pattern
Medallion: Bronze, Silver, Gold
Formats
Delta Lake, Apache Iceberg
Built before
Healthcare lakehouse on Databricks and AWS

Medallion lakehouse with a quality gate

Switch the environment to see the services it runs on. Play the walkthrough, or pick a step.

Deterministic codeAI stepHuman reviewYour existing systemAI assist, approved by a person
    Medallion lakehouse with a quality gate

    Open-format storage in your cloud, Bronze to Gold, with a quality gate, quarantine and governance end to end.

    Delivered on Databricks and AWS: Auto Loader, Delta Lake, Unity Catalog, and Snowflake publishing.

    See it run

    A nightly load from landing to Gold. The quality gate passes with warnings, quarantines 318 rows with reasons, and the change summariser explains the batch.

    Pipeline runNightly lakehouse load, Bronze to GoldIllustrative example
    Rows landed1,204,388
    Silver rows1,203,977
    Quarantined with reasons318
    Gold tables refreshed2
    Run time14 min 22 s

    Change summariser: "Claims volume is 3% above the 8-week average, driven by one partner's late February resubmission. No schema changes."

    Component view

    The same system as an exploded 3D drawing. Each plate is one step, and each block is one component in one of four materials.

    Drag to rotate. Select a layer to inspect it.

    Parts list

    Numbered bottom to top
    1. Files land in your own cloud storage. Arrival events start the pipeline, so nothing waits for a schedule unless it should.

      • Object storage Your existing system
      • Arrival events Deterministic code
      • Landing checks Deterministic code

      Typical toolingAmazon S3, Azure Data Lake Storage, Google Cloud Storage, EventBridge, Event Grid

    2. Data exactly as received, stored as tables with ingestion time, source and file hash. Nothing is transformed, so any later step can be re-run from here.

      • Raw tables Deterministic code
      • Audit columns Deterministic code
      • Schema registry Deterministic code

      Typical toolingDelta Lake or Apache Iceberg, Auto Loader, Glue crawlers

    3. IDs, codes, dates and names are standardised and history is tracked. When a batch looks unusual, an LLM summarises what changed for the engineer on call.

      • Standardisation Deterministic code
      • SCD Type 2 history Deterministic code
      • Change summariser AI step

      Typical toolingPySpark, SQL MERGE, Delta change data feed

    4. Data moves from Silver to Gold only after it passes its checks. Failing rows are quarantined, and data owners decide on exceptions.

      • Expectation suites Deterministic code
      • Quarantine Deterministic code
      • Owner sign-off Human review

      Typical toolingGreat Expectations, Lakeflow (DLT) expectations, Soda, dbt tests

    5. Business aggregates such as member months, claims roll-ups and risk scores, stored and clustered for the filters people use most.

      • Business aggregates Deterministic code
      • Clustered tables Deterministic code
      • Feature tables Deterministic code

      Typical toolingDelta Lake, Z-ordering or liquid clustering, feature tables

    6. One catalog controls who can see what, down to rows and columns, and records lineage from raw file to report. An LLM drafts table and column descriptions for owners to approve.

      • Catalog and permissions Deterministic code
      • Lineage Deterministic code
      • Description drafter AI step

      Typical toolingUnity Catalog, AWS Lake Formation, Microsoft Purview, Dataplex

    7. Gold tables are shared with the warehouse, partners, BI tools and machine learning without copying data more than needed.

      • Warehouse sync Deterministic code
      • Delta Sharing Deterministic code
      • BI and ML consumers Your existing system

      Typical toolingDelta Sharing, Snowflake COPY INTO, Iceberg REST catalog, BI connectors

    Deterministic codeAI stepHuman reviewYour existing system
    Sheet06 of 08
    Layers7
    SystemData lake and lakehouse
    Drawn byA. Adhikari
    IssuedSeptember 2026
    ScaleNot to scale

    Open formats, in your own storage

    Your data stays in your cloud account in open table formats, so any engine can read it and you are not tied to one vendor.

    ChoiceOptionsHow we decide
    Table formatDelta Lake, Apache IcebergWhich engines must read and write the tables
    EngineDatabricks, Spark, Athena or Trino, Snowflake, BigQueryWorkload size, team skills and existing contracts
    CatalogUnity Catalog, AWS Glue and Lake Formation, Polaris, DataplexWhere permissions and lineage need to live
    QualityGreat Expectations, Lakeflow expectations, Soda, dbt testsWhere checks run and who acts on failures

    Where it runs

    The drawing stays the same. The services change with the environment you already run, and the feasibility analysis picks the fit.

    Environment
    What we use
    When it fits

    Databricks on AWS, Azure or Google Cloud

    Built before
    What we useDelta Lake, Auto Loader, Unity Catalog, Lakeflow
    When it fitsThe healthcare standardisation lakehouse ran here.

    AWS native

    What we useS3, Glue, Athena, Lake Formation, Apache Iceberg
    When it fitsServerless, pay-per-query lakes.

    Azure

    What we useADLS, Fabric OneLake, Azure Databricks, Purview
    When it fitsMicrosoft-centred teams.

    Google Cloud

    What we useCloud Storage, BigLake, Dataplex, Dataproc
    When it fitsBigQuery-centred teams.

    Snowflake

    What we useIceberg tables on external volumes
    When it fitsOpen files in your storage, queried from Snowflake.

    Feasibility first

    Before anything is built, we check whether this system is worth building for you, and where it should run.

    • What data you receive, its formats and volumes, and how long it must be kept.
    • Who consumes it: BI, data science, partners, regulators.
    • Cloud, table format and catalog choices, and what they cost at your volume.
    • Security, residency and retention rules.
    • The migration path from any existing lake or warehouse.
    What you receiveA lakehouse design with layer contracts, quality gates and access model, a running-cost estimate, and a go or no-go recommendation.

    Built before

    Work delivered by Ashish Adhikari, who leads engineering at YoursSherpa.

    • Medallion lakehouse on Databricks and AWS for US healthcare data: Bronze, Silver and Gold Delta tables, SCD Type 2 for members and providers, Great Expectations gates, a schema registry, Z-ordered Gold tables and Snowflake publishing.Abacus Insights
    • Databricks Certified Data Engineer Associate, covering Spark, Delta Lake, Delta Live Tables, Workflows and Unity Catalog.Databricks, 2024

    Start with a feasibility call

    Tell us about one process or data problem. In the first call we will say which parts we would automate with code, which need an agent, and which we would leave alone.

    Send a short note through the contact form and we will set up the call.

    What helps us prepare

    • The process or system you have in mind, and who works on it today.
    • Where the data lives: cloud, platform and main tools.
    • Security or hosting rules we need to work within.