SystemsSheet 01 of 08
RAG and agentic RAG on your own knowledge
Search that answers from your documents, shows where each answer came from, and respects who is allowed to see what. We first delivered this on Databricks Apps with Databricks tooling. The same design runs on AWS, Azure, Google Cloud, Snowflake or a private cloud, and the feasibility analysis decides which.
- First delivered on
- Databricks Apps
- Runs on
- Databricks, AWS, Azure, Google Cloud, Snowflake, private cloud
- Knowledge options
- RAG, OKF bundles, or both
Agentic RAG reference architecture
Switch the environment to see the services it runs on. Play the walkthrough, or pick a step.
Delivered on Databricks Apps with Vector Search, Model Serving, Unity Catalog permissions and MLflow tracing.
See it run
A question that needs both policy text and claims data. The agent plans, retrieves under the user's permissions, runs one SQL query, and cites every claim.
- Planned two sub-questions: what changed in policy, and how many requests were affected0.9 s
- Searched the policy library with your permissions: 14 passages, 5 kept after reranking1.6 s
- Ran one SQL query on claims_gold.prior_auth_requests for Q2 20261.1 s
- Checked citations: 3 of 3 point to real passages0.2 s
Two rules changed on 1 February 2026. Lumbar spine MRI now needs six weeks of documented conservative therapy 1, and a repeat scan within 12 months needs a specialist referral 2. In Q2 2026, 1,284 outpatient MRI requests were reviewed under the new rules and 212 were returned for missing documentation 3.
RAG designs we build
Four patterns, from simple to agentic. We start with the simplest one that passes your evaluation set, and add steps only where it falls short.
Component view
The same system as an exploded 3D drawing. Each plate is one step, and each block is one component in one of four materials.
Drag to rotate. Select a layer to inspect it.
This drawing needs WebGL, which is turned off in this browser. The parts list describes every layer.
Parts list
Numbered bottom to topWe connect to what you already have and leave it where it is. Connectors read on a schedule or when content changes, and every document keeps a link back to its source.
- Wikis and SharePoint Your existing system
- PDFs and contracts Your existing system
- Tickets and email Your existing system
- Business tables Your existing system
Typical toolingUnity Catalog volumes, S3, Azure Blob, Cloud Storage, SharePoint and Confluence APIs
Code, no model. PDFs, slides, HTML and scans become clean text with headings, tables and page numbers kept. Sensitive fields are masked before anything is embedded.
- Layout-aware parser Deterministic code
- PII masking Deterministic code
- De-duplication Deterministic code
Typical toolingUnstructured, PyMuPDF, Azure Document Intelligence, Amazon Textract, Presidio
Chunks follow the document's own sections. An LLM adds tags such as document type, product and effective date, which become filters at query time. Curated knowledge can also be kept as an OKF bundle.
- Section chunker Deterministic code
- Metadata tagger AI step
- OKF concept pages Deterministic code
Typical toolingSection-aware splitters, a small LLM for tagging, OKF concept pages in Git
Each chunk is embedded and stored with its metadata and permissions. A keyword index sits beside the vector index, because exact terms such as policy numbers and codes are found better by keyword.
- Embedding model AI step
- Vector index Deterministic code
- Keyword index Deterministic code
Typical toolingDatabricks Vector Search, OpenSearch, Azure AI Search, pgvector, Qdrant
The user's permissions filter results before ranking, so nobody sees a passage they could not open in the source system. Hybrid search finds candidates and a reranker orders them.
- Permission filter Deterministic code
- Hybrid search Deterministic code
- Reranker AI step
Typical toolingDocument-level ACLs, hybrid BM25 and vector search, cross-encoder rerankers
This layer turns RAG into agentic RAG. The agent splits a hard question into parts, picks the source or tool for each part, and retrieves again when the evidence is thin. Simple questions skip it.
- Planner AI step
- SQL and API tools Deterministic code
- Self-check AI step
Typical toolingMosaic AI Agent Framework, LangGraph, tool calling, MCP tools
The model answers only from the retrieved passages. Code then checks that every citation points to a real passage, and guardrails block answers that break policy or expose restricted data.
- Answering model AI step
- Citation check Deterministic code
- Guardrails Deterministic code
Typical toolingDatabricks Model Serving, Amazon Bedrock, Azure OpenAI, open-weight models
Delivered as a Databricks App, a web app, a Teams or Slack bot, or an API. Every answer is traced. An evaluation set and reviewer feedback are run before each release, so changes are measured before users see them.
- Chat app or API Deterministic code
- Evaluation set Deterministic code
- Reviewer feedback Human review
Typical toolingDatabricks Apps, MLflow tracing and evaluation, FastAPI, Streamlit, React
Classic RAG or agentic RAG
Most teams need less agent than they expect. We start with the simpler pipeline and add agent steps where the evaluation set shows it misses.
Classic RAG
One retrieval pass, one answer. Layer 6 is skipped.
- Policy and procedure lookup
- Product and support documentation
- Questions answered by one or two passages
- Predictable cost and response time per question
- Easier to test, because the path is fixed
Agentic RAG
Plans, retrieves, checks, and retrieves again.
- Questions that span several documents or systems
- Questions that need a number from a table and a rule from a document
- Research tasks with follow-up steps
- Higher cost per question, capped by step limits
- Every step traced, so reviewers can see why it answered as it did
The OKF option for curated knowledge
OKF and RAG solve different problems, and they work well together.
The Open Knowledge Format (OKF) is an open, vendor-neutral specification published by Google Cloud in June 2026, with version 0.2 following in July. Knowledge is written as Markdown files with YAML frontmatter and grouped into bundles, with index pages that link concepts together. An agent reads the index, follows links to the concepts it needs, and loads only those pages.
RAG finds similar text in large collections. OKF states how curated knowledge fits together: definitions, business rules, metric logic and runbooks. We usually recommend both. An OKF bundle holds the curated core, and RAG covers the long tail of documents, tickets and email. The bundle can also be indexed in the vector store, so both paths share one source of truth in Git.
knowledge/
index.md
claims/
index.md
member-months.md
ibnr-reserve.md
policies/
index.md
prior-authorization.md--- type: Metric title: Member months description: Members enrolled in each month. The denominator for every PMPM figure. tags: [enrollment, finance] status: stable --- Counted from enrollment spans, with retroactive changes applied. Related: [PMPM](pmpm.md), [IBNR](ibnr-reserve.md)
type is required by the spec.OKF fits
- Definitions and rules that must be answered the same way every time
- Knowledge a team owns and can review through pull requests
- Data dictionaries and metric logic that agents use to write SQL
RAG fits
- Large collections that change often
- Content nobody will curate by hand
- Search across PDFs, tickets and email
Where it runs
The drawing stays the same. The services change with the environment you already run, and the feasibility analysis picks the fit.
Databricks
Built beforeAWS
Azure
Google Cloud
Snowflake
Private cloud or on-premises
Feasibility first
Before anything is built, we check whether this system is worth building for you, and where it should run.
- Where the documents live, how often they change, and who owns them.
- How permissions work in the source systems, and whether they can be carried into the index.
- Which questions people actually ask. We collect 50 to 100 real ones to build the evaluation set.
- Data residency, and which models are approved for your region and data.
- Expected volume, response time, and the monthly running cost at that volume.
- Whether classic RAG is enough, which questions need agent steps, and what belongs in an OKF bundle.
Built before
Work delivered by Ashish Adhikari, who leads engineering at YoursSherpa.
- RAG and agentic RAG built and deployed on Databricks Apps, using Databricks tools for retrieval, model serving and governance.Databricks
- Graphify, an open-source skill that turns code and documents into a queryable knowledge graph for AI coding assistants. Tested at about 50% fewer tokens on large codebases.Open source
Start with a feasibility call
Tell us about one process or data problem. In the first call we will say which parts we would automate with code, which need an agent, and which we would leave alone.
Send a short note through the contact form and we will set up the call.
What helps us prepare
- The process or system you have in mind, and who works on it today.
- Where the data lives: cloud, platform and main tools.
- Security or hosting rules we need to work within.