SystemsSheet 07 of 08
AI harnessing and fine-tuning
Most of an AI system's quality comes from the harness around the model: the instructions, tools, retrieval, guardrails and evaluations. We build and measure the harness first. When the evaluations show a model still falls short, we fine-tune one on your data and serve it privately.
- Order
- Harness, then knowledge, then fine-tuning
- Methods
- LoRA and QLoRA, supervised tuning, DPO, distillation
- Serving
- Private endpoints or your own GPUs
The improvement flywheel
Switch the environment to see the services it runs on. Play the walkthrough, or pick a step.
Mosaic AI training and Model Serving, with MLflow evaluation and tracing next to governed data.
See it run
The same golden set scores every version, so the decision to fine-tune, and which adapter to ship, rests on numbers.
Task accuracy on a 200-case golden set
Training loss, LoRA v2
Rank 16 adapters on 3,400 curated examples. Test cases kept out of training.
Component view
The same system as an exploded 3D drawing. Each plate is one step, and each block is one component in one of four materials.
Drag to rotate. Select a layer to inspect it.
This drawing needs WebGL, which is turned off in this browser. The parts list describes every layer.
Parts list
Numbered bottom to topWe choose the model by task, cost, latency and where it is allowed to run: a hosted model behind a private endpoint, or an open-weight model you control.
- Hosted model AI step
- Open-weight model AI step
Typical toolingClaude, GPT, Gemini, Llama, Mistral, Qwen, gpt-oss
Versioned instructions, tools with typed inputs, and output schemas the code can validate. This is where most quality gains come from, and it is ordinary, testable code.
- Instructions and skills Deterministic code
- Tool definitions Deterministic code
- Output schemas Deterministic code
Typical toolingPrompt templates in Git, JSON Schema, structured outputs, MCP tools
Inputs and outputs are checked against your rules. Deterministic filters run first, and a small classifier model handles the cases rules cannot describe.
- Input filters Deterministic code
- PII masking Deterministic code
- Policy classifier AI step
Typical toolingPresidio, allow-lists and pattern rules, safety classifiers such as Llama Guard
A fixed set of real tasks with expected results. Every change to the instructions, model or tools is scored before release. An LLM judge scales the scoring, and experts calibrate the judge.
- Golden test set Deterministic code
- LLM judge AI step
- Expert rating Human review
Typical toolingMLflow evaluation, promptfoo, Ragas, custom scorers
Training examples come from production traces and expert corrections. Code removes duplicates and sensitive data, and keeps test examples out of training.
- Trace mining Deterministic code
- Expert labelling Human review
- Split and de-duplicate Deterministic code
Typical toolingTrace tables, Label Studio or Argilla, dataset versioning
Parameter-efficient fine-tuning such as LoRA teaches a model your formats, vocabulary and decisions. Preference tuning such as DPO aligns it with expert choices. Adapters are versioned like code.
- Supervised tuning AI step
- Preference tuning AI step
- Adapter registry Deterministic code
Typical toolingHugging Face TRL and PEFT, Databricks Mosaic AI training, SageMaker, Vertex AI tuning
The tuned model is served on a private endpoint, with the previous version kept for rollback. Quality, cost and drift are tracked against the same evaluation set.
- Private endpoint Deterministic code
- Version routing Deterministic code
- Drift monitor Deterministic code
Typical toolingvLLM, Databricks Model Serving, SageMaker endpoints, MLflow
The order we try things
Each step is cheaper and easier to reverse than the one after it, so we only move on when the evaluation set shows we need to.
Improve the harness
Instructions, tools, output schemas and examples. The cheapest thing to change and the easiest to test.
Add knowledge
Retrieval or an OKF bundle when the model lacks facts. Facts that change should be looked up, not trained in.
Fine-tune
When the model has the facts but not the behaviour: a strict output format, domain vocabulary, consistent decisions, or a smaller model that must match a larger one.
Distil and serve
Train a smaller model on a larger model's approved outputs to lower cost and response time, then serve it privately.
When fine-tuning is worth it
Worth it
- Output must follow a strict format every time
- Domain language that general models handle poorly
- High volume, where a small tuned model can replace a large one
- Classification or extraction with thousands of labelled examples
Not yet
- Facts that change often. Use retrieval instead.
- Fewer than a few hundred good examples
- No evaluation set to measure the improvement
Where it runs
The drawing stays the same. The services change with the environment you already run, and the feasibility analysis picks the fit.
Databricks
AWS
Azure
Google Cloud
Your own GPUs
Feasibility first
Before anything is built, we check whether this system is worth building for you, and where it should run.
- The task, its current error rate, and what a good answer looks like.
- Whether an evaluation set exists, or can be built from real cases.
- How many good examples are available, and who can label more.
- Model licensing, hosting location and data restrictions.
- Serving volume, response time and cost targets.
Built before
Work delivered by Ashish Adhikari, who leads engineering at YoursSherpa.
- RAG and agentic RAG on Databricks Apps, with model serving and evaluation on Databricks tools.Databricks
- Research on generative adversarial networks for anomaly and malware detection, published in Advances in Artificial Intelligence Research (2024).Publication
- World finalist in Microsoft Imagine Cup 2021 with a computer vision and sensor system on Azure.Microsoft Imagine Cup
Start with a feasibility call
Tell us about one process or data problem. In the first call we will say which parts we would automate with code, which need an agent, and which we would leave alone.
Send a short note through the contact form and we will set up the call.
What helps us prepare
- The process or system you have in mind, and who works on it today.
- Where the data lives: cloud, platform and main tools.
- Security or hosting rules we need to work within.