YoursSherpa

SystemsSheet 07 of 08

AI harnessing and fine-tuning

Most of an AI system's quality comes from the harness around the model: the instructions, tools, retrieval, guardrails and evaluations. We build and measure the harness first. When the evaluations show a model still falls short, we fine-tune one on your data and serve it privately.

Order
Harness, then knowledge, then fine-tuning
Methods
LoRA and QLoRA, supervised tuning, DPO, distillation
Serving
Private endpoints or your own GPUs

The improvement flywheel

Switch the environment to see the services it runs on. Play the walkthrough, or pick a step.

Deterministic codeAI stepHuman reviewYour existing systemAI assist, approved by a person
    The improvement flywheel

    Traces, curation, tuning, evaluation, deployment and monitoring, turning around one golden evaluation set.

    Mosaic AI training and Model Serving, with MLflow evaluation and tracing next to governed data.

    See it run

    The same golden set scores every version, so the decision to fine-tune, and which adapter to ship, rests on numbers.

    Evaluation resultsSame golden set, every versionIllustrative example

    Task accuracy on a 200-case golden set

    0%25%50%75%100%Prompt v1: 64%64%Prompt v1Harness v2: 78%78%Harness v2+ Retrieval: 86%86%+ RetrievalLoRA v1: 90%90%LoRA v1LoRA v2: 93%93%LoRA v2

    Training loss, LoRA v2

    2.00.3step 0400

    Rank 16 adapters on 3,400 curated examples. Test cases kept out of training.

    VersionAccuracyFormat adherenceResponse time, p50Cost per 1,000 requests
    Prompt v1, large hosted model64%81%2.9 s$14.20
    Harness v2, same model78%97%3.1 s$15.10
    + Retrieval86%97%3.6 s$16.40
    LoRA v2, 8B model, private93%99%0.9 s$2.10

    Component view

    The same system as an exploded 3D drawing. Each plate is one step, and each block is one component in one of four materials.

    Drag to rotate. Select a layer to inspect it.

    Parts list

    Numbered bottom to top
    1. We choose the model by task, cost, latency and where it is allowed to run: a hosted model behind a private endpoint, or an open-weight model you control.

      • Hosted model AI step
      • Open-weight model AI step

      Typical toolingClaude, GPT, Gemini, Llama, Mistral, Qwen, gpt-oss

    2. Versioned instructions, tools with typed inputs, and output schemas the code can validate. This is where most quality gains come from, and it is ordinary, testable code.

      • Instructions and skills Deterministic code
      • Tool definitions Deterministic code
      • Output schemas Deterministic code

      Typical toolingPrompt templates in Git, JSON Schema, structured outputs, MCP tools

    3. Inputs and outputs are checked against your rules. Deterministic filters run first, and a small classifier model handles the cases rules cannot describe.

      • Input filters Deterministic code
      • PII masking Deterministic code
      • Policy classifier AI step

      Typical toolingPresidio, allow-lists and pattern rules, safety classifiers such as Llama Guard

    4. A fixed set of real tasks with expected results. Every change to the instructions, model or tools is scored before release. An LLM judge scales the scoring, and experts calibrate the judge.

      • Golden test set Deterministic code
      • LLM judge AI step
      • Expert rating Human review

      Typical toolingMLflow evaluation, promptfoo, Ragas, custom scorers

    5. Training examples come from production traces and expert corrections. Code removes duplicates and sensitive data, and keeps test examples out of training.

      • Trace mining Deterministic code
      • Expert labelling Human review
      • Split and de-duplicate Deterministic code

      Typical toolingTrace tables, Label Studio or Argilla, dataset versioning

    6. Parameter-efficient fine-tuning such as LoRA teaches a model your formats, vocabulary and decisions. Preference tuning such as DPO aligns it with expert choices. Adapters are versioned like code.

      • Supervised tuning AI step
      • Preference tuning AI step
      • Adapter registry Deterministic code

      Typical toolingHugging Face TRL and PEFT, Databricks Mosaic AI training, SageMaker, Vertex AI tuning

    7. The tuned model is served on a private endpoint, with the previous version kept for rollback. Quality, cost and drift are tracked against the same evaluation set.

      • Private endpoint Deterministic code
      • Version routing Deterministic code
      • Drift monitor Deterministic code

      Typical toolingvLLM, Databricks Model Serving, SageMaker endpoints, MLflow

    Deterministic codeAI stepHuman reviewYour existing system
    Sheet07 of 08
    Layers7
    SystemAI harness and fine-tuning
    Drawn byA. Adhikari
    IssuedSeptember 2026
    ScaleNot to scale

    The order we try things

    Each step is cheaper and easier to reverse than the one after it, so we only move on when the evaluation set shows we need to.

    1. Improve the harness

      Instructions, tools, output schemas and examples. The cheapest thing to change and the easiest to test.

    2. Add knowledge

      Retrieval or an OKF bundle when the model lacks facts. Facts that change should be looked up, not trained in.

    3. Fine-tune

      When the model has the facts but not the behaviour: a strict output format, domain vocabulary, consistent decisions, or a smaller model that must match a larger one.

    4. Distil and serve

      Train a smaller model on a larger model's approved outputs to lower cost and response time, then serve it privately.

    When fine-tuning is worth it

    Worth it

    • Output must follow a strict format every time
    • Domain language that general models handle poorly
    • High volume, where a small tuned model can replace a large one
    • Classification or extraction with thousands of labelled examples

    Not yet

    • Facts that change often. Use retrieval instead.
    • Fewer than a few hundred good examples
    • No evaluation set to measure the improvement

    Where it runs

    The drawing stays the same. The services change with the environment you already run, and the feasibility analysis picks the fit.

    Environment
    What we use
    When it fits

    Databricks

    What we useMosaic AI training, Model Serving, MLflow
    When it fitsTuning and serving next to governed data.

    AWS

    What we useSageMaker, Bedrock custom models
    When it fitsManaged training with private endpoints.

    Azure

    What we useAzure AI Foundry fine-tuning
    When it fitsAzure OpenAI and open models.

    Google Cloud

    What we useVertex AI tuning
    When it fitsGemini and open models.

    Your own GPUs

    What we useHugging Face TRL and PEFT, vLLM
    When it fitsFull control, and no data leaves your network.

    Feasibility first

    Before anything is built, we check whether this system is worth building for you, and where it should run.

    • The task, its current error rate, and what a good answer looks like.
    • Whether an evaluation set exists, or can be built from real cases.
    • How many good examples are available, and who can label more.
    • Model licensing, hosting location and data restrictions.
    • Serving volume, response time and cost targets.
    What you receiveAn evaluation set, baseline scores for the current harness, a recommendation on whether to fine-tune, and a training and serving plan if the answer is yes.

    Built before

    Work delivered by Ashish Adhikari, who leads engineering at YoursSherpa.

    • RAG and agentic RAG on Databricks Apps, with model serving and evaluation on Databricks tools.Databricks
    • Research on generative adversarial networks for anomaly and malware detection, published in Advances in Artificial Intelligence Research (2024).Publication
    • World finalist in Microsoft Imagine Cup 2021 with a computer vision and sensor system on Azure.Microsoft Imagine Cup

    Start with a feasibility call

    Tell us about one process or data problem. In the first call we will say which parts we would automate with code, which need an agent, and which we would leave alone.

    Send a short note through the contact form and we will set up the call.

    What helps us prepare

    • The process or system you have in mind, and who works on it today.
    • Where the data lives: cloud, platform and main tools.
    • Security or hosting rules we need to work within.