💳 Secure Payment

Full-Service Web & Software Agency · Klamath Falls and Redding

From fine-tune to release

Work with Sean

A trained model is released the way a feature launches: against evidence named before the work began. This lesson follows one small language model through the pipeline the Sean Dinwiddie’s Webmastery team trains with, from choosing a base model to the gate that decides whether it ships.

Each stage answers to something fixed in advance: a comparison, a split, an oracle, a record of its inputs, or the model it would replace. The training stages, which need a GPU, and the conversion that follows them appear as sketches; everything around them runs here.

From base to release Seven stages down a highlighted line: choose a base, generate the data, SFT then DPO, check the inputs against their record, merge and convert, evaluate as served, and the promotion gate, then release. Beside each stage is what it answers to, fixed in advance: one fair comparison, disjoint splits, a frozen reference, a SHA-256 digest per file, the quantization type, every oracle, and the parent, the model it would replace. stagefixed in advancechoose a baseone fair comparisongenerate the datadisjoint splitsSFT, then DPOa frozen referencecheck the inputsa SHA-256 per filemerge and convertquantization typeevaluate as servedevery oraclepromotion gatethe parentrelease
Each stage answers to something fixed before it runs, named on the right. The model moves down the violet line, and the promotion gate, against the model it would replace, decides whether it ships.

Choosing the base by a fair comparison

Fine-tuning starts from a small open-weight model, and choosing one is an experiment like any other. Every candidate answers the same prompts, through the same prompt template and runtime, with the same decoding settings, scored by the same scorer. Results from two different evaluations can’t be ranked against each other, so the ranking refuses them.

Within one evaluation, the rule is fixed before any result is seen and applied in order:

  1. the fewest answers that fail their oracle;
  2. then the closest agreement with the reference answers;
  3. then speed.
# pipeline/release/select.py
"""Choose a base model only from results that one evaluation produced."""

from __future__ import annotations

from dataclasses import dataclass
from typing import Iterable


@dataclass(frozen=True)
class Result:
    candidate: str
    evaluation: str  # names the prompt, decoding settings and scorer together
    failures: int  # answers that failed their oracle
    overlap: float  # agreement with the reference answers, 0 to 1
    seconds: float  # median time to answer


def rank(results: Iterable[Result]) -> tuple[str, ...]:
    results = tuple(results)
    evaluations = sorted({r.evaluation for r in results})
    if len(evaluations) != 1:
        raise ValueError(f"results from different evaluations: {evaluations}")
    ordered = sorted(results, key=lambda r: (r.failures, -r.overlap, r.seconds, r.candidate))
    return tuple(r.candidate for r in ordered)
from pipeline.release.select import Result, rank

results = (
    Result("candidate-a", "eval-1", failures=3, overlap=0.71, seconds=0.9),
    Result("candidate-b", "eval-1", failures=1, overlap=0.64, seconds=1.4),
    Result("candidate-c", "eval-1", failures=1, overlap=0.69, seconds=1.1),
)
print(rank(results))
try:
    rank((*results, Result("candidate-d", "eval-2", failures=0, overlap=0.80, seconds=0.5)))
except ValueError as error:
    print(error)
('candidate-c', 'candidate-b', 'candidate-a')
results from different evaluations: ['eval-1', 'eval-2']

Fixing the rule in advance keeps a favourite from winning on whichever measure flatters it. A base chosen for training is still only a starting point: what ships is decided at the end, by comparing the trained model with its parent.

Deterministic data in disjoint splits

Training data is generated, never written by hand: catalogs in JSON, combined with made-up situations and varied by index, as Python, the team’s way shows.

Every row about one situation lands in the same split (training, validation, development or release), so a situation seen in training never comes back, reworded, in a split that measures the model. Generated files are never edited by hand; a change to the data is a change to the generator.

# pipeline/release/splits.py
"""Disjoint splits, assigned by situation, and the one rule that guards release."""

from __future__ import annotations

from typing import Iterable, Mapping

PATTERN = ("train", "train", "train", "validation", "train", "train", "development", "release")


def split_of(situation_index: int) -> str:
    """Every row about one situation lands in the same split, on every run."""
    return PATTERN[situation_index % len(PATTERN)]


def training_rows(rows: Iterable[Mapping[str, object]]) -> tuple[Mapping[str, object], ...]:
    rows = tuple(rows)
    leaked = sorted({str(r["id"]) for r in rows if r["split"] in {"development", "release"}})
    if leaked:
        raise ValueError(f"held-out rows offered to training: {', '.join(leaked)}")
    return rows
from pipeline.release.splits import split_of, training_rows

print([split_of(i) for i in range(8)])
try:
    training_rows([
        {"id": "row-040", "split": "train"},
        {"id": "row-041", "split": "release"},
    ])
except ValueError as error:
    print(error)
['train', 'train', 'train', 'validation', 'train', 'train', 'development', 'release']
held-out rows offered to training: row-041

The release split is held back for the release decision and nothing else. It never guides training, tuning or the choice of a checkpoint. Once its evidence has informed a decision, it is retired for a fresh set the model has never been measured on, so it can’t slowly become one more thing the model was fitted to.

Oracles as the hard boundary

Overlap with a reference answer measures wording, and an oracle measures meaning. Each row carries one, and it names:

  • the facts the answer must state;
  • what it must never say;
  • whether the facts are missing, so the honest answer is that it doesn’t know;
  • a length it must stay within.

An answer that fails its oracle fails, whatever its other scores say:

# pipeline/release/oracle.py
"""A semantic oracle: what an answer must say, must not say, and must admit."""

from __future__ import annotations

import re
from dataclasses import dataclass


@dataclass(frozen=True)
class Oracle:
    required: tuple[str, ...] = ()  # each must appear
    forbidden: tuple[str, ...] = ()  # patterns that must not match
    unknown: bool = False  # the facts are missing, so the answer must say so
    max_words: int = 60


def oracle_errors(oracle: Oracle, answer: str) -> tuple[str, ...]:
    text = answer.lower().replace("\u2019", "'")  # a curly apostrophe counts too
    admits = "don't know" in text or "do not know" in text
    return (
        *(f"missing {p!r}" for p in oracle.required if p.lower() not in text),
        *(f"says {p!r}" for p in oracle.forbidden if re.search(p, answer, re.IGNORECASE)),
        *(() if admits or not oracle.unknown else ("does not admit it does not know",)),
        *(() if len(answer.split()) <= oracle.max_words else ("too long",)),
    )
{
  "id": "trail-0000-d4",
  "split": "train",
  "prompt": "Answer in a plain, brief, warm, cautious voice. The posted notice says the north trail is closed until Friday, because of a washed-out bridge. Is the north trail open?",
  "chosen": "Thanks for asking. The north trail is closed until Friday, because of a washed-out bridge. That is what the posted notice says.",
  "rejected": "Please note that the north trail is closed until Friday. The reason is a washed-out bridge. Plan around it.",
  "oracle": {"required": ["closed", "Friday", "washed-out bridge"], "forbidden": ["is open"]}
}
import json
from pathlib import Path

from pipeline.release.oracle import Oracle, oracle_errors

row = json.loads(Path("data/pair.json").read_text())
oracle = Oracle(**{key: tuple(value) for key, value in row["oracle"].items()})
print("chosen:  ", oracle_errors(oracle, row["chosen"]))
print("rejected:", oracle_errors(oracle, row["rejected"]))
print("negative:", oracle_errors(oracle, "Good news: the north trail is open."))
print("unknown: ", oracle_errors(Oracle(unknown=True), "The annex opens at nine."))
chosen:   ()
rejected: ()
negative: ("missing 'closed'", "missing 'Friday'", "missing 'washed-out bridge'", "says 'is open'")
unknown:  ('does not admit it does not know',)

The row above is a voice pair from Geometric reasoning in model training, so both of its answers pass the oracle: they differ only in voice. A hard negative is the other kind of pair, an answer built to be wrong, and it has to fail its oracle on meaning (here, it says the trail is open), not merely share fewer words with the reference.

Supervised, then preference training

Training runs in two stages on one LoRA adapter: a small set of added weights trained over a base model that stays fixed, held in 4-bit form so a model of a few billion parameters trains on one GPU.

  1. Supervised fine-tuning (SFT) teaches the answers, with the loss counted on the answer and not on the prompt.
  2. Direct preference optimization (DPO) then trains the same adapter on the pairs. It makes the chosen answer more likely than the rejected one relative to a frozen reference, the adapter as SFT left it, so the model learns the preference without drifting from what SFT taught.

The libraries are transformers, peft and trl, and every setting comes from the composed JSON config:

# A sketch of the two training stages, not run here: it needs a GPU and the
# model's weights. Every value comes from the composed JSON config, whose parts
# here also set "load", "adapter", "sft" and "dpo", and a "pairs" data file.
from pathlib import Path

from datasets import load_dataset
from peft import LoraConfig, PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from trl import DPOConfig, DPOTrainer, SFTConfig, SFTTrainer

from pipeline.config.compose import compose

config = compose(Path("config/train.json"))
name, revision = config["model"]["name"], config["model"]["revision"]
tokenizer = AutoTokenizer.from_pretrained(name, revision=revision)


def base():
    load = BitsAndBytesConfig(**config["load"])  # the base held in 4 bits
    return AutoModelForCausalLM.from_pretrained(name, revision=revision, quantization_config=load)


def rows(split):
    return load_dataset("json", data_files=config["data"][split], split="train")


# 1. Supervised fine-tuning: a new LoRA adapter, loss on the answer only.
#    completion_only_loss needs rows with a "prompt" and a "completion".
sft = SFTTrainer(
    model=base(),
    args=SFTConfig(output_dir="runs/sft", completion_only_loss=True, **config["sft"]),
    train_dataset=rows("train"),
    eval_dataset=rows("validation"),
    processing_class=tokenizer,
    peft_config=LoraConfig(**config["adapter"]),
)
sft.train()
sft.save_model("runs/sft/adapter")

# 2. Preference training: the same adapter learns from chosen and rejected
#    answers, measured against a frozen copy of where it started.
policy = PeftModel.from_pretrained(base(), "runs/sft/adapter", is_trainable=True)
dpo = DPOTrainer(
    model=policy,
    args=DPOConfig(output_dir="runs/dpo", **config["dpo"]),
    train_dataset=rows("pairs"),
    processing_class=tokenizer,
)
dpo.train()
dpo.save_model("runs/dpo/adapter")

In current trl, a model that already carries an adapter is trained against a frozen copy of that adapter, made when training starts. The run doesn’t take that on trust. Before the first step it records, as named evidence, that the adapter being trained and its reference start identical, and that the reference has nothing left to train.

Evidence that names the inputs

Before a trained adapter goes any further, the run shows what it was trained on. Its record names the inputs exactly:

  • the config;
  • the model’s pinned revision;
  • the package versions;
  • each data file, by its path, its size in bytes, its count of lines and a SHA-256 digest of its contents.

Each is checked again as a named check, what was expected against what was found, and any failure stops the run with the names of what failed:

# pipeline/release/promote.py
"""Promotion as named checks: each says what was expected and what was found."""

from __future__ import annotations

from dataclasses import dataclass
from typing import Any, Mapping, Sequence

from pipeline.evaluation.gate import gate_errors


@dataclass(frozen=True)
class Check:
    name: str
    expected: Any
    actual: Any

    @property
    def passed(self) -> bool:
        return self.expected == self.actual
# pipeline/release/promote.py, continued
def require(checks: Sequence[Check]) -> None:
    failed = [c for c in checks if not c.passed]
    if failed:
        raise ValueError("; ".join(f"{c.name}: expected {c.expected!r}, found {c.actual!r}" for c in failed))
# pipeline/release/inputs.py
"""Evidence that names the exact inputs a run used, checked again before release."""

from __future__ import annotations

import hashlib
from pathlib import Path

from pipeline.release.promote import Check


def describe(path: Path) -> dict[str, object]:
    data = path.read_bytes()
    return {
        "path": path.as_posix(),
        "bytes": len(data),
        "lines": data.count(b"\n"),
        "sha256": hashlib.sha256(data).hexdigest(),
    }


def input_checks(recorded: list[dict[str, object]]) -> tuple[Check, ...]:
    return tuple(
        Check(f"{entry['path']} unchanged", entry, describe(Path(str(entry["path"]))))
        for entry in recorded
    )
from pathlib import Path

from pipeline.release.inputs import describe, input_checks
from pipeline.release.promote import require

recorded = [describe(Path("data/train.jsonl"))]  # written when training starts
print(recorded)
require(input_checks(recorded))
print("inputs match the record")

rows = Path("data/train.jsonl")  # one id rewritten after training: same size, same lines
rows.write_text(rows.read_text().replace("row-001", "row-009"))
try:
    require(input_checks(recorded))
except ValueError as error:
    print(error)
[{'path': 'data/train.jsonl', 'bytes': 54, 'lines': 3, 'sha256': '1f1ea15caf02b22c8fbda2bb33f632e9ef7d231dc416c281fd870b88d9925ab8'}]
inputs match the record
data/train.jsonl unchanged: expected {'path': 'data/train.jsonl', 'bytes': 54, 'lines': 3, 'sha256': '1f1ea15caf02b22c8fbda2bb33f632e9ef7d231dc416c281fd870b88d9925ab8'}, found {'path': 'data/train.jsonl', 'bytes': 54, 'lines': 3, 'sha256': '474f637b110d99640405c7f06a0ea5c593c705bfc7d35a1e6966a2c62e0067c6'}

A failed check says which input changed and how. Here the size and the line count still match; only the digest catches the rewritten row. So the record can’t quietly describe a different run from the one that made the model.

Conversion, and checks at runtime

The released adapter is merged into the base weights at full precision, converted to GGUF with llama.cpp’s converter and quantized with llama-quantize, the file format and runtime the model ships in. The quantization type is a setting like any other, read from the config:

# A sketch, not run here: merge at full precision, then convert and quantize.
# full_precision_base() loads the same pinned base without the 4-bit form.
merged = PeftModel.from_pretrained(full_precision_base(), "runs/dpo/adapter").merge_and_unload()
merged.save_pretrained("runs/merged")
tokenizer.save_pretrained("runs/merged")

$ python llama.cpp/convert_hf_to_gguf.py runs/merged --outfile runs/model-f16.gguf --outtype f16
$ llama.cpp/build/bin/llama-quantize runs/model-f16.gguf runs/model.gguf <type>

Quantizing changes the weights, so it can change answers. The evaluation that decides the release runs on the converted file, through the runtime that will serve it, never only on the weights that came out of training.

The runtime gets checks of its own:

  • the model loads and answers;
  • output held to a grammar parses as the JSON it promises;
  • generation stops where it should;
  • a request that runs past its deadline ends cleanly;
  • requests made at the same time each get their own answer.

The promotion gate

Promotion compares the candidate with its parent under the same evaluation, and every condition is a named check:

  • the same evaluation as the parent;
  • no empty answers;
  • no answer that repeats its instructions;
  • the corner gate from the last lesson, with every corner measured, none worse and at least one better.

Any failure blocks promotion and says why:

# pipeline/release/promote.py, continued
def promotion_checks(candidate: Mapping[str, Any], parent: Mapping[str, Any]) -> tuple[Check, ...]:
    return (
        Check("same evaluation as the parent", parent["evaluation"], candidate["evaluation"]),
        Check("empty answers", 0, sum(not a.strip() for a in candidate["answers"])),
        Check("answers that repeat the instructions", 0,
              sum(candidate["instructions"] in a for a in candidate["answers"])),
        Check("corner gate", (), gate_errors(candidate["corners"], parent["corners"])),
    )
# tests/test_promote.py
import pytest

from pipeline.data.cube import corners
from pipeline.release.promote import promotion_checks, require

PARENT = {
    "evaluation": "eval-1",
    "instructions": "You answer questions about posted notices.",
    "answers": ("The north trail is closed until Friday.",),
    "corners": {c: 0.70 for c in corners(5)},
}


def candidate(**changes):
    return {**PARENT, "corners": {c: 0.71 for c in corners(5)}, **changes}


def test_a_candidate_better_at_every_corner_is_promoted():
    require(promotion_checks(candidate(), PARENT))


def test_an_empty_answer_blocks_promotion_by_name():
    with pytest.raises(ValueError, match="empty answers: expected 0, found 1"):
        require(promotion_checks(candidate(answers=("",)), PARENT))


def test_an_answer_that_repeats_its_instructions_blocks_promotion():
    leaked = ("You answer questions about posted notices. The trail is closed.",)
    with pytest.raises(ValueError, match="answers that repeat the instructions"):
        require(promotion_checks(candidate(answers=leaked), PARENT))


def test_one_worse_corner_blocks_promotion():
    worse = {**candidate()["corners"], "10110": 0.69}
    with pytest.raises(ValueError, match="corner gate"):
        require(promotion_checks(candidate(corners=worse), PARENT))


def test_a_different_evaluation_blocks_promotion():
    with pytest.raises(ValueError, match="same evaluation as the parent"):
        require(promotion_checks(candidate(evaluation="eval-2"), PARENT))
$ python -m pytest -v --no-header tests/test_promote.py
============================= test session starts ==============================
collecting ... collected 5 items

tests/test_promote.py::test_a_candidate_better_at_every_corner_is_promoted PASSED
tests/test_promote.py::test_an_empty_answer_blocks_promotion_by_name PASSED
tests/test_promote.py::test_an_answer_that_repeats_its_instructions_blocks_promotion PASSED
tests/test_promote.py::test_one_worse_corner_blocks_promotion PASSED
tests/test_promote.py::test_a_different_evaluation_blocks_promotion PASSED

============================== 5 passed in 0.01s ===============================

Each check is plain data, a name with what was expected and what was found, and the lecture Practical Applications of Functional Programming treats expected failure the same way: as a value to report, never an exception to swallow.

That is Module 5 in one model: Python written as pure functions over plain data, a cube that names every behavior to train and measure, and a release decided corner by corner against the model it would replace. Webmasters at any stage of the craft work under the Sean Dinwiddie’s Webmastery name, and Joining the team sets out the terms.

Copyright Sean Paul Payne Dinwiddie
All Rights Reserved