💳 Secure Payment

Full-Service Web & Software Agency · Klamath Falls and Redding

Python, the team’s way

Work with Sean

Python lets a program do almost anything from anywhere: change a global, read a file halfway through a calculation, shuffle with whatever seed the process happens to hold. The Sean Dinwiddie’s Webmastery team gives most of that freedom up, gladly, for something better: the same inputs give the same outputs on every machine, and a failing test can say exactly which rule broke.

Eight habits get it there. Each comes with a small example that runs on the standard library and pytest.

One layout, run as modules

Code is grouped by what it does, one package per domain, and nothing sits loose at the top of the repository. Configuration and data live beside the package, never inside it:

requirements.in      what the project uses, unpinned
requirements.lock    every package at an exact version, compiled from it
pytest.ini
config/train.json    a manifest of small config parts
config/parts/
data/
pipeline/
  config/compose.py
  data/build_rows.py
  data/contract.py
  evaluation/
  release/
tests/
bin/lock
bin/verify

Every entry point runs as a module from the repository’s root, so its imports are absolute and resolve the same way from a terminal, a test or a scheduled job. An entry point’s main() returns an exit code, and the module hands it to the shell:

# pipeline/data/build_rows.py
"""Print training rows as JSON lines, one per row."""

from __future__ import annotations

import json
import sys

from pipeline.data.contexts import CATALOG
from pipeline.data.order import rotate


def build_rows(count: int) -> tuple[dict[str, str], ...]:
    return tuple(
        {"id": f"row-{index:03d}", "place": rotate(CATALOG, index).place}
        for index in range(count)
    )


def main(argv: list[str] | None = None) -> int:
    args = sys.argv[1:] if argv is None else argv
    count = int(args[0]) if args else 4
    for row in build_rows(count):
        print(json.dumps(row, sort_keys=True))
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
$ python -m pipeline.data.build_rows
{"id": "row-000", "place": "the north trail"}
{"id": "row-001", "place": "the east parking lot"}
{"id": "row-002", "place": "the library annex"}
{"id": "row-003", "place": "the north trail"}

Run it as a file instead, and the same code falls over. Python puts the file’s own folder first on the import path, not the root:

$ python pipeline/data/build_rows.py 2>&1 | tail -1
ModuleNotFoundError: No module named 'pipeline'
Where an import starts looking A folder tree: the repository's root holds pipeline/, which holds data/, which holds build_rows.py. Run with python -m from the root, Python looks first in the current directory, the root, and a highlighted line shows import pipeline finding pipeline/ there. Run as a file, it looks first in data/, which holds no pipeline/, so the import fails. the rootpipeline/data/build_rows.pypython -mlooks here firstimport pipeline: founda file run looks herefirst: no pipeline/
Python puts one folder first on its import path. With python -m it is the current directory, here the root, so import pipeline finds the package (the violet line). Run as a file, it is the file’s own folder, data/, and there is no pipeline/ in there.

A locked environment

Each project has its own virtual environment, installed from a lock. People edit requirements.in, a short list of what the project uses.

uv compiles it into requirements.lock: every package and every dependency pinned, for a fixed Python version, with nothing published after a fixed date. Compile it again next month and you get the same lock:

#!/bin/sh
# Compile requirements.in into requirements.lock. With --check, fail if the lock is stale.
set -eu
compile() {
  uv pip compile requirements.in --quiet --no-header \
    --python-version 3.11 --exclude-newer 2026-09-01 --output-file "$1"
}
case "${1:-}" in
  --check)
    fresh=$(mktemp)
    trap 'rm -f "$fresh"' EXIT
    compile "$fresh"
    diff -u --label requirements.lock --label "compiled now" requirements.lock "$fresh" &&
      echo "requirements.lock is current" ;;
  *)
    compile requirements.lock ;;
esac

With --check, it compiles a fresh lock beside the real one and fails on any difference. Add a package to requirements.in without a new lock, and the run stops before anything installs it:

$ bin/lock --check
requirements.lock is current
$ echo six >> requirements.in
$ bin/lock --check; echo "exit $?"
--- requirements.lock
+++ compiled now
@@ -8,3 +8,5 @@
     # via pytest
 pytest==9.1.1
     # via -r requirements.in
+six==1.17.0
+    # via -r requirements.in
exit 1

The environment is made from the lock alone. --no-deps installs exactly what the lock names, and nothing it doesn’t:

$ python3 -m venv .venv
$ .venv/bin/python -m pip install --no-deps -r requirements.lock

Pure functions, the same answer every time

A function that reads only its arguments and returns a new value can be tested with plain data and trusted on any machine. Data generation holds to that strictly.

A row’s content is chosen by its index, turning through a fixed catalog like a dial; nothing is drawn at random. Where an order has to be shuffled, the shuffle gets a generator of its own, seeded from the run’s settings. The global one belongs to every library in the process, and any of them can disturb it:

# pipeline/data/order.py
"""Deterministic choices: the same inputs give the same rows on every machine."""

from __future__ import annotations

import random
from typing import Sequence, TypeVar

T = TypeVar("T")


def rotate(items: Sequence[T], index: int) -> T:
    """The item for row `index`: a fixed rotation, never a random draw."""
    return items[index % len(items)]


def seeded_order(count: int, seed: int, epoch: int) -> tuple[int, ...]:
    """Row indices in a shuffled order that a seed and an epoch fix."""
    rng = random.Random(f"order:{seed}:{epoch}")  # its own generator, not the global one
    indices = list(range(count))
    rng.shuffle(indices)
    return tuple(indices)

The hash() trap

For strings, Python’s built-in hash() is salted afresh in every process, so anything derived from it changes from run to run. The seeded order doesn’t budge:

$ for run in 1 2 3; do python -c "print(hash('north trail') % 1000)"; done
928
563
205
$ for run in 1 2 3; do python -c "from pipeline.data.order import seeded_order; print(seeded_order(8, seed=11, epoch=0))"; done
(0, 2, 6, 3, 7, 1, 4, 5)
(0, 2, 6, 3, 7, 1, 4, 5)
(0, 2, 6, 3, 7, 1, 4, 5)

A little housekeeping

Three small rules let two runs be compared line by line:

  • output is sorted;
  • JSON is written with sorted keys;
  • floats are rounded before they are compared.

Heavy libraries are imported inside the functions at the edge that need them, so a test of the core never loads a GPU stack.

The lecture What Is a Function? sets out what makes a function pure, and Functional Programming in Other Languages draws the line between the pure core and the shell around it.

Frozen data

Records are frozen dataclasses. A frozen record can’t change after it is made, so a value passed into a function is the same value when the function returns. A change is a new record, made with dataclasses.replace:

# pipeline/data/contexts.py
"""The made-up situations that training rows are written about."""

from __future__ import annotations

from dataclasses import dataclass


@dataclass(frozen=True)
class Context:
    place: str
    state: str
    reason: str


CATALOG = (
    Context("the north trail", "closed until Friday", "a washed-out bridge"),
    Context("the east parking lot", "open from 7 in the morning", "new gate hours"),
    Context("the library annex", "closed on Mondays", "a change in staffing"),
)
from dataclasses import FrozenInstanceError, replace

from pipeline.data.contexts import CATALOG

trail = CATALOG[0]
try:
    trail.state = "open"
except FrozenInstanceError as error:
    print("refused:", error)

reopened = replace(trail, state="open")
print(reopened)
print(trail)
refused: cannot assign to field 'state'
Context(place='the north trail', state='open', reason='a washed-out bridge')
Context(place='the north trail', state='closed until Friday', reason='a washed-out bridge')

The trail in the catalog is still closed until Friday; the reopened one is a new record beside it.

JSON as the one authority for config

Every value a run depends on lives in JSON: the model and its pinned revision, the data files, the seed, the training settings. Nothing repeats it. Code reads it, and documents point to it rather than copying its values.

A long config is hard to review, so it is composed from small parts that a manifest names, each set at a key or merged at the top:

{
  "parts": [
    {"file": "parts/base.json", "operation": "set", "target": "model"},
    {"file": "parts/data.json", "operation": "set", "target": "data"},
    {"file": "parts/run.json", "operation": "merge"}
  ]
}

The composer is strict. It refuses:

  • a key set twice;
  • a part outside the config’s folder;
  • a manifest that includes itself.

So composing never quietly overwrites a value or reaches for a file it shouldn’t:

# pipeline/config/compose.py
"""Compose one JSON config from small part files that a manifest names."""

from __future__ import annotations

import json
import sys
from pathlib import Path
from typing import Any


def load_json(path: Path) -> Any:
    return json.loads(path.read_text(encoding="utf-8"))


def place(document: dict[str, Any], entry: dict[str, Any], part: Any) -> dict[str, Any]:
    match entry["operation"]:
        case "set":
            key = entry["target"]
            if key in document:
                raise ValueError(f"duplicate target {key!r}")
            return {**document, key: part}
        case "merge":
            overlap = sorted(set(document) & set(part))
            if overlap:
                raise ValueError(f"duplicate keys {overlap}")
            return {**document, **part}
        case other:
            raise ValueError(f"unknown operation {other!r}")


def compose(manifest: Path, seen: tuple[Path, ...] = ()) -> dict[str, Any]:
    manifest = manifest.resolve()
    if manifest in seen:
        raise ValueError(f"cycle through {manifest.name}")
    folder = manifest.parent
    document: dict[str, Any] = {}
    for entry in load_json(manifest)["parts"]:
        path = (folder / entry["file"]).resolve()
        if not path.is_relative_to(folder):
            raise ValueError(f"part outside the config folder: {entry['file']}")
        raw = load_json(path)
        part = compose(path, (*seen, manifest)) if "parts" in raw else raw
        document = place(document, entry, part)
    return document


def main(argv: list[str] | None = None) -> int:
    args = sys.argv[1:] if argv is None else argv
    print(json.dumps(compose(Path(args[0])), indent=2, sort_keys=True))
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
$ python -m pipeline.config.compose config/train.json
{
  "data": {
    "train": "data/train.jsonl",
    "validation": "data/validation.jsonl"
  },
  "epochs": 2,
  "model": {
    "name": "example/base-model",
    "revision": "pinned"
  },
  "seed": 11
}
One config composed from parts The manifest, config/train.json, names three parts. base.json is set at the key model, data.json at the key data, and run.json is merged at the top, adding epochs and seed. Highlighted lines carry each part down into one composed config that holds model, data, epochs and seed. train.jsonbase.jsonset at modelmodeldata.jsonset at datadatarun.jsonmerged at topepochs, seedthe composed config
The manifest names each part and where it goes: two are set at a key, and one is merged at the top. The violet lines carry each part into the one config a run reads.

The lecture Practical Applications of Functional Programming builds configuration from explicit parts too, with an explicit precedence and a frozen result; the composer here is stricter and refuses any key set twice.

Contracts that report every fault

Data crosses a boundary only after it is checked against its contract:

  • the keys a row must carry, and the one it may (the oracle that From fine-tune to release adds);
  • the allowed values;
  • the rules between fields.

The checks are independent of each other, so all of them run, and the validator raises once with every fault, not just the first. Whoever fixes the data file gets the whole list in one go.

Counts and other derived facts are recomputed from the data, never trusted as declared:

# pipeline/data/contract.py
"""Check training rows against their contract and report every fault at once."""

from __future__ import annotations

import json
import sys
from collections import Counter
from pathlib import Path
from typing import Any, Iterable, Mapping

ROW_KEYS = frozenset({"id", "split", "prompt", "chosen", "rejected"})
OPTIONAL_KEYS = frozenset({"oracle"})  # what its answers must and must not say
SPLITS = frozenset({"train", "validation", "development", "release"})


def check(passed: bool, message: str) -> tuple[str, ...]:
    return () if passed else (message,)


def row_errors(row: Mapping[str, Any]) -> tuple[str, ...]:
    keys = set(row)
    return (
        *check(keys <= ROW_KEYS | OPTIONAL_KEYS, f"unexpected keys {sorted(keys - ROW_KEYS - OPTIONAL_KEYS)}"),
        *check(keys >= ROW_KEYS, f"missing keys {sorted(ROW_KEYS - keys)}"),
        *check(row.get("split") in SPLITS, f"unknown split {row.get('split')!r}"),
        *check(bool(str(row.get("prompt", "")).strip()), "empty prompt"),
        *check(row.get("chosen") != row.get("rejected"), "chosen and rejected are the same"),
    )


def count_errors(declared: Mapping[str, int], rows: Iterable[Mapping[str, Any]]) -> tuple[str, ...]:
    """Recount the rows in each split rather than trusting the declared counts."""
    found = Counter(row.get("split") for row in rows)
    return tuple(
        f"{split}: declared {declared.get(split, 0)}, found {found.get(split, 0)}"
        for split in sorted(set(declared) | set(found), key=str)
        if declared.get(split, 0) != found.get(split, 0)
    )


def validate(rows: Iterable[Mapping[str, Any]], declared: Mapping[str, int]) -> None:
    rows = tuple(rows)
    errors = (
        *(f"{row.get('id', '?')}: {error}" for row in rows for error in row_errors(row)),
        *count_errors(declared, rows),
    )
    if errors:
        raise ValueError("\n".join(errors))


def main(argv: list[str] | None = None) -> int:
    rows_path, counts_path = sys.argv[1:] if argv is None else argv
    rows = [json.loads(line) for line in Path(rows_path).read_text().splitlines() if line]
    validate(rows, json.loads(Path(counts_path).read_text()))
    print(f"{len(rows)} rows match their contract")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
from pipeline.data.contract import validate

rows = (
    {"id": "row-000", "split": "train", "prompt": "Is the north trail open?",
     "chosen": "It is closed until Friday.", "rejected": "Closed till Friday!"},
    {"id": "row-001", "split": "trian", "prompt": " ", "note": "draft",
     "chosen": "Open from 7 a.m.", "rejected": "Open from 7 a.m."},
)
try:
    validate(rows, declared={"train": 2})
except ValueError as error:
    print(error)
row-001: unexpected keys ['note']
row-001: unknown split 'trian'
row-001: empty prompt
row-001: chosen and rejected are the same
train: declared 2, found 1
trian: declared 0, found 1

The lecture Practical Applications of Functional Programming accumulates errors the same way in TypeScript.

Tests named for the behavior

Each test’s name states the rule it proves, so a failing run reads as a list of broken rules.

A contract is tested by breaking a valid row one way at a time and checking that each break is refused, by name. Several breaks at once have to come back in one error:

# tests/test_contract.py
import pytest

from pipeline.data.contract import validate

VALID = {
    "id": "row-000",
    "split": "train",
    "prompt": "Is the north trail open?",
    "chosen": "It is closed until Friday.",
    "rejected": "Closed till Friday!",
}

BROKEN = (
    ("unknown-split", {"split": "trian"}, "unknown split 'trian'"),
    ("empty-prompt", {"prompt": " "}, "empty prompt"),
    ("same-answers", {"rejected": VALID["chosen"]}, "chosen and rejected are the same"),
    ("extra-key", {"note": "draft"}, r"unexpected keys \['note'\]"),
)


def test_a_valid_row_passes():
    validate((VALID,), declared={"train": 1})


@pytest.mark.parametrize("change, error", [case[1:] for case in BROKEN], ids=[case[0] for case in BROKEN])
def test_each_broken_row_is_refused_by_name(change, error):
    with pytest.raises(ValueError, match=error):
        validate(({**VALID, **change},), declared={"train": 1})


def test_a_wrong_declared_count_is_refused():
    with pytest.raises(ValueError, match="train: declared 2, found 1"):
        validate((VALID,), declared={"train": 2})


def test_every_fault_is_reported_in_one_error():
    with pytest.raises(ValueError) as raised:
        validate(({**VALID, "split": "trian", "prompt": ""},), declared={"train": 1})
    assert len(str(raised.value).splitlines()) == 4

A test that has to fail

A test that has never failed has proved nothing, so a test can break the code on purpose. Here one test swaps in a broken rotation and expects the rotation test to fail. The last two hold the shuffled order to its seed and its epoch:

# tests/test_rows.py
import pytest

import pipeline.data.build_rows as build
from pipeline.data.contexts import CATALOG
from pipeline.data.order import seeded_order


def test_rows_rotate_through_the_catalog():
    places = [row["place"] for row in build.build_rows(4)]
    assert places == [context.place for context in (*CATALOG, CATALOG[0])]


def test_the_rotation_test_fails_when_the_rotation_breaks(monkeypatch):
    monkeypatch.setattr(build, "rotate", lambda items, index: items[0])  # broken on purpose
    with pytest.raises(AssertionError):
        test_rows_rotate_through_the_catalog()


def test_the_same_seed_and_epoch_give_the_same_order():
    assert seeded_order(8, seed=11, epoch=0) == seeded_order(8, seed=11, epoch=0)


def test_a_new_epoch_gives_a_new_order_of_the_same_rows():
    first, second = seeded_order(8, seed=11, epoch=0), seeded_order(8, seed=11, epoch=1)
    assert first != second and sorted(first) == sorted(second)

The run itself

Every run reads its settings from pytest.ini at the root: where the tests are, no cache left behind, and the classic output without a progress percentage:

[pytest]
testpaths = tests
addopts = -p no:cacheprovider
console_output_style = classic
$ python -m pytest -v --no-header tests/test_contract.py tests/test_rows.py
============================= test session starts ==============================
collecting ... collected 11 items

tests/test_contract.py::test_a_valid_row_passes PASSED
tests/test_contract.py::test_each_broken_row_is_refused_by_name[unknown-split] PASSED
tests/test_contract.py::test_each_broken_row_is_refused_by_name[empty-prompt] PASSED
tests/test_contract.py::test_each_broken_row_is_refused_by_name[same-answers] PASSED
tests/test_contract.py::test_each_broken_row_is_refused_by_name[extra-key] PASSED
tests/test_contract.py::test_a_wrong_declared_count_is_refused PASSED
tests/test_contract.py::test_every_fault_is_reported_in_one_error PASSED
tests/test_rows.py::test_rows_rotate_through_the_catalog PASSED
tests/test_rows.py::test_the_rotation_test_fails_when_the_rotation_breaks PASSED
tests/test_rows.py::test_the_same_seed_and_epoch_give_the_same_order PASSED
tests/test_rows.py::test_a_new_epoch_gives_a_new_order_of_the_same_rows PASSED

============================== 11 passed in 0.02s ==============================

A skipped test is a test that didn’t run, and a green run with a skip in it hides that. So conftest.py fails any run with a skip:

# tests/conftest.py
"""A skipped test is a test that did not run, so a run with one fails."""


def skipped(config):
    reporter = config.pluginmanager.get_plugin("terminalreporter")
    return len(reporter.stats.get("skipped", ())) if reporter else 0


def pytest_terminal_summary(terminalreporter, config):
    if skipped(config):
        terminalreporter.write_line(f"{skipped(config)} skipped: every test must run")


def pytest_sessionfinish(session):
    if skipped(session.config):
        session.exitstatus = 1

With a skipped test added, the run reports it and exits non-zero:

$ python -m pytest -q; echo "exit $?"
....................s..................
1 skipped: every test must run
38 passed, 1 skipped in 0.08s
exit 1

One script verifies everything

One script runs every check that must hold before any paid compute starts:

  • the lock is current;
  • the config composes;
  • the data meets its contract;
  • every test passes.

It reads and never writes. Regenerating data is a separate command, so verifying can never change what it verifies:

#!/bin/sh
# Everything that must hold before any paid compute starts. It reads and never writes.
set -eu
cd "$(dirname "$0")/.."
bin/lock --check
python -m pipeline.config.compose config/train.json > /dev/null
echo "config composes"
python -m pipeline.data.contract data/rows.jsonl data/counts.json
python -m pytest -q
echo "verified"
$ bin/verify
requirements.lock is current
config composes
6 rows match their contract
......................................
38 passed in 0.11s
verified

Next, these habits go to work on the cube that the team’s training data and evaluation share.

Copyright Sean Paul Payne Dinwiddie
All Rights Reserved