The first four modules build an app: a user story, the scenarios that test it, the state behind its screens and the typed API behind them. This module carries the same principles somewhere new. The Sean Dinwiddie’s Webmastery team also trains and evaluates small language models, and that work is written in Python.
Module 5 teaches the Python practice the team holds to, the geometry its training and evaluation are built on, and the path a trained model takes before it is released.
Why Python
The stack follows the project. The team builds its apps in Haskell, TypeScript and Rust, and the work of training a model lives in Python. The libraries that load a model, fine-tune it, measure it and convert it for a runtime are written for Python first, and the research behind them is published with Python beside it.
So the training work is written in Python, and the team’s principles come along for the ride. The lecture Functional Programming in Other Languages makes the same case for carrying one discipline from language to language.
The same principles
- Behavior is named before it is built. A user story names what a feature does for someone before any code is written. Here, the behavior a model must show is named first: the corners of an experiment, and the oracle each answer must pass, are fixed before any data is generated or any training runs.
- Pure functions sit at the core, and effects at the edges. Generating data, scoring answers and comparing two models are pure functions over plain data. Reading files, loading a model and running a GPU sit in thin layers around them, so the logic is tested without any of the three.
- Tests carry names. Each test names the rule it proves, the way a scenario names a behavior, so a failing run reads as a list of broken rules, and a run with a skipped test fails.
- Evidence comes before a release. A trained model is released only when its evidence shows it is no worse than the model it replaces, wherever it is measured, the way a feature launches only after the review against its written scope.
What geometric reasoning means here
A model’s answers have traits that come in pairs: plain or formal, brief or expansive, warm or reserved, cautious or bold. Four such traits give sixteen combinations, and they sit at the corners of a four-dimensional cube, where two corners are as far apart as the number of traits they differ in.
Training and evaluation both reason about that cube:
- Training pairs are built between corners at every distance, from neighbours to opposites, so the model learns each trait apart from the facts it states.
- Evaluation measures every corner, adds whether the evidence in the prompt is stated or contested, and reads the results as effects and interactions rather than one average score.
- A model is released only when no corner is worse than before.
How the module runs
Three lessons, each building on the one before:
- Python, the team’s way sets out the practices, each with a small example that runs.
- Geometric reasoning in model training builds the cube, the training pairs and the analysis in plain Python.
- From fine-tune to release follows one model from the choice of its base to the gate that decides whether it ships.
The examples run on the Python standard library and pytest, except two sketches, marked as such: the training stages, which need a GPU, and the conversion that follows them.
Module 5 sits beside the course’s three subjects as the team’s Python and model-training practice. It assumes no machine learning, only the habits the earlier modules teach: name the behavior, keep the core pure and let the evidence decide.