The first four modules build an app: a user story, the scenarios that test it, the state behind its screens and the typed API behind them. Module 5 puts an agent loop to work on a business process, with tools, a skill and a person approving anything that leaves the building. This module opens up every join that loop crosses, and the model at its center, one boundary at a time, with the same principles.
The Sean Dinwiddie’s Webmastery team works on both sides of those joins: the apps and APIs the course builds, and the small language models it trains and evaluates in Python. The meetup talk, next, runs the module’s first half in twenty-five minutes (the SI in its title is simply the name federal agencies now use for AI), and The AI protocol map starts with the joins themselves.
What geometric reasoning and tandem mean here
Geometric reasoning, in this course, is discrete: the behaviors to build and measure are points in a declared design space, one level for each factor, with a distance between any two of them. A reply’s three traits, brief or detailed, formal or casual, cautious or bold, make one such space: eight points, and two replies that differ in one trait are one step apart. A statistician calls the space a two-level full factorial; a coder calls the distance Hamming distance. Geometric reasoning as data draws the line between this sense and the others in circulation.
Tandem means deterministic code runs beside a model on every request, so the model proposes and the code decides.
A contract at every join
One rule runs through the module: every join between an app, an agent, a model and the tools around it carries a contract, and something checks it. Here is each join the module opens, what crosses it and what checks it, with the lesson that teaches it.
- App to tool. A tool call and its result, checked against the tool’s input and output schemas (The AI protocol map).
- Client to API. A request and its response, typed by one Servant type, with a drift check holding the committed OpenAPI document to that type (The AI protocol map) and a decoder on every response the client receives (Endpoints at the boundary).
- Design file to reader. Factors and a table of points, read by one shared reader in each language that refuses a broken design and reports every fault at once (Geometric reasoning as data).
- Harness to model. A prompt out, and a choice and a reply back: the choice checked against one offer list, and the reply sent exactly as written, or held back with every reason that applies (The tandem harness).
- Server to app. Two rounds of one booking in MCP’s result types, with a signed request state used once and every reply decoded strictly (Server and client in tandem).
- Harness to extension. A new rule, skill or tool, added only at a documented extension point, each with a scenario, and a new kind of question added to one list the compiler holds every screen to (Extending the harness).
- File to program. Data rows and config, checked by contracts that report every fault, the config read once into a frozen record (Python, the team’s way).
- Model to measurement. Answers to paired items, each pair scored together (Measuring reasoning on the geometry).
- Serving to training. The prompts serving sent, logged as sent and read back by training, never rendered a second time (Geometric reasoning in model training).
- Training to release. A trained model, checked by a gate against the model in service: both scored on the same pairs, and released only when, at every point, a drop past a declared margin is ruled out (From fine-tune to release).
To try it, write one more entry for a join in a project of your own: a booking form that posts to a calendar, say, or a price list a script reads each morning. Name what crosses it and what checks it today. A join with nothing to name in that last part is the next one to give a contract.
The same principles
- Behavior is named before it is built. A user story names what a feature does for someone before any code is written. Here, a join’s contract is written before either side of it, and the behavior a model must show is named first: the points of an experiment, and the oracle each answer must pass, are fixed before any data is generated or any training runs.
- Pure functions sit at the core, and effects at the edges. Composing a prompt, generating data, scoring answers and comparing two models are pure functions over plain data. Calling a tool, crossing the wire and running a model sit in thin rings around them, so the logic is tested without any of the three.
- Tests carry names. Each test names the rule it proves, the way a scenario names a behavior, so a failing run reads as a list of broken rules, and in the Python suites a run with a skipped test fails.
- Evidence comes before a release. A trained model is released only when its evidence rules out, at every point it is measured, a drop past a declared margin from the model it replaces, the way a feature launches only after the review against its written scope.
The course’s pure core grows outward here, one ring for each kind of join, and each ring is checked where something crosses into it.
Two halves, in this order
The stack follows the project. The team builds its apps in Haskell, TypeScript and Rust, and the work of training a model lives in Python: the libraries that load a model, fine-tune it, measure it and convert it for a runtime are written for Python first, and the research behind them is published with Python beside it. So the module runs in two halves.
- The joins, in TypeScript and Haskell. The AI protocol map names the joins, Geometric reasoning as data builds what crosses them, The tandem harness checks it, Server and client in tandem carries it over the wire, and Extending the harness shows where both kinds of harness grow.
- The model, in Python. Python, the team’s way lays the foundation, Measuring reasoning on the geometry writes the measurement before any training, as a scenario comes before its code, Geometric reasoning in model training builds the training set on the geometry, and From fine-tune to release decides whether a model ships.
Each lesson leans on the one before. The lecture Functional Programming in Other Languages makes the same case for carrying one discipline from language to language.
How the module runs
Nine lessons after this opener and its meetup talk, each building on the one before, with examples that run: Vitest 4 or later for TypeScript (checked here with 5.0.3), Hspec for Haskell and pytest for Python, on the standard library wherever it will do. The TypeScript runs on Node.js 22.18 or later on the 22 line, or 24.3 or later, which run it directly and without a warning. A step that can’t run on a laptop, such as a GPU training stage, is shown as a sketch and labelled one. Every business, price, score and setting in an example is invented; the numbers a protocol defines, such as its error codes, come from its specification.
Every page that teaches a protocol names the revision it teaches: the Model Context Protocol (MCP) at revision 2026-07-28, beside the earlier, handshake-based revisions its specification calls legacy, and the Agent2Agent protocol (A2A) at 1.0. The protocols are taught as the ecosystem a webmaster works in, and where a lesson describes the team’s own practice, it says so.
The module assumes no machine learning, only the habits the earlier modules teach: name the behavior, keep the core pure and let the evidence decide. From fine-tune to release closes the module, and the course, on a model released only on rules written before it ran.