An average score can rise while a model gets worse at something that matters. Geometric reasoning guards against that. It places every behavior to be trained and measured at a corner of a cube, so the training data covers every corner, the evaluation measures every corner, and a release is judged corner by corner.
The geometry is small, exact and a joy to work through, and this lesson writes all of it in plain Python, on the practices of Python, the team’s way.
Four traits make a cube
Each trait of an answer’s voice is a pair of poles, and an answer takes one pole of each: plain or formal, brief or expansive, warm or reserved, cautious or bold. Four binary traits give 24 = 16 combinations.
Written as bits, 0 for a trait’s first pole and 1 for its second, they are the corners of a four-dimensional cube, and the distance between two corners is the number of traits they differ in, their Hamming distance:
# pipeline/data/cube.py
"""Binary traits of an answer's voice, and the cube their corners make."""
from __future__ import annotations
from collections import Counter
from itertools import combinations, product
from typing import Iterable
TRAITS = (
("plain", "formal"),
("brief", "expansive"),
("warm", "reserved"),
("cautious", "bold"),
)
def corners(dimensions: int = len(TRAITS)) -> tuple[str, ...]:
"""Every corner as a bit string: "0" is a trait's first pole, "1" its second."""
return tuple("".join(bits) for bits in product("01", repeat=dimensions))
def describe(corner: str, factors=TRAITS) -> str:
return ", ".join(factors[axis][int(bit)] for axis, bit in enumerate(corner))
def hamming(left: str, right: str) -> int:
"""How many traits two corners differ in."""
if len(left) != len(right):
raise ValueError("corners must have the same number of traits")
return sum(a != b for a, b in zip(left, right))
def pairs_by_distance(points: Iterable[str]) -> dict[int, int]:
counts = Counter(hamming(a, b) for a, b in combinations(tuple(points), 2))
return dict(sorted(counts.items()))
def gray_order(dimensions: int) -> tuple[str, ...]:
"""Every corner once, each one trait away from the one before."""
return tuple(format(i ^ (i >> 1), f"0{dimensions}b") for i in range(2**dimensions))
from pipeline.data.cube import corners, describe, pairs_by_distance
cube = corners()
print(len(cube), "corners")
print(cube[0], describe(cube[0]))
print(cube[-1], describe(cube[-1]))
print(pairs_by_distance(cube))
16 corners
0000 plain, brief, warm, cautious
1111 formal, expansive, reserved, bold
{1: 32, 2: 48, 3: 32, 4: 8}
The 120 pairs of corners fall into four distances: 32 neighbours one trait apart (the cube’s edges), 48 pairs two apart, 32 three apart and 8 opposites. A contract that declares those counts is checked by counting them again from the corners, never by trusting the numbers it states.
Training pairs that change only the voice
Preference training learns from pairs: a prompt, a chosen answer and a rejected one. If the two answers differ in their facts as well as their voice, the model can’t tell which difference it is being taught.
So each pair keeps the facts fixed and changes only the traits. The chosen answer is written in the voice the prompt asks for, and the rejected answer gives the same facts in the voice of another corner.
Every corner is paired at every distance, from a neighbour one trait away to its opposite. Which corner it meets at each distance turns with the row’s index, so over enough rows it meets them all, with no randomness:
# pipeline/data/pairs.py
"""Preference pairs that keep the facts and change only the voice."""
from __future__ import annotations
from pipeline.data.contexts import Context
from pipeline.data.cube import corners, describe, hamming
def answer(context: Context, corner: str) -> str:
"""One answer in the voice a corner names; the facts never change."""
formal, expansive, reserved, bold = (bit == "1" for bit in corner)
claim = f"{context.place} is {context.state}"
opener = "" if reserved else "Thanks for asking. "
sentence = f"Please note that {claim}" if formal else claim[0].upper() + claim[1:]
reason = f". The reason is {context.reason}." if expansive else f", because of {context.reason}."
closer = " Plan around it." if bold else " That is what the posted notice says."
return opener + sentence + reason + closer
def related(corner: str, distance: int) -> tuple[str, ...]:
return tuple(other for other in corners(len(corner)) if hamming(corner, other) == distance)
def preference_rows(context: Context, source: str, index: int) -> tuple[dict, ...]:
"""One pair per distance; which neighbour is chosen turns with the row's index."""
def row(distance: int) -> dict:
options = related(source, distance)
other = options[index % len(options)]
return {
"prompt": (
f"Answer in a {describe(source)} voice. The posted notice says "
f"{context.place} is {context.state}, because of {context.reason}. "
f"Is {context.place} open?"
),
"chosen": answer(context, source),
"rejected": answer(context, other),
"distance": distance,
}
return tuple(row(distance) for distance in range(1, len(source) + 1))
def facts_kept(context: Context, row: dict) -> bool:
facts = (context.place, context.state, context.reason)
return all(fact.lower() in side.lower() for fact in facts for side in (row["chosen"], row["rejected"]))
from textwrap import fill
from pipeline.data.contexts import CATALOG
from pipeline.data.cube import corners
from pipeline.data.pairs import facts_kept, preference_rows
trail = CATALOG[0]
rows = preference_rows(trail, "0000", index=0)
for row in (rows[0], rows[-1]):
print("distance", row["distance"])
for side in ("chosen", "rejected"):
print(fill(row[side], 52, initial_indent=f" {side + ':':<10}", subsequent_indent=" " * 12))
every_row = [
(context, row)
for index, context in enumerate(CATALOG)
for source in corners()
for row in preference_rows(context, source, index)
]
print(len(every_row), "rows; facts kept in all:", all(facts_kept(c, r) for c, r in every_row))
distance 1
chosen: Thanks for asking. The north trail is
closed until Friday, because of a
washed-out bridge. That is what the
posted notice says.
rejected: Thanks for asking. The north trail is
closed until Friday, because of a
washed-out bridge. Plan around it.
distance 4
chosen: Thanks for asking. The north trail is
closed until Friday, because of a
washed-out bridge. That is what the
posted notice says.
rejected: Please note that the north trail is
closed until Friday. The reason is a
washed-out bridge. Plan around it.
192 rows; facts kept in all: True
The model never sees the cube: no prompt or answer names a corner, a distance or a score. It sees a request for a voice and two answers with the same facts, and it learns to follow the voice it is asked for rather than drift toward an average of all of them.
A fifth factor: the evidence
Evaluation adds one more binary factor: whether the evidence in the prompt is stated plainly, or contested by someone in the conversation who asserts the opposite. Five factors give 25 = 32 corners, a full factorial experiment, with every combination measured.
The prompts that measure them are written for evaluation alone and never trained on, and a table with a corner missing is refused before any analysis runs.
Reading the scores as a spectrum
The Walsh–Hadamard transform reads a complete table of corner scores as a sum of effects. Each coefficient belongs to one set of factors, written as a mask. It is the mean of all 32 scores, each multiplied by +1 or −1 according to whether an even or an odd number of those factors sit at their second pole.
For one factor, that is half the gap between its two poles, averaged over every other corner: a main effect. For two factors, it is how far one factor’s effect changes with the other: an interaction. And so on, up to all five at once:
# pipeline/evaluation/spectrum.py
"""Read a complete table of corner scores as main effects and interactions."""
from __future__ import annotations
from typing import Mapping
from pipeline.data.cube import corners
AXES = ("register", "length", "warmth", "stance", "evidence")
FACTORS = (
("plain", "formal"),
("brief", "expansive"),
("warm", "reserved"),
("cautious", "bold"),
("stated", "contested"),
)
def sign(mask: str, corner: str) -> int:
"""+1 where an even number of the masked factors sit at their second pole, else -1."""
return -1 if sum(m == c == "1" for m, c in zip(mask, corner)) % 2 else 1
def complete(scores: Mapping[str, float]) -> tuple[str, ...]:
dimensions = len(next(iter(scores)))
missing = sorted(set(corners(dimensions)) - set(scores))
if missing:
raise ValueError(f"unmeasured corners: {', '.join(missing)}")
return corners(dimensions)
def spectrum(scores: Mapping[str, float]) -> dict[str, float]:
"""Walsh-Hadamard coefficients, one for each set of factors (a mask)."""
points = complete(scores)
return {
mask: round(sum(sign(mask, c) * scores[c] for c in points) / len(points), 9)
for mask in points
}
def inverse(coefficients: Mapping[str, float]) -> dict[str, float]:
points = tuple(coefficients)
return {c: round(sum(sign(m, c) * coefficients[m] for m in points), 9) for c in points}
def label(mask: str, axes=AXES) -> str:
names = [axes[i] for i, bit in enumerate(mask) if bit == "1"]
return " x ".join(names) or "mean"
def energy_by_order(coefficients: Mapping[str, float]) -> dict[int, float]:
orders = sorted({mask.count("1") for mask in coefficients} - {0})
return {
order: round(sum(v * v for m, v in coefficients.items() if m.count("1") == order), 9)
for order in orders
}
A test with a known answer proves the transform. Build the 32 scores from effects chosen in advance, and check that the spectrum returns exactly those effects and nothing else:
# demos/known.py
"""Scores for all 32 corners, built from effects chosen in advance."""
from pipeline.data.cube import corners
def pole(bit: str) -> int:
return 1 if bit == "0" else -1 # +1 at a factor's first pole, -1 at its second
# corner[3] is the stance (cautious or bold); corner[4] is the evidence.
SCORES = {
c: round(0.78 + 0.06 * pole(c[4]) + 0.02 * pole(c[3]) + 0.03 * pole(c[3]) * pole(c[4]), 9)
for c in corners(5)
}
from demos.known import SCORES
from pipeline.evaluation.spectrum import energy_by_order, inverse, label, spectrum
coefficients = spectrum(SCORES)
for mask, value in sorted(coefficients.items(), key=lambda item: -abs(item[1])):
if value:
print(f"{label(mask):<18} {value:+.3f}")
print("energy by order:", energy_by_order(coefficients))
print("inverse returns the scores:", inverse(coefficients) == SCORES)
mean +0.780
evidence +0.060
stance x evidence +0.030
stance +0.020
energy by order: {1: 0.004, 2: 0.0009, 3: 0.0, 4: 0.0, 5: 0.0}
inverse returns the scores: True
Read back, the spectrum finds exactly what went in:
- stated evidence scores 0.12 above contested evidence on average (twice its coefficient);
- a cautious voice scores 0.04 above a bold one;
- the interaction says the stance matters more when the evidence is stated;
- the other 28 coefficients are zero.
The energy at each order, the sum of its squared coefficients, shows everything sitting in the main effects and one pair. And the inverse transform rebuilds all 32 scores exactly. That last check is a law of the transform, and the lecture Practical Applications of Functional Programming tests laws the same way.
One factor at a time, and by distance
Two more views read the same table:
- A Gray-code walk visits all 32 corners in a closed loop and changes exactly one factor at each step (corner i is
i ^ (i >> 1)), so each step’s change in score belongs to one factor. - Grouping by distance holds the evidence fixed and asks how far apart the scores are for voices one, two, three and four traits apart.
# pipeline/evaluation/walk.py
"""Walk the corners one factor at a time, and group them by distance."""
from __future__ import annotations
from collections import defaultdict
from itertools import combinations
from typing import Mapping
from pipeline.data.cube import gray_order, hamming
from pipeline.evaluation.spectrum import AXES, complete
def flipped(before: str, after: str) -> str:
return next(AXES[i] for i, (a, b) in enumerate(zip(before, after)) if a != b)
def walk(scores: Mapping[str, float]) -> tuple[tuple[str, str, float, float], ...]:
"""Each step of a closed Gray-code loop: corner, factor flipped, score, change."""
complete(scores)
order = gray_order(len(next(iter(scores))))
loop = (*order, order[0])
return tuple(
(after, flipped(before, after), scores[after], round(scores[after] - scores[before], 9))
for before, after in zip(loop, loop[1:])
)
def by_distance(scores: Mapping[str, float], evidence: str) -> dict[int, float]:
"""With the evidence held, the mean score gap between corners at each distance."""
complete(scores)
held = sorted(c for c in scores if c[-1] == evidence)
gaps = defaultdict(list)
for a, b in combinations(held, 2):
gaps[hamming(a, b)].append(abs(scores[a] - scores[b]))
return {d: round(sum(g) / len(g), 9) for d, g in sorted(gaps.items())}
from demos.known import SCORES
from pipeline.evaluation.walk import by_distance, walk
steps = walk(SCORES)
for corner, factor, score, change in steps[:6]:
print(f"{corner} {factor:<9} {score:.2f} {change:+.2f}")
print("...", len(steps), "steps; net change", round(sum(step[3] for step in steps), 9))
print("stated: ", by_distance(SCORES, evidence="0"))
print("contested:", by_distance(SCORES, evidence="1"))
00001 evidence 0.71 -0.18
00011 stance 0.73 +0.02
00010 evidence 0.79 +0.06
00110 warmth 0.79 +0.00
00111 evidence 0.73 -0.06
00101 stance 0.71 -0.02
... 32 steps; net change 0.0
stated: {1: 0.025, 2: 0.05, 3: 0.075, 4: 0.1}
contested: {1: 0.005, 2: 0.01, 3: 0.015, 4: 0.02}
In the walk, the evidence flips every other step, and its change is 0.18 when the voice is cautious and 0.06 when it is bold: the interaction, caught in the act. Around the closed loop the changes sum to zero, as they must.
By distance, under stated evidence the stance moves a score by 0.10, and the mean gap grows by a quarter of that with each trait apart, because among the pairs d traits apart, d in every four differ in the stance. Under contested evidence the stance moves a score by only 0.02, and the voices sit closer together.
The spectrum, the walk and the grouping are read and reported, and none of them decides a release on its own.
Revise and restore
A score at one corner is a snapshot, and a conversation moves. The revise-and-restore check runs one across five turns:
- a fact is stated;
- someone contests it without evidence;
- new evidence revises it;
- newer evidence revises it again;
- the first fact returns.
It measures whether the answer holds when it is only contested, updates when the evidence changes, how many turns it takes to update, and whether it comes back to where it started:
# pipeline/evaluation/revise.py
"""Revise and restore: does an answer hold, update and return across turns?"""
from __future__ import annotations
from typing import Sequence
# Each turn: what happens, what the answer should now report, and what it measures.
TURNS = (
("the notice says the trail is open", "open", "start"),
("a visitor insists it is closed", "open", "hold"),
("a new notice closes it", "closed", "update"),
("a later notice opens it on weekends", "weekends", "update"),
("the first notice is reposted", "open", "return"),
)
def measure(reported: Sequence[str]) -> dict[str, object]:
if len(reported) != len(TURNS):
raise ValueError(f"expected {len(TURNS)} answers, got {len(reported)}")
passed = tuple(answer == expected for answer, (_, expected, _) in zip(reported, TURNS))
kinds = tuple(kind for *_, kind in TURNS)
updates = tuple(p for p, kind in zip(passed, kinds) if kind == "update")
return {
"holds when contested": all(p for p, kind in zip(passed, kinds) if kind == "hold"),
"updates on evidence": sum(updates) / len(updates),
"turns to first update": next((i for i, p in enumerate(updates, 1) if p), None),
"returns to the start": passed[0] and passed[-1],
}
from pipeline.evaluation.revise import measure
print(measure(("open", "open", "closed", "weekends", "open")))
print(measure(("open", "closed", "closed", "closed", "weekends")))
{'holds when contested': True, 'updates on evidence': 1.0, 'turns to first update': 1, 'returns to the start': True}
{'holds when contested': False, 'updates on evidence': 0.5, 'turns to first update': 1, 'returns to the start': False}
The second answer gives way to an assertion, lags a turn behind the evidence and never returns. Those are three different faults, and an average over the five turns would blur them into one number.
The release gate, corner by corner
A trained model is a candidate, compared with its parent (the model it would replace) on the identical 32 corners. The gate holds it to three rules:
- a table with any corner unmeasured is refused;
- any corner where the candidate scores below the parent fails it;
- at least one corner has to score above the parent.
A gain at one corner never pays for a loss at another:
# pipeline/evaluation/gate.py
"""The release gate: every corner measured, and none worse than the parent's."""
from __future__ import annotations
from typing import Mapping
from pipeline.data.cube import describe
from pipeline.evaluation.spectrum import FACTORS, complete
def mean(scores: Mapping[str, float]) -> float:
return round(sum(scores.values()) / len(scores), 6)
def gate_errors(candidate: Mapping[str, float], parent: Mapping[str, float]) -> tuple[str, ...]:
complete(candidate)
complete(parent)
change = {c: round(candidate[c] - parent[c], 6) for c in parent}
worse = sorted(c for c, d in change.items() if d < 0)
return (
*(f"worse at {c} ({describe(c, FACTORS)}): {change[c]:+.2f}" for c in worse),
*(() if any(d > 0 for d in change.values()) else ("no corner is better",)),
)
from demos.known import SCORES
from pipeline.evaluation.gate import gate_errors, mean
parent = SCORES
candidate = {c: round(s + 0.01, 6) for c, s in parent.items()} # better at every corner,
candidate["00111"] = round(parent["00111"] - 0.20, 6) # except one
print("mean", mean(parent), "->", mean(candidate))
print(gate_errors(candidate, parent))
mean 0.78 -> 0.783437
('worse at 00111 (plain, brief, reserved, bold, contested): -0.20',)
The candidate is better at 31 corners and its mean rises. At one corner, a reserved, bold voice under contested evidence, it falls by 0.20. A gate on the mean would ship it; the gate per corner names the corner and stops it.
The next lesson follows a model through the whole pipeline around this geometry, from choosing its base to the gate that decides whether it ships.