An agent is software: a model in a loop with tools. The model asks for a tool, the application runs it and hands back the result, and the loop goes round until the model answers or stops for a reason the loop names, such as a limit. On the Sean Dinwiddie’s Webmastery team the loop is built like the rest of the practice: a pure core that decides, effects at the edge, and tests named after scenarios.
This lesson builds one for a bakery, an invented one. A customer asks whether two dozen sourdough rolls and six croissants can be picked up on Saturday at 9, and the agent reads the menu and the hours and drafts the answer. Its tools only read, and the draft goes nowhere on its own: Automating a business process with a person in the loop puts a person between a draft and the customer.
The loop, as the API runs it
Anthropic’s documentation calls tool use a contract between the application and the model: the application says which tools exist and what their inputs look like, and the model decides when to call them. The model never runs anything. It asks, and the application does the work (How tool use works).
On the Claude Messages API, one round of the loop goes like this:
- The request lists the tools, each a name, a description and a JSON Schema for its input.
- When the model wants a tool, the response stops with
stop_reason: "tool_use"and carries one or moretool_useblocks, each with anid, the tool’snameand itsinput. - The application runs every call and sends back one user message of
tool_resultblocks, each naming thetool_use_idit answers, withis_errorset when the tool failed. - The loop goes round again while the stop reason is
"tool_use". Any other reason ends this lesson’s loop, among them an answer ("end_turn"), a reply cut off atmax_tokens, and a"refusal".pause_turn, which only server tools return, asks the application to send the reply back and go on.
Anthropic’s Building effective agents (December 2024) describes agents the same way: usually just a model using tools on feedback from its environment, in a loop, with a stopping condition such as a maximum number of iterations to keep it in hand. This lesson’s loop has three ways to stop besides the answer, and each one is a value the caller can switch on.
decide chooses, the tools run and their results go back to the model. Every way the loop stops leads to one value of the type Stop; a request to the model that fails outright rejects runAgent’s promise instead.A model, as the loop sees it
The loop never imports a provider’s SDK. It talks to a Model: one conversation that takes a turn and gives back a reply. A turn is the task, or the results of the last calls; a reply is an answer, some tool calls or a stop, each with the tokens it cost.
// agent/model.ts
// What the loop needs from a model, and nothing about any one provider.
export type ToolCall = { readonly id: string; readonly name: string; readonly input: unknown }
export type ToolResult = { readonly callId: string; readonly content: string; readonly isError: boolean }
export type ModelTurn =
| { readonly kind: 'task'; readonly text: string }
| { readonly kind: 'results'; readonly results: readonly ToolResult[] }
export type ModelReply =
| { readonly kind: 'answer'; readonly text: string; readonly tokens: number }
| { readonly kind: 'toolCalls'; readonly calls: readonly ToolCall[]; readonly tokens: number }
| { readonly kind: 'stopped'; readonly reason: 'refused' | 'truncated' | 'unfinished'; readonly tokens: number }
// One conversation. The first turn carries the task; each later turn carries
// the results of the calls the model asked for last.
export interface Model {
readonly reply: (turn: ModelTurn) => Promise<ModelReply>
}
Three kinds of reply and no more, so the compiler can tell when a switch over them misses one. The interface is a seam: a scripted model fills it in the tests, and Claude fills it in production, through one small adapter at the end of this lesson.
The imports name their .ts files, so they run and check as written with:
- Node.js 22.18 or later on the 22 line, or 24.3 or later, the first releases on their lines to run them directly, with no flag and no warning;
tsc, which checks them with the strict settingstsc --initwrites from TypeScript 5.9 on (checked here with 7.0.2), plusallowImportingTsExtensions,noEmitand"types": ["node"]in place of the empty list it writes, and"type": "module"inpackage.json;- the installs
typescript,@types/nodefor thenode:imports,vitest4 or later, the first to print tests the way this module’s printouts show (checked here with 5.0.3), and@anthropic-ai/sdk0.115 or later, the first to typefallbacks: 'default'(checked here with 0.131.0).
Tools are typed functions with contracts
A tool has three parts. The spec is what the model reads: a name, a description and a JSON Schema for the input. The decoder turns whatever arrived into a typed value or a typed failure, the way Endpoints at the boundary decodes a response before the app trusts it. Then comes the run itself, which only ever sees a value that decoded.
// agent/tool.ts
import type { ToolCall, ToolResult } from './model.ts'
export type Failure = { readonly kind: 'badInput' | 'notFound' | 'unavailable'; readonly message: string }
export type Outcome<T> = { readonly ok: true; readonly value: T } | { readonly ok: false; readonly failure: Failure }
export const ok = <T>(value: T): Outcome<T> => ({ ok: true, value })
export const fail = (kind: Failure['kind'], message: string): Outcome<never> => ({ ok: false, failure: { kind, message } })
// What the model sees: a name, a description and a JSON Schema for the input.
export type ToolSpec = {
readonly name: string
readonly description: string
readonly inputSchema: {
readonly type: 'object'
readonly properties: Readonly<Record<string, object>>
readonly required: readonly string[]
readonly additionalProperties: false
}
}
export type Tool = ToolSpec & { readonly call: (input: unknown) => Promise<Outcome<unknown>> }
// The input is decoded before the tool runs, so run only ever sees a typed value.
export const defineTool = <I, O>(
spec: ToolSpec,
decode: (input: unknown) => Outcome<I>,
run: (input: I) => Promise<Outcome<O>>,
): Tool => ({
...spec,
call: async (input) => {
const decoded = decode(input)
return decoded.ok ? run(decoded.value) : decoded
},
})
// Pure: every outcome, success or failure, becomes a result the model can read.
export const toResult = (call: ToolCall, outcome: Outcome<unknown>): ToolResult =>
outcome.ok
? { callId: call.id, content: JSON.stringify(outcome.value), isError: false }
: { callId: call.id, content: `${outcome.failure.kind}: ${outcome.failure.message}`, isError: true }
A failure is plain data: badInput, notFound or unavailable, with a message the model can act on. toResult turns every outcome into a result, so a failed tool becomes a result with isError set, never an exception. Handle tool calls asks for exactly that: an error result with is_error, and a message that says what went wrong and what to try next.
The lecture Practical Applications of Functional Programming treats expected failure the same way. Here are the bakery’s two tools:
// agent/bakery.ts
// Two tools for a bakery. Both only read: nothing here changes an order or sends a word.
import { defineTool, fail, ok, type Outcome, type Tool } from './tool.ts'
const days = ['monday', 'tuesday', 'wednesday', 'thursday', 'friday', 'saturday', 'sunday'] as const
type Day = (typeof days)[number]
export type Item = { readonly name: string; readonly priceCents: number; readonly noticeDays: number }
export type Bakery = {
readonly hours: Readonly<Record<Day, { readonly opens: string; readonly closes: string } | 'closed'>>
readonly menu: readonly Item[]
}
const field = (input: unknown, key: string): unknown =>
typeof input === 'object' && input !== null ? (input as Record<string, unknown>)[key] : undefined
// Even strict tool use doesn't promise an enum value's capitals, so the day matches in any case.
const decodeDay = (input: unknown): Outcome<Day> => {
const wanted = field(input, 'day')
const day = days.find((d) => typeof wanted === 'string' && d === wanted.toLowerCase())
return day === undefined ? fail('badInput', `day must be one of: ${days.join(', ')}.`) : ok(day)
}
const decodeName = (input: unknown): Outcome<string> => {
const name = field(input, 'name')
return typeof name === 'string' && name.trim() !== ''
? ok(name.trim().toLowerCase())
: fail('badInput', 'name must be the item as the customer wrote it.')
}
export const bakeryTools = (bakery: Bakery): readonly Tool[] => [
defineTool(
{
name: 'opening_hours',
description: "The bakery's opening hours on one day of the week.",
inputSchema: {
type: 'object',
properties: { day: { type: 'string', enum: days } },
required: ['day'],
additionalProperties: false,
},
},
decodeDay,
async (day) => ok({ day, hours: bakery.hours[day] }),
),
defineTool(
{
name: 'menu_item',
description: 'One item on the menu: its price in cents and the days of notice an order of it needs.',
inputSchema: {
type: 'object',
properties: { name: { type: 'string', description: 'The item, as the customer wrote it.' } },
required: ['name'],
additionalProperties: false,
},
},
decodeName,
async (name) => {
const item = bakery.menu.find((i) => i.name === name)
return item === undefined
? fail('notFound', `the menu has no ${name}. It has: ${bakery.menu.map((i) => i.name).join(', ')}.`)
: ok(item)
},
),
]
// An invented bakery for the lesson: its hours, items and prices are made up.
export const exampleBakery: Bakery = {
hours: {
monday: 'closed',
tuesday: { opens: '07:00', closes: '15:00' },
wednesday: { opens: '07:00', closes: '15:00' },
thursday: { opens: '07:00', closes: '15:00' },
friday: { opens: '07:00', closes: '15:00' },
saturday: { opens: '08:00', closes: '13:00' },
sunday: 'closed',
},
menu: [
{ name: 'sourdough rolls', priceCents: 150, noticeDays: 2 },
{ name: 'rye loaf', priceCents: 900, noticeDays: 1 },
{ name: 'cinnamon buns', priceCents: 400, noticeDays: 1 },
],
}
The bakery is passed in, so a test chooses its data, and in production it comes from wherever the bakery keeps its menu. The schema’s enum tells the model which days exist, and decodeDay checks again anyway: a tool doesn’t trust whoever called it. It matches the day in any case, so "Saturday" reads Saturday’s hours, while "Saturday morning" is still bad input.
A failed lookup says what the menu does have, so the model’s next turn can offer something real instead of guessing.
A pure core decides
decide takes the model’s reply, what the run has spent and the limits, and returns a Decision: run these calls, or stop. It awaits nothing and reads no clock, so the same arguments give the same decision every time.
// agent/decide.ts
import type { ModelReply, ToolCall } from './model.ts'
export type Limits = { readonly maxTurns: number; readonly maxTokens: number }
export type Spent = { readonly turns: number; readonly tokens: number }
export type Stop =
| { readonly kind: 'answered'; readonly text: string }
| { readonly kind: 'outOfTurns' }
| { readonly kind: 'outOfTokens' }
| { readonly kind: 'modelStopped'; readonly reason: 'refused' | 'truncated' | 'unfinished' }
export type Decision =
| { readonly kind: 'runTools'; readonly calls: readonly ToolCall[] }
| { readonly kind: 'stop'; readonly stop: Stop }
const stop = (s: Stop): Decision => ({ kind: 'stop', stop: s })
// Pure: from the model's reply and what the run has spent, go round again or stop.
// A limit only keeps the loop from going round again; an answer that arrives is kept.
export const decide = (reply: ModelReply, spent: Spent, limits: Limits): Decision => {
switch (reply.kind) {
case 'answer':
return stop({ kind: 'answered', text: reply.text })
case 'stopped':
return stop({ kind: 'modelStopped', reason: reply.reason })
case 'toolCalls':
if (spent.turns >= limits.maxTurns) return stop({ kind: 'outOfTurns' })
if (spent.tokens >= limits.maxTokens) return stop({ kind: 'outOfTokens' })
return { kind: 'runTools', calls: reply.calls }
}
}
It’s the pure core of FRP fundamentals in software development, and the same split as the thin handlers in The API: Haskell Servant and Nile: pure code decides, and the code around it does the waiting. The lecture Functional Programming in Other Languages names the pattern: a functional core in an imperative shell.
Every stop is a value of Stop: answered, outOfTurns, outOfTokens, or modelStopped with its reason. A caller switches over four cases instead of searching a log for what happened.
The order of the checks is a rule too. An answer is always kept, even one that arrives over budget, since it’s already paid for. A limit only keeps the loop from going round again, so a limit of 3 turns means the model is asked at most three times.
The shell runs the effects
runAgent is the only code that waits on anything. It asks the model, adds up turns and tokens, lets decide choose, runs the calls and goes round with the results. Each turn hands what the run has spent to the next as an argument, so there’s no counter to mutate, and recursion does the work of a while loop.
// agent/loop.ts
import { decide, type Limits, type Spent, type Stop } from './decide.ts'
import type { Model, ModelTurn, ToolCall, ToolResult } from './model.ts'
import { fail, toResult, type Tool } from './tool.ts'
export type Step = { readonly call: ToolCall; readonly result: ToolResult }
export type Run = { readonly stop: Stop; readonly spent: Spent; readonly steps: readonly Step[] }
// The shell: the only code that waits on the model or a tool. decide makes every choice.
export const runAgent = (model: Model, tools: readonly Tool[], task: string, limits: Limits): Promise<Run> => {
const go = async (turn: ModelTurn, before: Spent, steps: readonly Step[]): Promise<Run> => {
const reply = await model.reply(turn)
const spent = { turns: before.turns + 1, tokens: before.tokens + reply.tokens }
const decision = decide(reply, spent, limits)
if (decision.kind === 'stop') return { stop: decision.stop, spent, steps }
const ran = await Promise.all(decision.calls.map((call) => runTool(tools, call)))
return go({ kind: 'results', results: ran.map((step) => step.result) }, spent, [...steps, ...ran])
}
return go({ kind: 'task', text: task }, { turns: 0, tokens: 0 }, [])
}
// An unknown name or a tool that throws is a failure like any other, never a crash.
const runTool = async (tools: readonly Tool[], call: ToolCall): Promise<Step> => {
const tool = tools.find((t) => t.name === call.name)
const outcome =
tool === undefined
? fail('notFound', `no tool is named ${call.name}. The tools are: ${tools.map((t) => t.name).join(', ')}.`)
: await tool.call(call.input).catch((error: unknown) => fail('unavailable', String(error)))
return { call, result: toResult(call, outcome) }
}
The model can ask for several tools in one reply. Lookups that only read are usually safe to run together, so Promise.all runs them at once, and the results go back in one turn, each matched to its call by id. Parallel tool use asks for one result per call, all in the next user message; this loop also keeps them in the order the model asked, so a test can read them.
An unknown tool name and a tool that throws both become failures, so a broken tool costs the model a turn, never the run. A request to the model that fails outright, after the SDK’s own retries, is the one exit that isn’t a Stop: runAgent’s promise rejects, and the caller handles it as it would any failed request.
One run, call by call
Here is the customer’s question run end to end, with a scripted model in the model’s seat. The script asks for three lookups in one turn, then answers; its token counts are made up, like the bakery:
// agent/trace.ts
// One run with a scripted model, printed call by call: what a real model would read back.
import { bakeryTools, exampleBakery } from './bakery.ts'
import { runAgent } from './loop.ts'
import { answers, asks, scriptedModel } from './scripted.ts'
const question = 'Can I pick up two dozen sourdough rolls and six croissants on Saturday at 9?'
const model = scriptedModel([
asks(
300,
{ id: 'c1', name: 'menu_item', input: { name: 'Sourdough rolls' } },
{ id: 'c2', name: 'menu_item', input: { name: 'Croissants' } },
{ id: 'c3', name: 'opening_hours', input: { day: 'saturday' } },
),
answers(
500,
"Two dozen sourdough rolls are $36 and need two days' notice, so please order by Thursday. " +
"We don't bake croissants; the menu has sourdough rolls, rye loaf and cinnamon buns. " +
'On Saturday we open at 8, so 9 works.',
),
])
const run = await runAgent(model, bakeryTools(exampleBakery), question, { maxTurns: 6, maxTokens: 20_000 })
for (const { call, result } of run.steps) {
console.log(`${call.id} ${call.name} ${JSON.stringify(call.input)}`)
console.log(` ${result.isError ? 'error' : 'ok'} ${result.content}`)
}
console.log(`${run.stop.kind} after ${run.spent.turns} turns and ${run.spent.tokens} tokens:`)
if (run.stop.kind === 'answered') console.log(run.stop.text)
$ node agent/trace.ts
c1 menu_item {"name":"Sourdough rolls"}
ok {"name":"sourdough rolls","priceCents":150,"noticeDays":2}
c2 menu_item {"name":"Croissants"}
error notFound: the menu has no croissants. It has: sourdough rolls, rye loaf, cinnamon buns.
c3 opening_hours {"day":"saturday"}
ok {"day":"saturday","hours":{"opens":"08:00","closes":"13:00"}}
answered after 2 turns and 800 tokens:
Two dozen sourdough rolls are $36 and need two days' notice, so please order by Thursday. We don't bake croissants; the menu has sourdough rolls, rye loaf and cinnamon buns. On Saturday we open at 8, so 9 works.
Read c2. The model asked for croissants, the menu has none, and what came back is an error that lists what the menu does have. A real model reads that on its next turn and can say so plainly, as the scripted answer does, instead of inventing a croissant.
Tests with a scripted model
The behavior comes first, as scenarios, in the Gherkin of Given-When-Then (Gherkin) syntax in BDD. Each one is a sentence an owner can read and a rule the loop has to keep:
# agent/pickup.feature
Feature: Pickup questions at the bakery
The agent drafts an answer from the bakery's own hours and menu.
Its tools only read, and the draft goes nowhere until a person sends it.
Scenario: The agent answers from the bakery's own data
Given a customer asks for two dozen sourdough rolls on Saturday at 9
When the model asks for the item and Saturday's hours in one turn
Then both results go back to it together, in the order it asked
And the run stops with its answer
Scenario: A failed lookup goes back to the model as a result
Given a customer asks for croissants, which the bakery doesn't bake
When the model looks them up
Then it reads an error that lists what the menu has
Scenario: Input that breaks the schema never reaches the tool
When the model asks for the hours on "Saturday morning"
Then it reads a badInput error that names the days the tool takes
Scenario: A day with a capital letter still finds the hours
When the model asks for the hours on "Saturday"
Then it reads Saturday's hours
Scenario: An unknown tool is a failure, not a crash
When the model asks for a tool named "place_order"
Then it reads a notFound error that names the tools it has
Scenario: A tool that throws is a failure, not a crash
Given the till that holds the menu is offline
When the model looks up an item
Then it reads an unavailable error that says why
Scenario: The loop stops at its turn limit
Given a model that asks for a tool on every turn
When the agent runs with a limit of 3 turns
Then the model is asked 3 times and no more
And the run stops at the turn limit
Scenario: The loop stops when its budget is spent
Given each reply costs 400 tokens
When the agent runs with a budget of 1000 tokens
Then the run stops over budget after the third reply
Scenario: A refusal is a stop, never an answer
Given the model declines the task
When the agent runs
Then the run stops with the reason "refused"
And no tool runs
A scripted model plays back replies written in advance and keeps every turn the loop sends it, so a test can check what a real model would have read. No test calls a real model, so every run is the same on every machine and costs nothing.
// agent/scripted.ts
import type { Model, ModelReply, ModelTurn, ToolCall } from './model.ts'
// A model that plays back a script, so no test ever calls a real one. It keeps every
// turn the loop sends it, so a test can check what a real model would have read.
export const scriptedModel = (script: readonly ModelReply[]): Model & { readonly seen: readonly ModelTurn[] } => {
const seen: ModelTurn[] = []
return {
seen,
reply: async (turn) => {
seen.push(turn)
const next = script[seen.length - 1]
if (next === undefined) throw new Error(`the script ran out after ${script.length} replies`)
return next
},
}
}
export const asks = (tokens: number, ...calls: readonly ToolCall[]): ModelReply => ({ kind: 'toolCalls', calls, tokens })
export const answers = (tokens: number, text: string): ModelReply => ({ kind: 'answer', text, tokens })
Each test carries its scenario’s name, and its comments follow the scenario’s lines. The last test reads the feature file and fails if a scenario has no test, so the two can’t drift apart:
// agent/loop.test.ts
import { readFileSync } from 'node:fs'
import { describe, expect, it } from 'vitest'
import { bakeryTools, exampleBakery } from './bakery.ts'
import { runAgent } from './loop.ts'
import type { ToolCall, ToolResult } from './model.ts'
import { answers, asks, scriptedModel } from './scripted.ts'
import { defineTool, ok, type Tool } from './tool.ts'
const tools = bakeryTools(exampleBakery)
const limits = { maxTurns: 6, maxTokens: 20_000 }
const noInput = { type: 'object', properties: {}, required: [], additionalProperties: false } as const
// The feature file and the tests name the same scenarios, or a test fails.
const feature = readFileSync(new URL('./pickup.feature', import.meta.url), 'utf8')
const written = [...feature.matchAll(/^\s*Scenario: (.+)$/gm)].map((match) => match[1])
const tested: string[] = []
const scenario = (title: string, test: () => Promise<void>) => {
tested.push(title)
it(`Scenario: ${title}`, test)
}
// One call, then the result the model reads back on its next turn.
const resultOf = async (call: ToolCall, using: readonly Tool[] = tools): Promise<ToolResult | undefined> => {
const model = scriptedModel([asks(100, call), answers(100, 'Thanks for asking.')])
await runAgent(model, using, 'A pickup question', limits)
const next = model.seen[1]
return next?.kind === 'results' ? next.results[0] : undefined
}
const asksEveryTurn = (tokens: number) =>
scriptedModel(
Array.from({ length: 5 }, (_, n) => asks(tokens, { id: `c${n}`, name: 'opening_hours', input: { day: 'friday' } })),
)
describe('Feature: Pickup questions at the bakery', () => {
scenario("The agent answers from the bakery's own data", async () => {
// Given a customer asks for two dozen sourdough rolls on Saturday at 9
const answer = "Yes. Two dozen sourdough rolls are $36, and we open at 8 on Saturday."
const model = scriptedModel([
// When the model asks for the item and Saturday's hours in one turn
asks(
300,
{ id: 'c1', name: 'menu_item', input: { name: 'Sourdough rolls' } },
{ id: 'c2', name: 'opening_hours', input: { day: 'saturday' } },
),
answers(400, answer),
])
const run = await runAgent(model, tools, 'Two dozen sourdough rolls on Saturday at 9?', limits)
// Then both results go back to it together, in the order it asked
expect(model.seen[1]).toEqual({
kind: 'results',
results: [
{ callId: 'c1', isError: false, content: '{"name":"sourdough rolls","priceCents":150,"noticeDays":2}' },
{ callId: 'c2', isError: false, content: '{"day":"saturday","hours":{"opens":"08:00","closes":"13:00"}}' },
],
})
// And the run stops with its answer
expect(run.stop).toEqual({ kind: 'answered', text: answer })
expect(run.spent).toEqual({ turns: 2, tokens: 700 })
})
scenario('A failed lookup goes back to the model as a result', async () => {
// Given a customer asks for croissants, which the bakery doesn't bake
// When the model looks them up
const result = await resultOf({ id: 'c1', name: 'menu_item', input: { name: 'Croissants' } })
// Then it reads an error that lists what the menu has
expect(result).toEqual({
callId: 'c1',
isError: true,
content: 'notFound: the menu has no croissants. It has: sourdough rolls, rye loaf, cinnamon buns.',
})
})
scenario('Input that breaks the schema never reaches the tool', async () => {
// When the model asks for the hours on "Saturday morning"
const result = await resultOf({ id: 'c1', name: 'opening_hours', input: { day: 'Saturday morning' } })
// Then it reads a badInput error that names the days the tool takes
expect(result).toEqual({
callId: 'c1',
isError: true,
content: 'badInput: day must be one of: monday, tuesday, wednesday, thursday, friday, saturday, sunday.',
})
})
scenario('A day with a capital letter still finds the hours', async () => {
// When the model asks for the hours on "Saturday"
const result = await resultOf({ id: 'c1', name: 'opening_hours', input: { day: 'Saturday' } })
// Then it reads Saturday's hours
expect(result).toEqual({
callId: 'c1',
isError: false,
content: '{"day":"saturday","hours":{"opens":"08:00","closes":"13:00"}}',
})
})
scenario('An unknown tool is a failure, not a crash', async () => {
// When the model asks for a tool named "place_order"
const result = await resultOf({ id: 'c1', name: 'place_order', input: { item: 'rye loaf' } })
// Then it reads a notFound error that names the tools it has
expect(result).toEqual({
callId: 'c1',
isError: true,
content: 'notFound: no tool is named place_order. The tools are: opening_hours, menu_item.',
})
})
scenario('A tool that throws is a failure, not a crash', async () => {
// Given the till that holds the menu is offline
const offline = defineTool(
{ name: 'menu_item', description: 'The menu, read from the till.', inputSchema: noInput },
ok, // every input decodes
() => Promise.reject(new Error('the till is offline')),
)
// When the model looks up an item
const result = await resultOf({ id: 'c1', name: 'menu_item', input: {} }, [offline])
// Then it reads an unavailable error that says why
expect(result).toEqual({ callId: 'c1', isError: true, content: 'unavailable: Error: the till is offline' })
})
scenario('The loop stops at its turn limit', async () => {
// Given a model that asks for a tool on every turn
const model = asksEveryTurn(100)
// When the agent runs with a limit of 3 turns
const run = await runAgent(model, tools, 'When are you open?', { ...limits, maxTurns: 3 })
// Then the model is asked 3 times and no more
expect(model.seen).toHaveLength(3)
// And the run stops at the turn limit
expect(run.stop).toEqual({ kind: 'outOfTurns' })
})
scenario('The loop stops when its budget is spent', async () => {
// Given each reply costs 400 tokens
const model = asksEveryTurn(400)
// When the agent runs with a budget of 1000 tokens
const run = await runAgent(model, tools, 'When are you open?', { ...limits, maxTokens: 1000 })
// Then the run stops over budget after the third reply
expect(run.stop).toEqual({ kind: 'outOfTokens' })
expect(run.spent).toEqual({ turns: 3, tokens: 1200 })
})
scenario('A refusal is a stop, never an answer', async () => {
// Given the model declines the task
const model = scriptedModel([{ kind: 'stopped', reason: 'refused', tokens: 50 }])
// When the agent runs
const run = await runAgent(model, tools, 'A pickup question', limits)
// Then the run stops with the reason "refused"
expect(run.stop).toEqual({ kind: 'modelStopped', reason: 'refused' })
// And no tool runs
expect(run.steps).toEqual([])
})
it('has a test for every scenario in pickup.feature', () => expect(tested).toEqual(written))
})
$ npx vitest run --reporter=verbose | grep -E 'โ|Tests'
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: The agent answers from the bakery's own data 3ms
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: A failed lookup goes back to the model as a result 1ms
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: Input that breaks the schema never reaches the tool 0ms
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: A day with a capital letter still finds the hours 0ms
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: An unknown tool is a failure, not a crash 0ms
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: A tool that throws is a failure, not a crash 0ms
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: The loop stops at its turn limit 1ms
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: The loop stops when its budget is spent 1ms
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: A refusal is a stop, never an answer 1ms
โ agent/loop.test.ts > Feature: Pickup questions at the bakery > has a test for every scenario in pickup.feature 0ms
Tests 10 passed (10)
A failing test names the rule that broke. Delete one character in decide, turning >= into >, and the model gets asked a fourth time:
$ sed -i.bak 's/turns >= limits/turns > limits/' agent/decide.ts
$ npx vitest run 2>&1 | grep -E 'FAIL|AssertionError'
FAIL agent/loop.test.ts > Feature: Pickup questions at the bakery > Scenario: The loop stops at its turn limit
AssertionError: expected [ { kind: 'task', โฆ(1) }, โฆ(3) ] to have a length of 3 but got 4
A scripted model tests the loop, never the model’s judgment: its answers were written before the run. Whether a real model asks for the right tools and drafts a good answer is a question for evals on real days, and Automating a business process with a person in the loop builds them for its own task.
The Claude adapter
One file speaks Claude’s API. It turns each ToolSpec into a tool definition, keeps the conversation, and maps each response to a ModelReply:
// agent/claude.ts
// The one file that speaks Claude's API. No test calls it; strict tsc checks it.
import type Anthropic from '@anthropic-ai/sdk'
import type { Model, ModelReply, ModelTurn } from './model.ts'
import type { ToolSpec } from './tool.ts'
export const claudeModel = (client: Anthropic, system: string, specs: readonly ToolSpec[]): Model => {
// The system prompt and the tools stay fixed for the conversation, and the
// messages only grow: each reply's content goes back exactly as it came.
const tools: Anthropic.Beta.BetaTool[] = specs.map((spec) => ({
name: spec.name,
description: spec.description,
input_schema: { ...spec.inputSchema, required: [...spec.inputSchema.required] },
strict: true,
}))
const messages: Anthropic.Beta.BetaMessageParam[] = []
return {
reply: async (turn) => {
messages.push(toMessage(turn))
const response = await client.beta.messages.create({
model: 'claude-opus-5-5',
max_tokens: 16000,
output_config: { effort: 'medium' },
system,
tools,
messages,
// If a safety classifier declines, the API retries on the model it recommends.
betas: ['server-side-fallback-2026-07-01'],
fallbacks: 'default',
})
messages.push({ role: 'assistant', content: response.content })
return toReply(response)
},
}
}
const toMessage = (turn: ModelTurn): Anthropic.Beta.BetaMessageParam =>
turn.kind === 'task'
? { role: 'user', content: turn.text }
: {
role: 'user',
content: turn.results.map((result) => ({
type: 'tool_result' as const,
tool_use_id: result.callId,
content: result.content,
is_error: result.isError,
})),
}
// Every stop reason the SDK knows maps to a reply. When the SDK adds one,
// this switch stops compiling until someone decides what it means.
const toReply = (response: Anthropic.Beta.BetaMessage): ModelReply => {
const tokens = response.usage.input_tokens + response.usage.output_tokens
switch (response.stop_reason) {
case 'end_turn':
case 'stop_sequence':
return { kind: 'answer', text: response.content.flatMap((b) => (b.type === 'text' ? [b.text] : [])).join(''), tokens }
case 'tool_use':
return {
kind: 'toolCalls',
calls: response.content.flatMap((b) => (b.type === 'tool_use' ? [{ id: b.id, name: b.name, input: b.input }] : [])),
tokens,
}
case 'refusal':
return { kind: 'stopped', reason: 'refused', tokens }
case 'max_tokens':
case 'model_context_window_exceeded':
return { kind: 'stopped', reason: 'truncated', tokens }
case 'pause_turn':
case 'compaction':
case null:
return { kind: 'stopped', reason: 'unfinished', tokens }
}
}
A few of its choices carry weight:
- The model is
claude-opus-5-5, Claude Opus 5.5, where Anthropic’s models overview suggests most work start. Its effort is set tomedium, the model’s own default, written out so a change shows in a diff (Effort). - Strict tools. With
strict: true, Claude’s tool inputs match the schema, and every tool it names is one it was offered, because the API constrains sampling to them (Strict tool use), apart from the cases Structured outputs lists: a refusal, a reply cut off atmax_tokens, and the capitalization of anenumvalue. The decoders still run, andrunToolstill answers an unknown name: the scripted model isn’t strict,decodeDaycompares the day without regard to case, as those docs advise, and a tool still doesn’t trust its caller. - The conversation only grows. The system prompt and the tools stay fixed, and each response’s content goes back exactly as it came, thinking blocks included. On Claude Opus 5.5 a thinking block stays valid only while everything sent before it is unchanged (Preserved thinking). The adapter keeps that transcript; the loop keeps only what it needs to decide.
- Refusals. Claude Opus 5.5’s safety classifiers can decline a request, and that arrives as an ordinary response with
stop_reason: "refusal", not an error. The adapter opts into the API’s server-side fallback, a beta: a declined request is retried on the model Anthropic recommends for that kind of refusal, and where there’s none, or that model is rate limited or overloaded, the refusal stands and the loop stops withmodelStopped(Refusals and fallback). - Truncation. A reply cut off at
max_tokensmaps totruncated, so a half-written tool call never runs.
Here is the conversation the adapter would keep if Claude made the same three calls:
tool_result answers the tool_use with its id (the violet lines), all three in one user message, and any thinking blocks go back as they came.toReply covers every stop reason the SDK knows, including pause_turn and compaction, which come only from server tools and the API’s compaction features, and this agent uses neither. The switch has no default, so when the SDK learns a new reason, the build stops until someone decides what it means. Take a case away and see:
$ sed -i.bak "/case 'compaction':/d" agent/claude.ts
$ npx tsc --noEmit
agent/claude.ts(52,57): error TS2366: Function lacks ending return statement and return type does not include 'undefined'.
The production wiring puts the pieces together. The tools the model reads and the tools the loop runs come from one list:
// agent/pickup.ts
// The same loop and tools the tests run, with Claude in the model's seat.
import type Anthropic from '@anthropic-ai/sdk'
import { bakeryTools, type Bakery } from './bakery.ts'
import { claudeModel } from './claude.ts'
import { runAgent, type Run } from './loop.ts'
const system =
"You draft answers to customers' pickup questions for the bakery. " +
'Take hours, items and prices from the tools, never from memory. A person reads every draft before it is sent.'
// A fresh conversation for every question; the limits are this lesson's own.
export const draftPickupAnswer = (client: Anthropic, bakery: Bakery, question: string): Promise<Run> => {
const tools = bakeryTools(bakery)
return runAgent(claudeModel(client, system, tools), tools, question, { maxTurns: 6, maxTokens: 50_000 })
}
The limits here are this lesson’s own; set yours from what real runs cost.
The loop gives the agent its hands. Writing an agentic skill gives it know-how: a skill it loads when a task matches, written from a user story and its scenarios, the same way.