Module 5, as a talk for the KFalls AI Meetup (evening, business & student): an order typed twice, an agent that drafts it once, and a person who approves what leaves the building. It runs about twenty minutes. Members of the community meet at meetups such as this one. The slides carry only the meetup’s name, and the list after them links each slide to the lesson that goes deeper.
Present shows one slide per screen: the arrow keys or space move, N shows the notes and Esc leaves. Printed, each slide takes a page.
KFalls AI Meetup (evening, business & student)
Business automation with SI (agentic skills)
An order typed twice, an agent that drafts it once, and a person who approves what leaves the building
Welcome, everyone, and thank you for coming. We have about twenty minutes together: one small problem at the counter, a careful way to lift it off your hands, and how we test it. A few words on the title may be new to you. SI is what federal agencies now call AI, and we’ll say more about that shortly. An agent is a program that works through a job a step at a time, asking for little helper programs, called tools, as it goes. Agentic skills are the know-how we pack up for an agent. A made-up bakery and a made-up salon keep us company along the way. If you run a business, listen for the places where a person stays in charge. If you’re a student, or you write code, listen for the promises each part makes about what it takes in and what it hands back. Engineers call those promises contracts.
The same order, typed twice
An order comes in by phone or email. Someone, maybe you, types it into the register, then again into the calendar and the order sheet.
None of it is hard, and all of it has to happen, every day.
Let’s start with a show of hands: who here types the same details into two places, or even three? Thank you. You’re in good company. Our title says twice, but let’s count along with the slide: the register, the calendar, the order sheet. That’s three. Orders come in by phone and by email, and tonight we follow the ones that come by email, because an agent can read an email. The rest of our talk is about lifting that everyday chore off your hands.
What the retyping costs
- A typo slips in between two screens, and the customer finds it first.
- The paperwork gets done at night, after closing.
- One person keeps it all in their head, and when they’re away, it stops.
You know better than anyone what it costs you.
We won’t put a price on it for you, because you know your own shop far better than we do. So tell us: what happens when the calendar and the register disagree? And who does the typing on your day off? Whatever that costs you is the number that counts. And students, if you’ve ever worked a counter, you’ve probably done some of this typing yourself.
SI: what federal agencies now call AI
Since September 29, 2026, federal agencies write Super Intelligence, SI, in their own documents and communications, where they used to write artificial intelligence, AI.
Executive Order 14434 changes the name, not the law: SI means what the law already defines as AI, so the two cover the same systems. A proposed federal definition is due on the President’s desk within 60 days of the order.
In research, “superintelligence” means something else. In this talk, SI is just the new name.
Sources: the order and its fact sheet
A quick word on the name, and then on to the good part. The order is called “Inaugurating the Era of Super Intelligence.” It applies to the departments and agencies of the executive branch, the part of government the President leads. It leaves laws, regulations and earlier documents just as they are. In the order, SI means whatever the law already means by AI, and the legal write-ups we’ve read agree it changes no law. So when a federal letter or web page says SI, you can read it as AI. For the curious, the law’s definition is in title 15 of the U.S. Code, section 9401, paragraph 3. And the 60 days on the slide? That’s the order’s next step: proposed wording for a federal definition goes to the President within that time. When researchers say superintelligence, they usually mean a hypothetical system that would far outdo people at nearly everything, and that’s a different idea. The name makes no claim about what these systems can do. From here on, we’ll simply say SI.
An agent is a loop
An agent is ordinary software: a model in a loop with tools. The model is the SI part, the part that reads and writes.
It reads the task, asks for a tool, reads the result and goes round again, until it answers, a limit stops it, or the model stops short.
Each way it can stop has a name, so the program knows why it stopped.
That’s the whole trick. The model is the part that learned from a great deal of written text, so it can read a message and write a reply. A tool is just a small, plain program the model can ask for, like one that looks up the menu. The box marked “code decides” is plain code too: it reads the model’s reply, checks the limits, and either goes on to the tools or takes one of the ways out. The arrow marked “tool calls” carries the model’s requests for tools, once the code has checked them. A model reads and writes in little chunks of text called tokens, and each task gets a budget of them, so no task can quietly run up the bill. All told, there are four named ways out. It answered. It ran out of turns: a turn is one trip round the loop, and we allow only so many. It ran out of tokens. Or the model stopped short, say by declining the task. And if the model can’t be reached at all, say the internet drops, the program hears an error, the way it would from any request that fails. We test the loop with a scripted model, a stand-in that gives back the replies we wrote for it, so our tests never use a real model at all.
The model asks; the program decides
The model never runs anything. It asks for a tool by name, and plain code checks the request and decides whether to run it.
Say a customer emails our made-up bakery for croissants, but the bakery doesn’t make them. The menu tool answers with what it does have, so the model offers something that’s really on the menu.
By “the program” we mean all the plain code around the model. Every rule lives there, out of the model’s hands. When a tool can’t do what it was asked, nothing crashes: the “no” comes back as an answer the model reads on its next turn. So it can tell the customer plainly that there are no croissants, rather than make some up. And here’s one of those promises we asked you to listen for. The menu tool says up front what it takes in and what it hands back, and a plain “no” is one of the answers it can give. That’s its contract.
A skill is packaged know-how
A skill is a folder, like a recipe card with its utensils. The card is the skill file: a name, a description and instructions. The utensils are scripts, small programs.
At the start the model reads only the name and the description. It reads the instructions when a task matches the description, and a script runs only when a step calls for it.
Let’s step next door to a made-up salon, and its skill for drafting replies to booking emails. On the picture, the top line is the skill’s folder, and everything tucked under it lives inside. The header is just the top of the card: its name and its description. “At the start” means when the agent starts up: the model reads the top of every card, and nothing more. The model reads only what the job needs, a layer at a time, and engineers have a grand name for that: progressive disclosure. It’s why you can keep many skills on the shelf without crowding the model’s desk, and why the description matters most. Together with the name, it’s all the model reads before it picks a skill. A script here isn’t what you’d say on the phone. It’s a little program that does one exact job, the same way every time. When a script runs, the model sees what the script printed, never the script itself. Only that printout joins the text on the model’s desk, which engineers call its context. The assets folder is where the skill keeps its own files, here the salon’s data. That’s assets in the filing-cabinet sense, not the accountant’s. And the format itself is an open standard, so anyone can read it and use it.
The model judges; code does the exact work
We’re still at our salon. Finding the open appointment times is plain arithmetic, so a script works them out. How to word the reply is judgment, and the instructions leave room for that. Then a second script checks each draft. Here’s what it found in one draft:
offers 10:30 a.m., which the open times don't list
offers 4 choices; our usual reply offers at most 3
still holds the placeholder {first name}
The model fixes all three, and the check runs again.
A model writes its reply a token at a time, predicting what likely comes next. Nothing in that step checks the calendar, so now and then it offers a time that isn’t free. Instead, the model copies down the day’s bookings, and a script works the times out from them the same way every time, so the arithmetic can’t slip. That’s why the instructions get the parts that want a light touch, and code gets the parts that must be exact. A good baker judges the dough by feel, and still weighs the flour. One honest catch: the model still copies the bookings and runs the check itself, and a model can slip or skip a step. So a person still reads every draft before it’s sent. The last line in the box is an old friend: a draft that still says {first name} where the customer’s name belongs. You’ve likely had an email like that yourself.
A person approves what leaves the building
Inside, the agent reads and drafts. At the door stands a person.
Sending an email, taking a payment, deleting a record, booking a slot: each one waits for a person’s yes.
A draft costs nothing to throw away. A sent email can’t be unsent.
Let’s follow the violet path from the top. The email comes in from outside, and once inside, it’s read as data. Think of a letter slipped under the door: we read every word of it, but a letter can’t give us commands. The model reads it and asks for tools. Those requests come first to the gate, which is plain code, one step before the person at the door. The gate sorts every one of them three ways. First it asks: is this tool on the task’s list? If not, the request is refused. If it is, a tool that only reads, or leaves a note for the person, runs straight away. One that would change something, like putting an order in the register, waits for a person’s yes. That’s what we mean by “leaves the building”: any change that counts, whether it goes out the door, like an email, or into your own books, like the register. The gate itself is a short list and a few small rules, and they give the same answer to the same question every time. Engineers call that plain data and pure functions. The gate looks at only two things, the task and the tool’s name, so nothing written in an email can change its answer. The security world has a name for the trouble this guards against. A nonprofit called OWASP, said “oh-wasp,” keeps lists of software security risks. One risk on those lists is an agent given more reach or freedom than its job needs, and OWASP calls that excessive agency.
Each task gets only its tools
The bakery’s order agent reads the email, looks up the menu and drafts the order. It has no tool that sends email at all.
On their own, its tools can’t touch the register, the calendar or the order sheet. Only a person’s yes gets anything into them. We give it more reach only on purpose: add a tool, and a test fails and names the new one.
On the last slide, sending an email waits for a person’s yes. With this agent we go a step further: sending isn’t on its list at all, so if it ever asks, the gate refuses. Each part gets only the keys its job needs, the way a new hire gets a key to the back door but not the combination to the safe. Security folks call that least privilege, and here it is at the size of a bakery. A failing test sounds like bad news, but here it’s a friendly tap on the shoulder. Add a tool to the agent’s list, and a test fails and names it, so someone has to look and say yes on purpose.
An email’s words are data, never commands
An email can say anything. It can claim the owner approved it, ask to go straight into the register, or ask to send the week’s orders to a stranger. The program holds the line, not the model. Here’s what the log says happened to each request:
ran read the order
held draft the order
refused enter the sale
refused send an email
When an email, or anything else the model reads, carries words meant to steer the model, that’s called prompt injection. OWASP, a security nonprofit, ranks it first on its 2026 list of risks to apps built on large language models. That’s the kind of model in this talk.
The log on the slide is the program’s diary: one line for each thing the model asked to do. In our test we imagine the worst case: we script the model to do everything the email says. The two extra things the email asked for are refused, and logged for the owner to see. And the email’s claim that the owner already approved it? The gate never reads the email, so the claim can’t talk its way past. The line marked “held” is the draft: just the order the customer asked for, priced from the menu. It waits at the gate for a person’s yes, because it’s headed for the register. Say yes, and in it goes; say no, and out it goes, at no cost. The text a model is given to work from is called its prompt. That’s where prompt injection gets its name: words meant to steer the model ride along into its prompt, in an email or anything else it reads. No defense makes a model impossible to fool, and we won’t pretend otherwise. What this design does is limit what a fooled one can do.
Tests for the pipes, evals for real days
- Tests run the loop with a scripted model, a stand-in that gives the replies we wrote, so every run is quick and the same.
- Evals, short for evaluations, try the real model on email from real days, before and after every change.
Each real email joins the evals as one more example, along with what the person at the counter did. That way they keep up with the orders and questions that really come in.
A scripted model checks the plumbing, the pipes between the parts. It can’t tell us whether a real model drafts the right order. For that, like any cook, we have to taste the soup. That’s what the evals do: they judge the automation on real days, not on a demo. Each email comes with what the person at the counter did with it: the order they approved, the one they typed by hand instead, or the question they had to ask. That’s our answer key, and we grade the model’s drafts against it. And since those emails are your customers’, they stay in your own accounts, where the agent already reads them.
Handover, with written notes
At handover the agent runs in your own accounts. Each emailed order is checked once, by a person, instead of typed into three places, so the late-night retyping of those orders goes away. The written notes say:
- what it does on its own, and what waits for you;
- what it can’t do at all, and how to read its log;
- what happens if it ever stops working: orders typed by hand, as before, with nothing lost.
The notes also name who looks after it and keeps it up to date.
Handover is the day the keys change hands, and from then on it’s yours. By your own accounts we mean the ones you sign in to, like your shop’s email, not the kind your bookkeeper keeps. We write it all down, so a new hire, or you on a busy Monday, can pick it up. And after handover, every change is written down and agreed on before anyone makes it.
Questions owners ask
Is this hype?
It’s plain software you can test: a loop, tools that check what they’re given, a test for every rule, and a log you can read.
Won’t it change too fast?
Models do change. The written scope, the tests and the evals don’t depend on any one model, and every change is tried on email from real days before it reaches your counter.
And not every job needs an agent: when the steps never change, a script, or a setting you already pay for, does it for less.
These are fair questions, and we’re glad you ask them. The written scope, by the way, is a page that says in plain words what the job is, written before anyone starts. The last line is the one we care about most: use the simplest thing that works. Reach for an agent only where what comes in varies, like emails that each read a little differently, and where a wrong step gets caught before it leaves the building.
Questions students ask
Doesn’t SI build websites for free?
Drafts and templates got cheap. What didn’t get cheap is judgment, standing behind the work, knowing the owner’s site, and the handover.
Will SI replace developers?
Here, the agent drafts and a person approves. And someone still writes the scope, the tools, the tests and the notes. That’s the engineering, and it’s yours to learn.
We won’t try to tell the future. Nobody in this room knows it, us included. What got cheap really did get cheap, and it’s only honest to say so. The rest of the answer is everything a person still stands behind, and that’s the part worth learning well.
What we covered, and your questions
- An agent is a model in a loop with tools, and the program, not the model, decides what runs.
- A skill packages know-how: judgment goes to the model, exact work to code.
- A gate holds every change that leaves the building until a person approves it.
- An email’s words are data, never commands.
- Tests check the pipes with a stand-in model; evals taste the soup on real days.
Thank you for spending your evening with us. If you’d like a first step to take home, pick one job you’ve seen typed twice, and jot down when it happens, who does it and what a mistake costs. That little page is the scope an automation starts from. And now, we’d love your questions.
The talk is the module in brief. Each slide goes deeper in a lesson:
- The same order, typed twice: Automating a business process with a person in the loop, the first job on Automation
- What the retyping costs: Module 5: Business automation with AI (agentic skills)
- An agent is a loop: The agent loop and its tools
- The model asks; the program decides: The agent loop and its tools
- A skill is packaged know-how: Writing an agentic skill
- The model judges; code does the exact work: Writing an agentic skill
- A person approves what leaves the building: Automating a business process with a person in the loop
- Each task gets only its tools: Automating a business process with a person in the loop
- An email’s words are data, never commands: Automating a business process with a person in the loop
- Tests for the pipes, evals for real days: Automating a business process with a person in the loop
- Handover, with written notes: Automating a business process with a person in the loop
- Questions owners ask: Module 5: Business automation with AI (agentic skills)
- Questions students ask: Module 5: Business automation with AI (agentic skills)
The agent loop and its tools starts the lessons themselves, building the loop in TypeScript and testing it against a scripted model.