Talk one of three · October 2026

Hands
and feet

What an agent is, and why git matters

AI in Systems Engineering · week 7 · Andrew Dunn

Then week 8, The ladder and the meter: meta models, graphs, retrieval, and the evaluation loop. Week 9, Structure over the trove: meta models, SysML v2, and one question asked four ways.

Scroll to run it Read on: the plates are printed at their finished state

Plate I

Correct the answer,
or correct the file

Most people patch the answer in front of them, and that correction dies when the session does.

1. Correct the answer, or correct the file

One

A shopping list comes back wrong, because the file that lists what is in the kitchen is out of date.

The list leaves off broccoli. pantry.md still says a bag is in the freezer, and the bag was eaten a week ago.

Two

Two fixes are available, and on the day they are the same size.

Correct the answer: type the missing item into the reply. Correct the file: open pantry.md and delete one line. Write down which one you would do. The weeks run forward when you scroll on, or when you move the slider.

Three

Six weeks later, only one of the two transcripts has stopped seeing the mistake.

A session is one conversation, from its first message to its last. A correction typed into a reply lives as long as that session. A correction written into the file is read again next Sunday, and every Sunday after it.

Fig. 1 A drawing of the same Sunday question six weeks running. Left, the reply was corrected; right, the file was.

Plate II

What an
agent is

An agent is a language model given tools, running a loop: read, decide, act, observe, repeat.

2. What an agent is

One

A language model, the program behind a chat window, reads what is in front of it and writes text back.

Ask a chat window to delete the broccoli line and you get a sentence saying it has been deleted. Nothing on disk has changed.

Two

An agent is that same model placed in a loop and handed tools.

A tool is a named action the model may ask for, carried out by ordinary software: read a file, write a file, run one named command. The model asks for one call at a time, the tool runs, and the result comes back as the next thing the model reads. One pass round the loop is a turn, and the figure counts them. Read, decide, act, observe, and round again until it needs a person.

Three

One real tool carries today's example.

pandoc is a converter that turns one document format into another, and it is older than any of this. It turns a recipe web page into a plain file once: curl -sL URL | pandoc -f html -t markdown > recipes/SLUG.md. Every week after that, the agent reads the file instead of the page. A page can change or go away; the file is the same every Sunday, and reading it needs no network.

Fig. 2 The loop, with the detour it takes at act. Beside it, the same model with no tools. The rack is drawn narrow: the software around the model, which plate IV names the harness, hands an agent exactly these three tools, and a call outside them is refused.

Plate III

What the
session forgets

Everything the agent knows fits one window, that window gets worse as it fills, and at the end of the session it goes to zero.

3. What the session forgets

One

A context window is everything the model can see at once.

Your words, the files it read, the results of the commands it ran, all of it counted in tokens, the pieces of text a model reads and is billed by. The window has a fixed size, a million tokens on some models today, and inside one session the room it takes is not given back.

Two

It gets worse before it is full.

Dex Horthy's rule of thumb, from his own use and said in a talk, not measured in a study: around the 40 percent line, depending on the task, the answers start to fall off, and he calls the far side the dumb zone. On a model with a million-token window he stops at around three or four hundred thousand, in an interview.

Past that mark, by his account, the answers get worse. The bar draws his rule: rough ink from the mark on.

Three

Then the session ends, and all of it is gone.

The next session opens on an empty window. The only part that survives is what was written to a file, and the folder beside the bar is exactly as full as it was before.

Fig. 3 One session's window, filling. The folder beside it is on disk, and it does not empty.

Plate IV

The folder
that loads first

The fix for a model that forgets is not a bigger window, it is a few plain files the agent reads before anything else.

4. The folder that loads first

One

A harness is everything around the model that decides what it knows and what it may do: the files it reads first, the tools on its rack, the gates where it stops for a person. Talk three measures one part by part.

This one is five things, four markdown files and a folder of three recipe cards, and one command that starts a Sunday. Markdown is plain text with a few marks in it for headings and lists. There is no application here, and nothing in it is code.

Two

Five things, in the order the lesson builds them.

CLAUDE.md says who the agent is and what it must never do. The plan-the-week skill is the workflow written as a document the agent follows, one file, SKILL.md. pantry.md is what is on the shelves. preferences.md is memory, one line per correction. recipes/ holds the cards, one file per dinner.

Three

One line of the real preferences file reads: cauliflower in any form, the kids refuse.

Somebody wrote that down once, and in this household no plan since has proposed it. Typing /plan-the-week starts the workflow, and all five are read before anything is proposed.

Fig. 4 The five things, and where they land. Every name on the plate is a real file in the folder.

Plate V

Every change
carries a name

Once the fix is to edit a plain file, you need to see what changed and be able to take it back. git already does that.

5. Every change carries a name

One

The folder behind today's example is a git repository.

git is a tool that records a folder's history, and a folder it records is a repository. A commit is one saved change to that folder, with a name, a date, the person who made it, and a short id to point at it. Every correction to a file arrives as one.

Two

In this household, preferences.md grows about a line a week.

Weeknight dinners should use at most one pan and one oven rack. If a recipe needs a marinade, save it for the weekend. Sumac, when a card offers sumac or za'atar. Each of these arrived on a Sunday, the frozen pizza line too, and the rail beside the card draws six Sundays.

Three

A correction can itself be wrong, so there has to be a way back.

The last entry on the rail takes one line back out of the file. None of this was built for AI. git is the ordinary tool underneath, and it turns out to be what a folder you correct by hand needed.

Fig. 5 Every line on the card is a line of the real file. The rail is a drawing of six Sundays, and so is the reversal at its end; the marinade line is in the real file today.

Plate VI

The two
gates

The loop only stays honest because it stops at fixed points for a person to decide.

6. The two gates

One

The workflow runs in four phases: sort what we have, propose the week, build the shopping list, hold a retrospective: what landed, what did not.

It is the skill from plate IV, one markdown file, and the agent works through it the way a person works through a checklist.

Two

Two of those phases stop and wait for a person.

The shopping list is not written until the week is approved. Nothing is written to memory until the diff, the list of what would change, is approved. The marker on the lane waits at each gate until you open it. A gate here is a stop for a person; talk three's gate is a model's own yes or no, and the word is the same because both decide whether the loop goes on.

Three

What you do at a gate is the part a model cannot do for you.

Approve, and the plan stands as proposed. Tweak, and the agent revises the week and waits again. At the retrospective, what landed and what did not come back as lines; at the second gate you approve that diff, and only then is a line written into preferences.md, where next Sunday will read it. On the plate the approve path writes nothing because nothing was corrected: the line that reaches memory is the tweak, read back. In this household it is about five minutes a week.

Fig. 6 Four phases, two gates. A line reaches preferences.md only past the second gate: one on the tweak path, none on the approve path.