One
A shopping list comes back wrong, because the file that lists what is in the kitchen is out of date.
The list leaves off broccoli. pantry.md still says a bag is in the freezer, and the bag was eaten a week ago.
Talk one of three · October 2026
What an agent is, and why git matters
AI in Systems Engineering · week 7 · Andrew Dunn
Then week 8, The ladder and the meter: meta models, graphs, retrieval, and the evaluation loop. Week 9, Structure over the trove: meta models, SysML v2, and one question asked four ways.
Scroll to run it Read on: the plates are printed at their finished state
Plate I
Most people patch the answer in front of them, and that correction dies when the session does.
One
A shopping list comes back wrong, because the file that lists what is in the kitchen is out of date.
The list leaves off broccoli. pantry.md still says a bag is in the freezer, and the bag was eaten a week ago.
Two
Two fixes are available, and on the day they are the same size.
Correct the answer: type the missing item into the reply. Correct the file: open pantry.md and delete one line. Write down which one you would do. The weeks run forward when you scroll on, or when you move the slider.
Three
Six weeks later, only one of the two transcripts has stopped seeing the mistake.
A session is one conversation, from its first message to its last. A correction typed into a reply lives as long as that session. A correction written into the file is read again next Sunday, and every Sunday after it.
Plate II
An agent is a language model given tools, running a loop: read, decide, act, observe, repeat.
One
A language model, the program behind a chat window, reads what is in front of it and writes text back.
Ask a chat window to delete the broccoli line and you get a sentence saying it has been deleted. Nothing on disk has changed.
Two
An agent is that same model placed in a loop and handed tools.
A tool is a named action the model may ask for, carried out by ordinary software: read a file, write a file, run one named command. The model asks for one call at a time, the tool runs, and the result comes back as the next thing the model reads. One pass round the loop is a turn, and the figure counts them. Read, decide, act, observe, and round again until it needs a person.
Three
One real tool carries today's example.
pandoc is a converter that turns one document format into another, and it is older than any of this. It turns a recipe web page into a plain file once: curl -sL URL | pandoc -f html -t markdown > recipes/SLUG.md. Every week after that, the agent reads the file instead of the page. A page can change or go away; the file is the same every Sunday, and reading it needs no network.
Plate III
Everything the agent knows fits one window, that window gets worse as it fills, and at the end of the session it goes to zero.
One
A context window is everything the model can see at once.
Your words, the files it read, the results of the commands it ran, all of it counted in tokens, the pieces of text a model reads and is billed by. The window has a fixed size, a million tokens on some models today, and inside one session the room it takes is not given back.
Two
It gets worse before it is full.
Dex Horthy's rule of thumb, from his own use and said in a talk, not measured in a study: around the 40 percent line, depending on the task, the answers start to fall off, and he calls the far side the dumb zone. On a model with a million-token window he stops at around three or four hundred thousand, in an interview.
Past that mark, by his account, the answers get worse. The bar draws his rule: rough ink from the mark on.
Three
Then the session ends, and all of it is gone.
The next session opens on an empty window. The only part that survives is what was written to a file, and the folder beside the bar is exactly as full as it was before.
Plate IV
The fix for a model that forgets is not a bigger window, it is a few plain files the agent reads before anything else.
One
A harness is everything around the model that decides what it knows and what it may do: the files it reads first, the tools on its rack, the gates where it stops for a person. Talk three measures one part by part.
This one is five things, four markdown files and a folder of three recipe cards, and one command that starts a Sunday. Markdown is plain text with a few marks in it for headings and lists. There is no application here, and nothing in it is code.
Two
Five things, in the order the lesson builds them.
CLAUDE.md says who the agent is and what it must never do. The plan-the-week skill is the workflow written as a document the agent follows, one file, SKILL.md. pantry.md is what is on the shelves. preferences.md is memory, one line per correction. recipes/ holds the cards, one file per dinner.
Three
One line of the real preferences file reads: cauliflower in any form, the kids refuse.
Somebody wrote that down once, and in this household no plan since has proposed it. Typing /plan-the-week starts the workflow, and all five are read before anything is proposed.
Plate V
Once the fix is to edit a plain file, you need to see what changed and be able to take it back. git already does that.
One
The folder behind today's example is a git repository.
git is a tool that records a folder's history, and a folder it records is a repository. A commit is one saved change to that folder, with a name, a date, the person who made it, and a short id to point at it. Every correction to a file arrives as one.
Two
In this household, preferences.md grows about a line a week.
Weeknight dinners should use at most one pan and one oven rack. If a recipe needs a marinade, save it for the weekend. Sumac, when a card offers sumac or za'atar. Each of these arrived on a Sunday, the frozen pizza line too, and the rail beside the card draws six Sundays.
Three
A correction can itself be wrong, so there has to be a way back.
The last entry on the rail takes one line back out of the file. None of this was built for AI. git is the ordinary tool underneath, and it turns out to be what a folder you correct by hand needed.
Plate VI
The loop only stays honest because it stops at fixed points for a person to decide.
One
The workflow runs in four phases: sort what we have, propose the week, build the shopping list, hold a retrospective: what landed, what did not.
It is the skill from plate IV, one markdown file, and the agent works through it the way a person works through a checklist.
Two
Two of those phases stop and wait for a person.
The shopping list is not written until the week is approved. Nothing is written to memory until the diff, the list of what would change, is approved. The marker on the lane waits at each gate until you open it. A gate here is a stop for a person; talk three's gate is a model's own yes or no, and the word is the same because both decide whether the loop goes on.
Three
What you do at a gate is the part a model cannot do for you.
Approve, and the plan stands as proposed. Tweak, and the agent revises the week and waits again. At the retrospective, what landed and what did not come back as lines; at the second gate you approve that diff, and only then is a line written into preferences.md, where next Sunday will read it. On the plate the approve path writes nothing because nothing was corrected: the line that reaches memory is the tweak, read back. In this household it is about five minutes a week.
The marker is waiting at the first gate: approve the week as proposed, or tweak it.
preferences.md only past the second gate: one on the tweak path, none on the approve path.