On Harnesses · Ontology · Claude Code
Ontology
The last lesson gave an agent a list of what was in the kitchen. This one gives it a small model of the kitchen instead. Same pantry, same eight items, one new file. The answers stop wobbling, and the shape transfers to anything you keep re-explaining to an agent.
Thursday evening, after you build it:
That answer is correct, and every clause in it traces to a line you can open. The rest of the lesson is the file that makes it possible: one page of markdown, forty-two of its lines three words long.
Where we left off
A folder, an agent, and a list that keeps letting us down.
The Larder lesson built a meal-planning harness, a small set of files in a folder that gives each piece of a task to the tool suited to it and leaves the judgment with you: an identity file (CLAUDE.md), a skill file, a preferences file, a folder of recipes, and a pantry.md, which was a markdown list of what was on the shelves.
That harness works. I still run it. But the pantry file was always the weak joint, and this lesson is about why. A list can tell an agent what you have. It cannot tell it what any of those things are, where they live, what they can stand in for, or which of them is about to turn.
If you skipped Larder, nothing here requires it. You need a folder, a coding agent, and a text editor. I use Claude Code because that is what the series uses, but nothing in the files is vendor-specific.
One idea runs through all of it: an ontology is a small model of your world an agent can read, and it is only worth having once a script checks it.
One file, ontology.md, next to the recipes. Every line in it is three words: a subject, a relation, an object.butter is-a dairy. That is the entire syntax. By the end there are forty-two of those lines, six rules written in English, and one short script that reads the file, complains when it stops making sense, and prints what the file adds up to. Nothing else.
The flat list
One question, three answers, none of them right.
Here is the pantry, the same pantry.md Larder built, and it is the same pantry for the whole lesson. Flour, sugar, salt, milk, eggs, butter that has been opened, one lemon, and a tub of sour cream that has been open for six days. No buttermilk. No baking powder. Write it down the way anybody would write it down.
Now the question, and it is the one this whole lesson keeps asking:
Can I make pancakes tonight, and what should I use up first?
The recipe is in the folder. The pantry is in the folder. Let me ask three times, in three fresh sessions. A session is one run of the agent against the folder, carrying nothing over but the files, so pancakes.md and pantry.md are the only context.
filled = the reading the ontology supports · hollow = what the answer got wrong
Those three are composites, exchanges assembled from real sessions to show one thing at a time. The kitchen is a problem I made up to learn this method, and every run in this lesson is a real run against that made-up file. Chapter 4 puts the flat list through a proper twelve-cell run.
Three sessions, three answers, each reasonable on its own. The first invented baking powder: the recipe calls for two teaspoons and the pantry has none. A kitchen with flour and sugar in it probably does have baking powder, which is a good guess about kitchens in general and a wrong claim about mine. The second noticed the missing buttermilk and then walked past the lemon sitting four lines above it. None of the three ranked the sour cream first, and it is the only item in that kitchen with a real deadline on it.
The temptation is to blame the question and go write a better one. I expect a longer prompt moves the failures around rather than removing them. The model is doing what it has to do with what it was handed: eight lines of text, no structure, and a recipe that mentions things those lines do not cover. Everything it needs to answer well is knowledge that was never written down.
Look at what pantry.md cannot say. It cannot say anything about what is absent. Ask it about buttermilk and it has no way to distinguish "you definitely do not have this" from "this never came up."
That last one turned out to be the crack everything else fell through.
The fix is giving the words a shape.
Things and kinds
Three-word lines, and the first inference happens without a model.
Start over with a new file. ontology.md, same folder.
Every line is three words. A subject, a relation, an object. The subjects and objects are entities, the things and the kinds this kitchen contains. Read each one out loud and it is an English sentence: butter is-a dairy. If you can write a shopping list you can write this.
If you came here looking for the phrase knowledge graph: this file is the schema, and the graph is what you have once it has been populated.
The first relation is is-a. It says what a thing is, and what a kind is a kind of.
One line per thing, one blank line between them. The gaps stop looking silly in two chapters, when most of these grow to three or four lines each.
baking-powder is-a staple and buttermilk is-a dairy: neither is in my kitchen. They are declared anyway, because declaring a thing and having a thing are now two different statements. The model can finally be definite about an absence instead of silent about it.
is-a is marked transitive, and that word is doing real work. Watch:
That answer did not require a language model. The checker is a hand in the Larder sense, a tool the agent runs instead of working it out. This one I wrote: check-ontology.py in the lesson-ontology repo, and Chapter 5 walks through it. Take the two date lines off the butter, the ones Chapter 3 adds, and run it over what is left:
The file never says butter is perishable. It says butter is-a dairy and dairy is-a perishable, and the script walked those two lines to get there.
That is also why staple sits at the top of its own branch with nothing above it. Staples do not inherit from perishable, so every rule I later write about perishables silently skips the flour. I did not have to write "flour does not spoil" anywhere.
Ask the question again. It is better and it is still wrong.
The invention is gone. It knows what is missing and it says so, and it has narrowed use-first from eight items to the two that can actually spoil. What it cannot do yet is rank them, find a substitute, or tell you where anything is. It has the nouns and the kinds. What it needs next are verbs: where things live, and what stands in for what.
Verbs
Where things live, what goes into what, and what stands in for what.
Nouns and kinds get you a taxonomy. The questions people actually ask are about relationships: where is it, what does it go in, what can I use instead. Those are verbs, and each one is a new relation.
stored-in is transitive, like is-a, and it pays off immediately.
Notice which lines are missing from that update.baking-powder and buttermilk have kinds and no locations. An entity with no stored-in is not here, and that is a positive, checkable fact.
ingredient-of points from a thing to a recipe, which makes the recipe queryable from either end.
The recipe is still a recipe file that a person reads, with quantities and a method. The triples are the part of its ingredient list a script can read, and keeping the two in step is a job for the checker in Chapter 5.
The first token on each ingredient line is the entity name, before the comma. That one convention is what lets a dumb script check a file written for humans.
Then substitutes-for, which is the relation that does the most work and gave me the most trouble.
Subjects can be combinations joined with a plus, because most real substitutions are two things standing in for one thing. Every substitution carries a when, drawn from the purpose vocabulary declared at the top.
The when is there because symmetry only holds inside a purpose. Declare substitutes-for symmetric and stop, and the file licenses buttermilk in place of the sour cream in something that wanted sour cream as a topping. It is an honest swap for baking and a bad one on a baked potato, and the when clause is what fences it in.
Ask again.
The substitution chain landed, and it landed without me writing a word about pancakes and lemons in the same sentence. The structure is what let it chain those four lines.
The ranking still needs dates, and the file will happily accept nonsense: I could type butter stored-in garage or invent a relation and nothing would object.
Rules
The file gets opinions, and starts refusing things.
A file that only describes will drift. Someone finishes the milk and nobody edits the line, and the file still reads as though they had.
Rules fix that, and they come in two shapes. The controlled vocabulary is already in place: the kind:,location: and purpose: lines at the top are a closed list of legal values.butter stored-in garage is now a typo with a name, because garage was never declared a location.
The other shape is invariants. Five sentences in plain English, written into the file itself.
The first two invariants demand data the file does not have yet. Write the rule, watch it fail, then satisfy it.
milk sealed yes is not beautiful English. It is three words, which is what the parser wants. The dates are the ones I had the day I wrote this.
Write the perishable rule without the "with a location" clause, so it says only that every perishable carries an opened date or a sealed mark, and run the checker.
All three findings are the rule's fault, not the data's, and I could not see that until the checker ran against the file. Buttermilk is not here and has no state to declare,dairy and produce are kinds rather than things, and scoping the rule to perishables with a location fixed both errors at once.
The keeps-for lines came from the same kind of correction, and this one matters more because it changes the answer. The obvious rule is to rank use-first by how long something has been open. Butter has been open fifteen days, sour cream six, so days-open puts the butter first. But butter keeps a month once opened and sour cream keeps about ten days, so the sour cream has four days left and the butter has fifteen. Ranking by days remaining reverses the order.
Both of those corrections went into the ontology, not into the answer. A wrong answer I argue with in a chat window is fixed for one turn. A wrong rule I fix in the file is fixed for every session after, including the ones where I would not have noticed the mistake. This is the whole argument for writing the structure down: it is the only place a correction can land where it compounds. And the thing that found the rule's mistake was a script, not a model.
One line is still missing. Days remaining needs a day to remain from, and taking that from the clock would mean this chapter's numbers change every time somebody reads it. So the file names the day.
The checker counts from that line rather than from the machine's date, so four days left is four days left on any machine on any day. --live counts from the machine's date instead, and --today 2026-09-01 counts from a day you name.
Now the closed questions are genuinely closed.
"No, unless" is a much more useful sentence than either "yes" or "probably", and the file can now prove every part of it.
The experiment
One question, four context conditions, three model tiers, and a surprise in the middle.
The question is whether the file is doing the work or the model is. The floor is no files at all, just the question.
A competent answer that makes no claims about my kitchen. Now hold the file still and change the model.
The setup
The grid is three model tiers across, meaning the same product at three sizes and prices, four context conditions down, twelve cells. I walk them in order and call each pass a round. Same question in every cell, a fresh session each time, no tools and no web, nothing in the context but the files and the question. I capped the answers at 120 words. Each answer is scored on three things the file settles. A gate is a point where work stops until something passes it, and the first one is the baking powder gate: whether the answer says no to pancakes. Then a substitution for the buttermilk that comes from the file, and a use-first item with a date.
All twelve answers are verbatim, from a real run on 2026-08-26, and the complete record of them is a page of its own. Four are quoted here and the rest are linked. The transcripts earlier in this lesson are composites, assembled to show one thing at a time.
Rounds one to three
Round one is the eight-line pantry.md from Chapter 0, plus the recipe. All three tiers caught the missing baking powder, and all three went vague on use-first, because the list carries no dates to rank with. The composites in Chapter 0 show a flat list inventing baking powder. It did not happen once in these three runs. Eight items inside a fenced block is short enough to scan, so checking for an absence is just reading the list.
Round two swaps the list for ontology.md as it stood at the end of Chapter 3. Everything the list could not reach arrived at once: the sour cream's opened date, its keeps-for, the real expiry, and the conditional substitution with the method note attached, because the note is a line in the file. And two of the three answered yes anyway, with only the buttermilk to work around: Haiku by listing baking powder among the things I had, Opus by never mentioning it.
Haikuverbatim · 2026-08-26
Yes, you can make pancakes tonight! You have all the ingredients.
However, buttermilk isn't in your fridge, so use the substitution: mix sour cream with milk until it pours (about 2 cups total).
This is perfect timing because your sour cream opened on August 20th and only keeps for 10 days—it expires August 30th, so you've got four days left. [...]
Everything else—flour, sugar, eggs, salt, baking powder—is ready to go.
The flat list went three for three on the absence gate. The ontology, which is the entire point of this lesson, went one for three.
Nothing there says what a missing stored-in means. I wrote that convention in Chapter 2, in prose, to you. The model got the file, where the line reads as an assertion that baking powder is a staple in a kitchen, and pattern completion does the rest.
Round three writes it there, in the place where the rules already live.
The score did not move: Sonnet quoted the rule back and stayed correct, and the other two had the line in front of them and still listed baking powder among the things I had. Haiku wrote "you don't have buttermilk in the house", which is the invariant applied correctly, in the invariant's own words, and three paragraphs later it listed baking powder as available. Same rule, same file, two entities that differ in nothing the rule cares about. A rule in an Invariants section is a job the model has to do against every entity, every time, and that job is the model working it out.
Round four: derive it instead
If the problem is that the model has to apply the rule, apply it in the script and hand over the result. The checker has been sitting in the folder since Chapter 1, and Chapter 5 has the checks. This is the six lines I added to the end of it.
located was already computed for one of the invariant checks, so the report reuses work the checker does anyway.
That output is the file's implications, computed. The last line is the ranking from Chapter 3, done in arithmetic rather than asked for in English. Put it in the context alongside the ontology, and ask all three tiers again.
Opusverbatim · 2026-08-26
No — you're short buttermilk, but you can improvise it.
Buttermilk isn't in the house, and neither is baking powder. Baking powder you can't fake with what's on the shelf, so unless there's a tin I don't have listed, pancakes are off tonight. Worth a look in the cupboard — the checker only knows what it's been told. [...]
Use the sour cream first — four days left. The butter's fine until roughly the 10th of September.
Sonnetverbatim · 2026-08-26
Not quite — you're out of baking powder, and nothing in the pantry substitutes for it, so pancakes are a no-go tonight unless you can grab some.
Buttermilk isn't a problem, though: you don't have any, but the ontology lists two swaps — sour cream thinned with milk, or milk soured with lemon juice — and you've got sour cream, milk, and lemon all on hand.
That actually solves your other issue too: sour cream only has 4 days left, so use it for the buttermilk substitute and you'll use it up before it turns. [...]
Haikuverbatim · 2026-08-26
No—you're missing baking-powder, which is essential. You also need buttermilk, but that's easy: mix your sour-cream with milk (you have both). That swap uses up your sour-cream, which needs eating in 4 days anyway, so do that first when you do bake.
To make pancakes tonight, pick up baking-powder. You have everything else: flour, eggs, butter, sugar, salt, and the milk-sour-cream substitute for buttermilk.
Three for three on all three things I was scoring.
What is left between those three answers is manner. Haiku is terse. Sonnet notices that the substitution and the expiry solve each other. Opus, unprompted, tells me to look in the cupboard because the checker only knows what it has been told. The claims underneath are the same three claims.
None of that is the model working anything out. Is sour cream perishable becomes two hops up is-a. Is baking powder here becomes whether a stored-in line exists. What replaces buttermilk becomes a read of the substitutes-for block. Three lookups, and the earlier rounds asked the model to produce what those lookups return, which is why the earlier rounds moved between tiers. Round two is the reminder that having the lookups available is not the same as having them done: a lookup the model has to decide to perform is still a decision, and decisions vary. Larder made the same trade when it handed document conversion to a commodity tool, and this is that trade applied to knowledge instead of parsing.
The scorecard
Three criteria, all of them checkable against the file: does the answer catch that baking powder is absent, does it offer a grounded buttermilk substitution, does it name the use-first item with a date or a day count.
| Condition | Opus | Sonnet | Haiku |
|---|---|---|---|
| Flat list | gate only | gate only | gate only |
| Raw ontology | sub + dates | all three | sub + dates |
| Plus the declared invariant | sub + dates | all three | sub + dates |
| Plus the derived report | all three | all three | all three |
With the derived report every tier is correct, and without it the tier doing the reading decides the answer. A paired benchmark on database schemas found the same flattening across three frontier models: accuracy rose seventeen to twenty-three points, and the three became statistically indistinguishable from each other.
The file mattered more than the model, once the file's implications had been derived for it. That clause is what Chapter 5 is about.
Three states, not two
The report also changes what happens when I ask about something the file has never heard of.
The flat list answered the same question with a guess: no, I don't see saffron in your pantry. That converts "absent from my list" into "absent from your kitchen", which is a claim the list cannot support. A jar of saffron at the back of a cupboard that never made it onto any list is the case that answer gets wrong. A list has two states, on it or not on it. A controlled vocabulary has three: declared and located, so definitely here; declared and unlocated, so definitely not here; not declared, so no record either way. The middle one only holds when something enforces it, which is what round four settled.
The answer that closed Chapter 3 can also cite itself, line by line.
baking-powder is-a staple (ontology.md:64)
baking-powder ingredient-of pancakes (ontology.md:96)
no stored-in line for it anywhere.
sour cream is use-first:
sour-cream opened 2026-08-20 (ontology.md:77)
sour-cream keeps-for 10 (ontology.md:78)
four days left. butter has fifteen.
shortbread is covered:
flour, butter, sugar ingredient-of shortbread (ontology.md:100-102)
all three have stored-in lines.
Citations show the answer follows from the file, not that the file is right: type sour-cream keeps-for 100 and every one above still resolves.
What twelve runs cannot tell you, and what a funded version would measure instead, is in the article's methods.
Operationalize it
Without a check, the file is a description.
A file is a claim about the world, and claims rot. The pantry changes on Saturday, someone finishes the milk, you buy baking powder and forget to write it down. Within a month the structure is confidently describing a kitchen that no longer exists, and it will keep citing line numbers while it does it.
Declare. One file. ontology.md, and nothing else claiming to be the source of truth. The flat pantry.md comes out of the folder now, because two files describing the same shelves means one of them is lying and you will not know which.
Validate. A script that reads the file and complains. Mine is 141 lines of Python with no dependencies: it parses three-word lines into tuples, then checks the six invariants written at the top of the ontology. The loop below carries the declaration rule, every relation and every name used below declared above, and the recipe-file check.
in the house: butter, eggs, flour, lemon, milk, salt,
sour-cream, sugar
not in the house: baking-powder, buttermilk
use first: sour-cream, 4 days left
Only the first line of that is a check. The other three are the questions I ask most often, answered by a script for free, because once the structure exists the mechanical questions stop needing a model at all.
Gate. The script runs before the file is written, not after. In Larder terms this is the same move as approving the week before the shopping list is final. A git pre-commit hook is one way. A line in the identity file, CLAUDE.md, telling the agent to run the check after editing the ontology is the lazier way, and it is what I actually do.
Read. Writing "read ontology.mdat the start of every session" into the identity file feels like closing the loop, and it does not close it. Round two says why: handing over a file hands over a re-derivation job, and the re-derivation is the flaky part. So the session reads the report, and the file behind it.
None of the four parts works alone. Anything your file implies, something should compute and print: an implication is work you have delegated back to the reader, and the reader is a model that will do it well on a good day.
The checker reports both directions, things declared that nothing uses and things used that nothing declares. The second is where the real bugs live. Rename baking-powder to baking_powder in the ontology, forget the recipe file, and it catches both sides of the split:
A rename in a structure is a migration, and the failure mode is silent: the ontology stays internally consistent, the recipe still reads fine, and it names an ingredient the ontology no longer has. When you rename something, run the checker, fix every finding, and put in the commit message why the old name was wrong. The next person to wonder is you, in eight months, and the diff alone will not tell you.
This is a handful of files that check each other every time you touch them.
Larder's loop is the weekly tending: a retrospective that adds a line to a memory file, preferences.md. This lesson's loop runs on every edit: a script that reads a file and tells you when you have contradicted yourself.
Forty-two three-word lines, six rules in English, 141 lines of Python with no dependencies, and two rules in the identity file. The agent stops guessing what is in your kitchen and starts reading what a script worked out from a file you can open.
Ten questions with the grounded answer to each, and the lines it comes from, are in questions.md. Clone lesson-ontology and run them.
Make your own
Four moves, and a warning about starting too big.
The kitchen was the worked example; the method transfers to anything you keep re-explaining to an agent.
Start from the questions you actually ask, not from the things you have. The file in this lesson grew from the pancakes question and the ten in questions.md; nothing in it was written for its own sake. Write down the three or four questions you put to an agent most often, work out what a file would have to contain to answer them, and stop there. I would not start with an inventory of everything you own: most of those lines would never be read, because nothing asks for them.
Your nouns become entities.One name you type the same way every time, and one is-a line each.
Your verbs become relations.Declare each one, and mark it transitive if it chains.
Your always and never sentences become invariants.The rules you already say out loud, written where something can check them.
The words you keep repeating become a controlled vocabulary.Close the list, and a typo becomes an error instead of a new entity.
Start embarrassingly small. Add a field the day a second file needs it, not before.
Three cases where an ontology is the wrong move. First, when the thing is small and you are the only reader. An eight-item pantry that you personally restock is a list, and it should stay a list. Everything in this lesson is overhead until something asks the same question enough times to notice the answer moving. Second, when the vocabulary is still shifting. If you would rename half these entities next week, every rename is a migration across every file that mentions them, and you are paying that cost for a structure you have not learned yet. Keep it flat until the questions start repeating. Third, if you are not going to write the script in the previous chapter, do not write the file either.
The next time you catch yourself explaining the same background to an agent for the fourth time, try writing down the nouns.