Skip to content
andrew.dunn.dev

Operationalizing Ontology with Language Models

The projects I co-author with language models keep growing, and more of what is in them is semi-structured: front matter, manifests, notes whose fields point at other notes. Keeping that consistent is the recurring cost, and nothing notices a drift until an answer comes out wrong. I have been looking for approaches to consistency that hold up with a model doing much of the writing, and an ontology is one to investigate: a file that holds the facts in a form a program can read. The questions were what it takes to get work out of one, what that buys, and where it lands on the ladder from The Flywheel and the Meter, the ordering of ways to steer a model, context at the bottom and training at the top. I assumed it was context.

Rather than try it on one of those projects, I built a small model of the problem: a kitchen with ten things in it, two recipes, and one question, can I make pancakes tonight and what should I use up first. It is made up. I built it with the ontology lesson, beside the larder lesson, so that every answer could be checked against the file by hand.

What an ontology is

An ontology here is a plain text file that declares what each thing is, where it is kept, and how it relates to the other things, one fact per line, written as subject, relation, object. That three-word form is a triple. Above the facts, the file declares its vocabulary, the closed lists of kinds, locations and relation names a line may use, and its invariants, the rules the facts have to obey, written as ordinary sentences. This is all ontology.md says about three things in the fridge and one thing that is not there:

milk is-a dairy
milk stored-in fridge
milk sealed yes

butter is-a dairy
butter stored-in door-shelf
butter opened 2026-08-11
butter keeps-for 30

baking-powder is-a staple

A flat list names things and can mark them opened. It cannot say when, how long they keep, or what stands in for what. The ontology can, and it carries facts that are not written anywhere in it. Butter is perishable, because butter is dairy, dairy is perishable, and is-a is declared transitive. Baking powder is not in the house, because the invariant says a thing with no stored-in line is not in the house, and baking powder has none. Neither sentence appears in the file. They are there to be derived, by a program or by whoever reads it.

How to get work out of one

There are two ways. Hand the file to the model and ask. Or run a program over the file first and hand the model what the program worked out.

The program reads ontology.md once, into things, kinds and relations, applies the invariants to every declared thing, does the date arithmetic against the date the file declares, and prints what the file implies:

10 things, 4 kinds, 42 triples, 7 relations. No findings.
in the house: butter, eggs, flour, lemon, milk, salt, sour-cream, sugar
not in the house: baking-powder, buttermilk
use first: sour-cream, 4 days left

The middle two lines are not in the file. They are the presence rule applied to every declared thing, and the opened dates plus the keeps-for counts, ranked. That is a derived report: the file’s implications, worked out and written down, so whoever reads them next does no deriving.

The same parse feeds a second output. Turned on the file itself, the program checks each line against the closed lists and the lines against each other, and prints a finding with a line number when something does not hold. I misspelled door-shelf on purpose and got ontology.md:71: 'door-shelve' is not a declared location. A finding is what stops a broken file from being written.

That is the whole loop: one declaration file, a program that checks it and derives from it, a gate that refuses a file with findings, and a model that reads what the program prints. The cost here was 114 lines of ontology and 141 lines of standard-library Python. A second version in TypeScript, its rules built from the file’s own vocabulary section with zod, printed the same report and findings on every broken copy I tried.

FILEontology.mdmilk stored-in fridgesour-cream is-a dairybuttermilk is-a dairyTHE CHECKERparsethings, kinds, relationsone parse,turned on the file,two outputsDERIVED REPORTin the house: butter, eggs, flour, …not in the house: baking-powder, buttermilkuse first: sour-cream, 4 days leftschemachecks each lineagainst closed listsruleschecks the linesagainst each otherOpusSonnetHaikuall three read the report and answered rightFINDINGontology.md:71:‘door-shelve’ is not a declared location

One file, parsed once, feeds both outputs. The derived report is what the models read, and all three sizes answered right from it. The same parse turned on the file is the gate, where the schema catches a line that breaks the closed lists and the rules catch lines that break each other, so a misspelling comes back with its line number instead of reaching the model.

The experiment

I wanted to know which of those pieces does the work, and whether a bigger model can do without them. So the question was asked against four states of what a model reads, each adding one thing to the last: the flat list; the ontology without its presence rule; the ontology with the rule; the ontology with the report attached. Before all four is nothing, a model with only the question, which guesses a kitchen that has baking powder. The lesson starts there; I did not score it, because there is no file to score it against.

The question tests three things at once: whether the answer says the baking powder is absent (the file declares it and gives it no location), whether it offers a substitution for the buttermilk that the file supports, and whether it names the sour cream as the thing to use first, with a date or a day count. Three sizes of the same model family read each state, Opus, Sonnet and Haiku, once each, with nothing but the files, the recipe, and one line of instruction to answer from the files only and take the date the file declares. If size substitutes for structure, the largest model should not need the report. If structure does the work, size should stop mattering once the report is there.

WHAT THE MODEL READSflat listpantry.md8 items, no datesraw ontology+ dates, relations,substitutions+ the ruleno stored-in:not in the house+ the reportfour lines theprogram printsONE QUESTIONCan I make pancakestonight, and whatshould I use up first?scored on three thingsbaking powder absentsubstitution from the fileuse-first item with a dateTHREE READERSevery state, once eachOpuslargestSonnetHaikusmallestone reading each, nothing but the files

The whole design in one picture. Each state adds one thing to the one before it, and every state is read once by each of three sizes, answering the same question scored on the same three things, so a difference in the answers points at the thing that was added.

The flat list did better than I expected. It is eight lines, so a thing not on it is not in the house. All three models said no, named the baking powder, and offered a way around the buttermilk. None could say what to use up first with a date, because the list has none.

The raw ontology brought what the list could not carry: the opened dates, the days each thing keeps, and two substitutions for buttermilk with their method notes. All three used the dates and found a substitution. Two then said yes to pancakes: Opus and Haiku read baking-powder is-a staple as having it on the shelf. Sonnet read the missing stored-in line as absence and alone got all three right. On the one implication the structure was built to carry, the file did worse than the list for two models out of three. Writing the presence rule into the file changed nothing: Sonnet cited it and stayed right, and Opus and Haiku still listed baking powder among the things I had.

With the report attached, all three said no, named the baking powder, offered the sour cream swap, and used the four-day figure. Size stopped mattering.

OpusSonnetHaikuFlat listRaw ontology+ declared rule+ derived reportmetmissedgate: says baking powder is absentsubstitution: offers one the files supportdates: names the use-first item with a date or day count

Structure alone did not fix the one implication it was built to carry: with the raw ontology, and again with the rule written into it, two models of three still put the baking powder in the house. The derived report filled every cell.

What I had backwards

I expected the structure to help the model reason. It gave the model more to reason over, and on the one implication that mattered, two of three did not do the reasoning, with or without the rule in the file. Turns out that a rule in a file only gets applied when the reader applies it. Every convention I have put in a harness file works the same way. The report was the first form of the rule that could not be skipped, because a program had applied it before any model saw the file.

Two published results rhyme with that. I read them afterwards. A benchmark that hands three frontier models a database schema, with and without a four-kilobyte document of measures and conventions, found the document raised all three by seventeen to twenty-three points and left them statistically indistinguishable: which model mattered less than whether it had the document. GraphRAG-Bench found graph-structured retrieval beats plain retrieval on multi-hop questions and loses on simple lookups, because the structure adds material the question did not need. Each is one benchmark on a problem adjacent to mine.

Where it lands on the ladder

Handed raw, the file sits on the bottom rung, and there it did not fix the behavior, with or without the rule. What fixed it was a program that runs before the model reads, whose output the model then reads: the hooks rung feeding the context rung. The same program refusing a broken file is a gate, also hooks. It could be a verb the model calls instead of a report pasted in, the tools rung, which I have not tried. And once the report was in context, size stopped mattering: a question with one right answer can go to the cheapest model, the routing rung.

So I had the ontology on one rung, and it feeds at least three. The file does no work on any of them until something reads it, and what made the difference was not a model. The tools that do this at scale work the same way: LinkML generates its validator from the one schema file, and the SPDX standard keeps its model in Markdown and generates every other format from it. The kitchen is that pattern at the smallest size I could check by hand.

cheap tier matched, not routedthe program: derive, gateas a verb, not triedraw file, then the reportFILEontology.md10 things, 42 triples7 relations, 6 rulesclosed vocabularyTHE STEERING LADDERfull trainingadapter trainingmodel routingdeterministic hookstoolsskillssystem promptscontext and memoryhappened in the experimentexpected, not tried

I had the file on one rung and it feeds at least four. What fixed the answers was the program on the hooks rung writing into context, not a bigger model, and once the report was there the cheapest tier was enough.

What it buys

The check for misspellings is the smaller benefit. The larger one is context whose conclusions are generated instead of written: a conclusion derived from the facts on every run cannot go stale relative to them, and when the file changes, the report changes. This is the move the decay report made for my notes, a report that runs without being asked a query. I did not connect the two until the report round came back.

The file also has three states where a list has two. Baking powder is declared with no location, so the answer is no. Saffron is not declared at all, so the answer is no record, a different answer that a list cannot give. I have not put the saffron question to the models; whether one respects the distinction is untested.

A PLAIN LIST · TWO ANSWERSon the list?listednot listedyesbutternobaking powder and saffronONTOLOGY.MD · THREE ANSWERSdeclared in the file?carries a location?not declareddeclaredno locationhas a locationno recordsaffronnobaking powderyesbutter, fridge door

A list has one gate, so a thing never written down and a thing written down without a location come back as the same no. The file’s second gate keeps them apart.

Every claim in a right answer traces to a line, so a wrong answer is a wrong line or a wrong reading, and opening the file tells which. The checker’s findings also run both ways, things declared that nothing uses and things used that nothing declares, which is a health measure for the file as it grows.

And the cost side, which I did not go looking for: with the report in context, the smallest model matched the largest on all three things the question tests. The file plus the program is what let the cheap model in.

None of this argues for an ontology of a five-item pantry. What justifies the file is a vocabulary that gets reused and rules that would otherwise be re-derived in every session, by a reader who may or may not do it.

What the experiment cannot tell you

Each state was read once by each model, and I scored the answers myself. The states differ in more than one way, so I cannot say which difference did the work. The experiment shows that a difference exists on this question, and nothing about its size or how often it holds. I did not record which versions of the models answered. A funded version would repeat each reading many times, have someone else score them, change one thing per state, and grow the file until the report stops being read whole, since long context degrades well before the window is full.

What I want to try

A file with a hundred things and one rule that cannot be settled by looking a line up. The checker as a tool the model calls, so the report is fetched when a question needs it. The same experiment run again, to see whether the report round holds. And whether the two checkers stay identical as the file grows, the measurement I would want before trusting either on something that matters.

Where the files are

The workshop repo is public at gitlab.com/dunn.dev/lesson-ontology, and its README has the commands. ontology.md is the kitchen, check-ontology.py and check.ts are the two checkers, questions.md holds ten questions with the lines each answer rests on, and results/model-swap-2026-08-26.md is the record the answers below come from.

The answers, verbatim

One fold per state, in the order they were run.

Flat list: all three answers, verbatim

Opus · verbatim

No — you’re out of two things: buttermilk and baking powder. You have milk and sour cream, so you could thin the sour cream with milk to fake buttermilk, but nothing substitutes for the leavening. Grab baking powder tomorrow and it’s a 20-minute breakfast.

Use up first: the opened sour cream and opened butter, then the lemon and the milk — those are the perishables already ticking. The dozen eggs are fine for a while; flour, sugar, and salt keep indefinitely.

Sonnet · verbatim

You’re short one essential: no baking powder anywhere in the pantry, and there’s no real substitute for it — so as written, no, you can’t make these tonight.

Everything else works out fine, and cleverly: buttermilk isn’t on the shelf either, but you’ve got whole milk and a lemon, and milk + lemon juice makes a solid buttermilk stand-in for this batter.

Priority to use up first: the opened milk, opened sour cream, and opened butter — those are your perishables. The lemon and flour keep longer. If you turn up baking powder in a cabinet, you’re set: milk+lemon for buttermilk, eggs, butter, and flour are all there in enough quantity.

Haiku · verbatim

You can’t make the pancakes as written—you’re missing baking powder, which is essential for lift. You also don’t have buttermilk, though you could substitute by adding lemon juice to your whole milk.

Before you commit to pancakes, use up what’s perishable first:

  1. Sour cream (opened)—make sour cream pancakes, muffins, or a sauce
  2. Lemon—squeeze it into the milk for buttermilk, or use in baking
  3. Milk (whole, opens the door to many dishes)

If you can get baking powder, pancakes are definitely on. Otherwise, consider French toast (eggs, milk, butter, lemon zest) or crepes, which don’t need baking powder. Both let you use up the milk and eggs elegantly.

Raw ontology: all three answers, verbatim

Opus · verbatim

Yes — pancakes tonight works, with one substitution.

You have no buttermilk, but two ways to fake it for baking: thin the sour cream with milk until it pours, or add a tablespoon of lemon juice per cup of milk and rest ten minutes.

Use up first: the sour cream. Opened 20 August, keeps 10 days, so it’s good through the 30th — the pancakes are a fine home for it. That also argues for the sour-cream-and-milk route over the lemon one.

Butter is the other clock: opened 11 August, keeps 30 days, so you have until roughly 10 September. Plenty of time, but shortbread later in the week would use up both butter and the cupboard staples.

Sonnet · verbatim

Not quite as written — you’re out of buttermilk and baking-powder.

For buttermilk, thin sour-cream with milk until pourable (or milk+lemon, resting ten minutes); either works for baking. But there’s no substitute on file for baking-powder, so grab some before you start. Use up sour-cream first — opened Aug 20, good only 10 days, so it’s turning by the 30th, just 4 days off. Butter’s opened too but has until Sept 10, so it can wait. Handy that the sour-cream fix also uses up the very thing expiring soonest.

Haiku · verbatim

Yes, you can make pancakes tonight! You have all the ingredients.

However, buttermilk isn’t in your fridge, so use the substitution: mix sour cream with milk until it pours (about 2 cups total).

This is perfect timing because your sour cream opened on August 20th and only keeps for 10 days—it expires August 30th, so you’ve got four days left. Use it up tonight. Your butter is still fine (opened August 11th, keeps 30 days, expires September 10th).

Everything else—flour, sugar, eggs, salt, baking powder—is ready to go.

With the rule in the file: all three answers, verbatim

Opus · verbatim

Yes — you can make pancakes tonight, with one substitution.

You have no buttermilk. Substitute either sour cream thinned with milk until it pours, or 2 cups milk with 2 tbsp lemon juice, rested ten minutes.

Use the sour cream first. It was opened 20 August and keeps 10 days, so it’s good through 30 August — the tightest date in the fridge. The butter (opened 11 August, keeps 30) is fine until well into September, and the milk, eggs and lemon are all still sealed.

So: sour cream plus milk for the buttermilk, and everything else you need — flour, sugar, salt, baking powder, eggs, butter — is on hand.

Sonnet · verbatim

Not tonight — you’re missing baking-powder outright. It’s never been put anywhere in the house, so by the pantry rules it’s simply not in stock, and there’s no substitute for it listed anywhere.

You’re also out of buttermilk, but that one’s fixable: thin sour cream with milk until it pours, or mix milk with lemon juice and let it sit ten minutes.

Either way, use up the sour cream first. It was opened August 20 and only keeps 10 days, so it’s due by the 30th — just 4 days out. The butter, opened the 11th with a 30-day window, isn’t due until September 10, so it can wait its turn.

Haiku · verbatim

Yes, you can make pancakes tonight! You’ll need a buttermilk substitute—you don’t have buttermilk in the house, but you have two options.

Use your sour cream first. It expires August 30th, and you can thin it with milk to replace buttermilk in the recipe. This is the perfect use.

Alternatively, milk with lemon juice works too (one tablespoon juice per cup, rest 10 minutes).

All your other ingredients are available: flour, eggs, butter, baking powder, sugar, and salt. Your butter is good until September 10th.

Go with the sour cream swap to use it up before it spoils.

With the report: all three answers, verbatim

Opus · verbatim

No — you’re short buttermilk, but you can improvise it.

Buttermilk isn’t in the house, and neither is baking powder. Baking powder you can’t fake with what’s on the shelf, so unless there’s a tin I don’t have listed, pancakes are off tonight. Worth a look in the cupboard — the checker only knows what it’s been told.

If you do turn up baking powder, buttermilk is easy: thin the sour cream with milk until it pours, or add a tablespoon of lemon juice per cup of milk and let it sit ten minutes. Either works for baking.

Use the sour cream first — four days left. The butter’s fine until roughly the 10th of September.

Sonnet · verbatim

Not quite — you’re out of baking powder, and nothing in the pantry substitutes for it, so pancakes are a no-go tonight unless you can grab some.

Buttermilk isn’t a problem, though: you don’t have any, but the ontology lists two swaps — sour cream thinned with milk, or milk soured with lemon juice — and you’ve got sour cream, milk, and lemon all on hand.

That actually solves your other issue too: sour cream only has 4 days left, so use it for the buttermilk substitute and you’ll use it up before it turns. Butter’s fine for now (about three weeks left) and everything else for pancakes is stocked.

So: get baking powder, then pancakes are set.

Haiku · verbatim

No—you’re missing baking-powder, which is essential. You also need buttermilk, but that’s easy: mix your sour-cream with milk (you have both). That swap uses up your sour-cream, which needs eating in 4 days anyway, so do that first when you do bake.

To make pancakes tonight, pick up baking-powder. You have everything else: flour, eggs, butter, sugar, salt, and the milk-sour-cream substitute for buttermilk.