Talk two of three · October 2026

The ladder
and the meter

AI in engineering practice: meta models, graphs, retrieval, and the evaluation loop

This talk lands in the week of 19 October 2026.

AI in Systems Engineering · week 8 · Andrew Dunn

Scroll to run it Read on: the plates are printed at their finished state

Plate I

The model behind
the model

A model pasted into a chat as prose, or the same model as a graph: which one keeps the links the work depends on?

1. The model behind the model

One

An engineering model is parts, requirements, and the links between them.

The rules for what counts as a part, what counts as a requirement, and which links are allowed between them are the meta model, the model behind the model. One rule on this plate: a line may join a part to a part, or a part to the requirement it satisfies. Every modelling tool enforces rules like it, and the work depends on them holding.

Two

Paste that model into a chat window and it arrives as prose.

The parts become nouns. The links become sentences. A reader can follow them. A program cannot, so nothing downstream can check that the valve the controller opens is the valve R2 names. Ask which part must open within two seconds, and what commands it: the answer is the valve and the controller, and the prose puts three sentences between the two facts.

Three

A knowledge graph is the same model written so a program can walk it.

Elements become nodes, links become edges, and each keeps its name and its kind: R2 is satisfied by the valve, so a program can ask which requirements nothing satisfies. Nothing is added here that the model did not already say. The difference is that the links are now things a machine can follow, one at a time.

Fig. 1 One small system model, twice. The elements on the right are the elements on the left, and the six lines are what the prose was carrying in its grammar.

Plate II

Context wins,
until we price it

Pasting more of the model into the context window buys score, and the bill arrives with every question.

2. Context wins, until we price it

One

A context window is everything the language model can see at once.

Your question, whatever was pasted in with it, and what it has already said. It is counted in tokens, the pieces of text a language model is billed by. Stuffing is pasting source text into the question whole, nothing left out and nothing condensed, and letting the language model read the lot.

Two

Three stops were measured on the same forty questions.

Score: a scorer marks each answer field right or wrong; the share right is the question's score, averaged over forty. The forty are on this course's Apollo 11 model, written in SysML v2.

Sonnet 5 and GPT-5 mini are two language models, the mini smaller and cheaper; the view and the slice ran on Sonnet 5, the model on the mini. This view is the completeness report; talk three scores another, the traceability matrix, 0.15.

A view, a selection of the model's elements by type, rendered as text before any question was asked, scores 0.17, about one field in six. The slice, seventeen of the model's files as source text pasted in whole, scores 0.59, nearly three in five. That jump is this talk's biggest single gain.

Three

The model, all twenty-eight files, reaches 0.68, and the bill arrives with every question.

Sonnet 5's run on the model covered five of the forty and never got a forty-task number, so this stop is the mini's, read against the mini's own slice.

It ran on the mini and bills 85,018 tokens a question, about a hundred and fifty of the mini's views, paid again each time somebody asks something. On the mini the slice scores 0.58: the model buys ten points more score, 0.58 to 0.68, for twice the tokens. Nobody can call that trade without a bar, the score the task needs. Plate six sets one.

walk the stopsa view

a view: score 0.17, the baseline: one view of it.

Fig. 2 Three stops, both meters. The hatched bars are the score; the dashed bars are tokens per question, indexed to a view on the same model and drawn on a ten-fold scale.

Plate III

Retrieval sends
the nearest text

Retrieval sends the model only the ten chunks of source text that rank nearest the question: about half the points, on about a seventh of the slice's tokens.

3. Retrieval sends the nearest text

One

Retrieval means cutting the slice into chunks of source text, ranking them against the question, and sending only the nearest.

Nothing is pasted whole. The chunks are scored for how close they sit to the question's words, the ten nearest go into the window, and the language model answers from those alone.

Two

On the same forty questions it scores 0.51, on about a seventh of the slice's tokens.

Retrieval 0.51 and the slice 0.59, both Sonnet 5, forty tasks on this deck: 10,276 tokens a question against 76,658.

Against the slice it keeps nearly nine tenths of the slice's score for about a seventh of the tokens: the first stop on this rail where paying less does not mean scoring less in proportion.

Three

What it fetches is one ranking, and the lines that did not rank stay missing.

Retrieval takes one pass and stops. If the answer sat in a chunk that did not rank near the question's words, nothing in the pass goes back for it, and the answer comes back confident and short of the fact. The R2 question ranks well. Ask next what opening the valve does to R1, and the answer sits on the tank's lines, which did not.

walk the stopsretrieval

retrieval: the ten chunks nearest the question, sent.

Fig. 3 The rail from plate two, with retrieval added at the end. Below it, plate one's graph once more: at the model's stop every element is in the window; at retrieval's, only what the nearest chunks held.

Plate IV

Tools close
the rest

A model calling tools over the graph takes as many turns as a question needs. Two frozen cards say what decided their scores.

4. Tools close the rest

One

Retrieval takes one pass. An agent, a model calling tools over the graph one call at a time, takes as many as the question needs.

Here it lists the parts, reads a property, traces a link, and renders what it found: four calls, each chosen after the last one came back, so the second question can be one the first answer raised.

Two

A frozen card is a result recorded elsewhere, quoted whole with its own corpus and model. The first reads 0.806 against 0.521, and both arms, the two set-ups compared, ran the same loop.

Frozen card, the Apollo suite under Sonnet 4.5. Both arms ran the same loop: a view rendered from the graph 0.806, against the agent assembling it 0.521.

Nothing about the language model changed between those two numbers. What changed is the tool it was handed. The arm that could render a view, drawn for the question in hand, scored where the agent assembling it did not. Chapter two's view was drawn before any question arrived; that is the difference between those two uses of the word.

Three

The second card, 0.900 against 0.569, is a different corpus, the body of text its questions ran over, and a different model.

Frozen card, the Eve corpus under Sonnet 4: a command-line tool against retrieval. This deck's own agent scores 0.90, Sonnet 5, the Apollo corpus, forty tasks.

It says the same thing about tools against one retrieval pass, and it is not the same measurement. Two cards, two labels, never averaged into one. The number this deck recorded for its own agent sits under them, labelled the same way. Its 0.900 and this deck's 0.90 are two measurements that happen to round the same, not one number twice.

Fig. 4 Four calls over the same model, each lighting what it reached, and two results recorded elsewhere. The cards keep their own corpus and model, and stay separate; the last call renders the same R2 answer chapter three fetched. This deck's own agent: 0.90, Sonnet 5, the Apollo corpus, forty tasks.

Plate V

Every number
came from a scorer

Every score so far exists because the same questions were graded against a fixed answer key. Run that loop on live work and it keeps producing evidence.

5. Every number came from a scorer

One

A scorer is a program that grades an answer against a fixed key.

The questions were written once and do not change. The answers are checked the same way every time. Swap the representation, what the language model is shown, rerun the same questions, and the number that moves is the only thing we keep.

Two

Without it, every claim about what worked is an impression.

0.17, 0.59, 0.51 and 0.90 are not opinions about the four arms. They are what this loop returned, on the same forty questions and the same language model, with only the representation changed between runs. The model's 0.68 came off those same forty questions on GPT-5 mini, and says so wherever it is printed.

Three

Run the same loop on live work and it keeps producing evidence.

Every answer accepted, corrected or discarded lands on disk as a graded example, and that pile is the corpus the next decision gets read out of. Mining it is not built yet. Before plate six, write one question you would score your own assistant on, and what a right answer must contain.

Fig. 5 The loop runs while this chapter is on screen, and stops when it leaves. The mix of verdicts on the stack is drawn even: what the mix is, is what running the loop tells you. Accepted matched the key, corrected was fixed then kept, discarded means the question was bad.

Plate VI

The lowest rung,
on evidence

Choosing between a view, the slice, retrieval and the agent is a scored decision, the cheapest row that clears the bar the task needs, and those four sit at the bottom of an eight-rung ladder.

6. The lowest rung, on evidence

One

Read the rows in cost order and stop at the first one that clears the bar the task needs.

Forty tasks on this deck, all on Sonnet 5. Tokens a question: a view 868, retrieval 10,276, the agent 41,514, the slice 76,658. The model's 0.68 has no row: it ran on GPT-5 mini at other prices, and Sonnet 5's run on the model covered five of the forty, so there is no forty-task bill.

Move the bar and the pick changes. At 0.4, retrieval is the cheapest row that clears it. Guess before the bar passes 0.51: the slice or the agent? Then move it past 0.51 and read the note under the slider. Above 0.90 nothing clears.

Cost per correct answer is the forty's bill over the questions answered correct, every field right. Cheapest first: a view, $0.115 over 2, is $0.058, and it answers only two of the forty fully; retrieval, $1.063 over 17, is $0.063; the agent, $3.64 over 34, is $0.107; the slice, $6.49 over 20, is $0.324.

Two

A rung is one way of steering a language model you did not train, and the article this talk borrows its ladder from counts eight.

Context, memory and retrieval sit on rung 1: what this deck pasted in or fetched. Rung 2, system prompts, standing instructions sent with every question. Rung 3, skills, written procedures the model loads for a task. Rung 4, tools and an interface, what chapter four's agent was handed. Rung 5, deterministic hooks, a program with no model in it that checks or blocks what the model did. Rung 6, model routing, sending each question to a different model by rule. Rung 7, adapter training, a small trained add-on to the model's weights. Rung 8, training the model itself. Each step up buys durability and costs more to undo.

Three

The rule is the lowest rung that durably fixes the behavior, on evidence from the corpus: durably meaning the fix keeps holding without anyone adding it back each session, the corpus meaning chapter five's pile of graded examples.

Placing this deck's four arms on rungs 1 and 4 is this talk's reading, not the article's. The article's own words for the bottom four: they change in minutes and roll back for free. The top four are the ones this talk did not run.

minimum score0.40

At 0.40, the cheapest row that clears the bar is retrieval.

Fig. 6 The four arms in cost order, then the article's eight rungs in the same box shapes. Tokens per question are indexed to a view on a ten-fold scale. Placing these four on rungs 1 and 4 is this talk's reading.