Talk three of three · October 2026

Structure
over the trove

Meta models, SysML v2, and one question asked four ways

This talk lands in the week of 26 October 2026.

AI in Systems Engineering · week 9 · Andrew Dunn

Scroll to run it Read on: the plates are printed at their finished state

Plate I

What a meta
model is

A meta model is the small set of types and relationships laid over a trove of information, and SysML v2 is one built for engineering.

1. What a meta model is

One

The model this talk runs on arrives as a pile of files.

Apollo 11 SysML v2 model, 28 files, 7,249 lines. Airbus, MPL-2.0. One question of its forty runs through every plate, AP13: the vacuum thrust of the F-1 and the J-2 engines, the push with no air around them, in kilonewtons, and the J-2's vacuum specific impulse, the seconds a unit of propellant weight keeps pushing. Three lines of TechnicalComponentsPackage.sysml hold the answers: thrustVacuum = 7770 ['kN'], = 1033 ['kN'], specificImpulseVacuum = 421 ['s'].

It is the Apollo 11 mission written as an engineering model: the spacecraft, the requirements it had to meet, the actions the crew and the vehicle perform, and the phases of the flight. Nothing in the pile tells a program which of those a given line is.

Two

A meta model does not read the pile. It types it.

From the shipped model: 194 part definitions, 296 requirement definitions, 110 action definitions, 16 state definitions, inside 2,544 typed elements.

A meta model is the small set of types and relationships a trove of information is allowed to use: this is a part, this is a requirement, this is an action, and these may connect to those. SysML v2 is one of them, fixed by a standard before anyone describes a spacecraft, so every tool reads the same vocabulary.

Before the stamps there was nothing to pick by: RocketEngine was a string and part def two words. Talk one ended by sending you to the larder lesson, where an agent was handed a pantry file, a preferences file and a folder of recipes: a list of what is in the kitchen, with nothing in it saying what any line is. Typing the pile is the other move.

Three

Typed, the pile sorts; and a selection by type, rendered, is a view.

The view for AP13 is four lines of 7,249: 179, 180, 185 and 186 of TechnicalComponentsPackage.sysml.

Each stamped line is an element, one typed thing. A view is a selection of elements by type, rendered, that is, written out as text; SysML v2 spells it expose, filter hastype, render, and a view satisfies a viewpoint, the question it was drawn for. The Apollo modellers left one in the model, RequirementsTableView, its expose and filter lines commented out.

Ours, for AP13: the part definitions that specialise RocketEngine, with what they redefine. Four lines come back out of the pile, and the three answers are on them. Line 180, the F-1's specific impulse, 304 s, came too: AP13 never asked for it, but the pick was by type, and a pick by type brings every value it covers, asked or not. That is why one view can answer the next question as well, and why a trap can sit beside an answer.

Stuffing is pasting source text into the question whole; the slice, seventeen of these files as source text, is what chapter four's ladder stuffs. Guess now, and hold it: does stuffing the whole pile into the question beat typing it first?

Fig. 1 Six lines of the Apollo model first, each stamped with what the meta model says it is; then the 28 files on the left and what the meta model finds inside them on the right. The counters are the shipped model's own, counted by type. Last, one view drawn from the types for AP13: the part definitions that specialise RocketEngine and what they redefine come back out of the pile as four lines, the three answers ringed; chapter four prices the view and runs it on the forty. Apollo 11 SysML v2 model, Airbus, MPL-2.0.

Plate II

The graph is parsed,
not guessed

Once elements are typed the links between them are just as real, and a parser builds that graph from the syntax without asking a language model to find anything.

2. The graph is parsed, not guessed

One

A parser reads the source text by its grammar, the way a compiler reads code.

It runs over the files once and returns what the syntax says is there: every definition, every usage, and every link one of them declares to another. The result is a knowledge graph, which is the same model written so a program can walk it: elements are nodes, relationships are edges, and each keeps its own name.

Two

The links carry as much of the engineering as the elements do.

From the shipped model: 272 satisfy links, 104 perform links, 14 phase states and 14 transitions, one of them the entry.

A satisfy link says which part answers which requirement. A perform link says which part carries out which action. The mission chain says which phase follows which, and what event ends the one before it. None of that was inferred. It was written down, and the parser read it.

Three

Nobody has to ask a language model where the parts are.

A codebase comparison put parsed graphs ahead of model-built ones; the model-built graph skipped 377 files on one repository.

Asking a model to extract the graph adds a step that can be wrong, and it is wrong in ways nothing downstream can see. Parsing is cheap, it is repeatable, and when the model changes the graph changes with it on the next run.

Fig. 2 The pipeline along the top runs once: text, parse, index, and then the four operations the graph answers. Below it the links draw in three waves, in the counts the shipped model carries.

Plate III

One question,
two shapes

Hold the question fixed and change only how the same files reach the model, and the answer moves.

3. The duel: one question, two shapes

One

A representation is what a prompt keeps of the facts, and the shape it keeps them in.

Not a different question. The same seventeen files of the model, handed over twice: once as the source text they are written in, once as a table rendered from them. A duel holds everything else still and changes only that.

Two

On the same forty questions the two shapes land 40 points apart.

Forty tasks on this deck, answered by one language model, Sonnet 5: the slice, the files as source text, 0.59 at 76,658 tokens in per question, on average, the rendered requirements table 0.19 at 7,906. Tokens are the pieces of text a model reads and is billed by. A score is the share of a question's answer fields answered right, averaged over the forty; counted correct, every field right, the slice answers 20 of 40 and the table 3, which is how the lesson counts.

The table is smaller, tidier, and easier to read. It also throws away most of what the questions ask about, and the score says so. Shape decides which facts are still in the room when the question is read.

Three

Spell one selection two ways and the answer moves far less.

A third selection, rows pulled from the graph, spelled two ways: as plain rows 0.12, in TOON, a compact row format, 0.14; Sonnet 5 over forty tasks.

Take one selection out of the graph and write it two ways and the answer moves by about two points. Representation is selection first, spelling second: the forty points in step two are what the table dropped; these two are the spelling. The ladder is about the first. Next, four ways to hand the model over: the matrix, the slice, retrieval, the agent. Pick one in your head: which rung wins on the forty, and which costs least per correct answer?

the source text: 0.59 (20 of 40 correct), 76,658 tokens in per question, on average.

Fig. 3 One question pinned, AP13. Both panes are made from the same seventeen files, the source text whole and a table rendered from them; press one and its pane opens with a line ringed, the lesson's hit or miss ring. The needle travels between the two positions the recording holds for all forty. Apollo 11 SysML v2 model, Airbus, MPL-2.0.

Plate IV

Four rungs, and
retrieval is one

There are four ways to hand a model this graph, and retrieval is a rung rather than the finish line.

4. The ladder: four rungs, and retrieval is one of them

One

Four ways to hand over the model, measured on the same forty questions.

The bottom rung is the matrix, the traceability matrix: every requirement with what satisfies it and what verifies it, a view drawn for traceability and not for AP13, so it has no engine row. Chapter three's table is a view too, and so was the modellers' own. The matrix averages 0.15, 1 of 40. The slice, seventeen files of source text pasted in whole, averages 0.59: the biggest single gain on the ladder, paid in tokens every time somebody asks.

Two

Retrieval fetches the part of the graph a question touches, and sends only that.

Retrieval 0.51 at 10,276 tokens in per question, on average, against the slice's 0.59 at 76,658. Sonnet 5, forty tasks. Sonnet 5 bills $2.00 per million tokens in and $10.00 per million out, the vendor's price as of September 2026; every Sonnet 5 dollar on these plates is tokens times that.

Retrieval, which the ML room calls RAG, cuts the slice into chunks of source text, ranks them against the question, and sends the ten nearest. The arm averages about half the points for about a seventh of the slice's tokens. Cost per correct answer is the forty's bill over the questions answered correct: the slice's $6.49 over 20 is $0.324. Per correct, retrieval is cheaper, $0.063 on 17, and stops there; the matrix's $0.265 buys 1. Read the count beside the price.

Three

The three obvious ways to improve that pass were each measured, and each came back null.

All three on the first corpus under Sonnet 4: vector search against plain search 0.880 to 0.880; graph traversal at two to three hops, links followed out from a match, 0.493 against search's 0.528, at about 69 times the tokens; planning tools +0.035 on hard tasks, underpowered.

The typed view, run beside the ladder and not on it: a view drawn by type for each question, by a rule written down before the run, pre-registered 2026-10-01, chapter one's four lines among AP13's rows. On AP13 it answers 3 of 3 for $0.0033, 1,366 tokens in, Sonnet 5; over the forty, 13 of 40, mean 0.42, at $0.0212 per correct, below retrieval's $0.063, the least of the four rungs; and it stops at 13, where the slice answers 20.

The top rung lands here: the agent's $3.64 over 34 is $0.107 per correct. Switch a probe on and the bar moves and settles back where the recording left it. Check both guesses. Chapter one's: the slice beats the matrix, 20 to 1, and beats the typed view on count, 20 to 13, though not per correct answer, the margin has the bill; the agent walking the typed model beats all three at 34, and chapter five is how. Step three's: the agent wins; of the four rungs retrieval costs least per correct, on 17 right, loses to the slice, and on AP13 fetched the declaration, not the redefinition.

the matrix 0.15 (1 of 40) · the slice 0.59 (20) · retrieval 0.51 (17) · the agent 0.90 (34), Sonnet 5; the count is questions answered correct, every field right. Switch a probe on: the retrieval rung moves and settles back where the recording left it.

Fig. 4 Four rungs, correctness against tokens per question, each labelled with its cost per correct answer over the forty; the top rung, the agent, is a model calling tools over the graph, chapter five's subject. Beside each rung, AP13's three values as lamps, lit where that rung's prompt held them and ringed dashed where it did not, the reading the four envelopes opened in the room give. The envelopes carry AP13's own bills, one question each; the plate's tokens are the forty's averages. The three probes on the retrieval rung are the paper's null results, each recorded on the first corpus under Sonnet 4 and printed with its own number.

Plate V

The agent walks
the graph

The top rung is not a bigger view. It is an agent making one call at a time against the graph and reading only what the question needs.

5. The agent walks the graph

One

An agent is a model that calls tools over the graph, one call at a time.

AP13 on this deck's agent, Sonnet 5, from the lesson's recording: three turns, each billed for everything before it, the question and every result so far; 2,842 tokens in after the first, 8,837 after the second, 18,753 after the third; 295 out; $0.040.

It is handed no view. It searches for an element, traces a link, inspects what came back, and renders what it found, choosing each call after the last one returned. Those four operations are the paper's archetype, the ones chapter two's index was built to answer; the deck's agent has seven tools over the parsed model and its files, and on AP13 it needed three calls, two searches and one file read, before it answered.

Two

On this deck the agent reaches 0.90, and the smaller model reaches 0.83.

Forty tasks on this deck, the Apollo model: the agent 0.90 (34 of 40 correct) on Sonnet 5 at 41,514 tokens in per question, on average, 0.83 (31 of 40) on GPT-5 mini at 28,432. A live instrument moves; both numbers are dated where they are printed.

It reads about half the slice's tokens and answers more of the questions than any single shape of it does. The gain comes from asking narrow questions instead of one wide one.

Three

Two frozen results sit either side of this one, and neither is the comparison the chapter just made.

Pre-registered: the rules and predictions were written down and committed before any call. Apollo replication, Sonnet 4.5, explanation tasks, both arms agent loops: a view rendered from the graph 0.806 against the agent assembling those facts call by call 0.521: g=0.83, a large gap, about four fifths of a standard deviation; exact p=0.0007, seven chances in ten thousand of a gap that size with no real difference. First corpus, Sonnet 4, discovery tasks: an agent over command line search 0.900 against vector retrieval 0.569, p=0.018, about two in a hundred. Each number licenses a claim about its own corpus and model, not about this deck.

Guess before the cards turn: a view drawn for this question, or the agent building it? Drawn for the question, the view wins; the agent beats the matrix, a view drawn for something else, chapter four's bottom rung, and beats vector retrieval on the second card. On this deck the agent beats our typed view too, 34 to 13: ours is drawn once by a fixed rule and sent in one call, where the card's view was rendered inside an agent loop; the two are not averaged. The registration carried a second effect, tool-selection guidance, which mostly vanished on replication and was reported as a failure to replicate, which is what registering it was for.

this deck's agent: 0.90 (34 of 40 correct), 41,514 tokens in per question, on average, Sonnet 5; 0.83 (31 of 40) on GPT-5 mini at 28,432. The cards: a view rendered from the graph 0.806 against the agent assembling it 0.521, Apollo, Sonnet 4.5; command line search 0.900 against vector retrieval 0.569, the first corpus, Sonnet 4.

Fig. 5 Chapter two's field again, its names on plate two. The marker walks the archetype's four operations while the score climbs; the tape under the field is the run the deck's agent recorded on AP13, Sonnet 5, one line per call, what it asked and what came back: two searches in one turn, then one read, lines 1 to 230 of TechnicalComponentsPackage.sysml, then the three values. Each turn is billed for everything before it, the question and every result so far; the bill is in step one's margin. The two frozen cards beside it keep their own corpus and model, their scores in the note above.

Plate VI

A fast model answers
a closed question

A fast model reads what retrieval fetched once and answers a closed question, yes or no, with a probability attached; it writes no prose, and its questions are written before the run.

6. A fast model answers a closed question

One

Open: what is the F-1's vacuum thrust? Closed: does what retrieval fetched hold it, yes or no, with a probability.

A vendor's label, September 2026: a glossary entry, one hosted product, one open-weight model of the same shape.

Chapter five's agent answers the open one in free text, as many calls as it needs. A closed question's answers are fixed before it is asked; the probability lives on that set. Because the set is fixed, every answer can be marked right or wrong against a label, here whether retrieval's own answer was right; the label is what lets the gate be scored, against a threshold written down before the run. An answer in free text has no such label. A fast model reads what retrieval fetched once and answers only closed questions. The vendor's label for it is system-one, after Kahneman's two systems; we take only the division of labour, a cheap judgment and an expensive one over the same fetched text.

Two

Measured independently it ranks well, agrees with human labels, and costs a fraction of a judge.

Guo et al. 2026, arXiv:2609.29429: zero-shot detection of ten alignment-failure types, 44 benchmarks, 7,193 instances; median AUROC 0.886, the chance a true case scores above a false one, where 0.5 is a coin and 1 is perfect; kappa 0.809, agreement with human labels beyond chance, about as much as two humans agree with each other, 0.811. One group, one model, their benchmarks and not this deck.

It ranks: what it scores higher is more often right. That is the shape of a verifier, a cascade or a router, each older than the label; the rail dates them. New this year is the fixed question set, with calibration a stated training objective.

Three

Its probabilities hold across a pool, drift within one dataset, and give no reason to audit.

Guo et al., appendix C: calibration error 0.047 over the pool, the stated probability off by about 5 points; within any one benchmark it drifts, median 0.168, about 17 points, against 0.074, about 7, expected.

Calibration is a group property: pooled, the probabilities look honest; inside one dataset the base rate shifts and they are off while the ranking holds. An answer cannot fall outside its question's set; it can be the wrong one. It earns its place as every arm here did: a labelled set, a threshold, a measured cost per correct.

enough to answer? runs no to yes; on AP13 the gate read 0.05, the needle at no, and the question walked. Which call next? one of the four operations, trace ringed as the call after search; does it follow? a score in five cells, the first on AP13. Never outside its question's set; its probability is calibrated over a pool, not within one dataset. The rail: cheap decisions before the expensive call.

Fig. 6 Chapter five's field, its names on plate two, with the elements retrieval fetched lit, what a fast model reads once. Beside it three closed questions with their answer spaces; the measurements are in the margin with their paper and denominators. The rail under the field dates the idea: cheap decisions before the expensive call. The open question on the left is chapter five's; the closed ones on the right are the fast model's, and the first is the gate's.

Plate VII

The harness is
the design unit

The unit to design and measure is a harness over the meta model: the meta model gives the structure, a fast model makes the closed calls, a reasoning model works within what those admit, and tools act; each part measured alone.

7. The harness is the design unit

One

Four kinds of part on one drawing: structure, fast calls, reasoning, tools.

Shipped harnesses already route between models, gate risky tool calls and classify agent state with a fast model (LangChain, 2026). Stacked imperfect modules can amplify each other's faults (Shefa et al. 2026, arXiv:2609.03230): measure each part alone.

The meta model is the part that never guesses; chapter two parsed its links from the syntax. The fast model answers closed questions over what retrieval fetched; the reasoning model does the work those answers admit; tools act.

Two

One gate, one closed question: cost per correct answer down 24.0% on this deck, no answer lost.

E1, the gate. Our pilot, pre-registered, run 27 September 2026 on one model snapshot: “is this enough to answer?” over 118 labelled fields, forty tasks, three repeats; AUC, the chance a true case scores above a false one, 0.954, where 0.5 is a coin and 1 is perfect; plain word matching, BM25, 0.702. Routed at 0.5 on Sonnet 5: 34 of 40 correct, 24.0% less per correct, on this pilot; on AP13 the gate read 0.05 and the question walked. The lesson's seventh step, #g=the-gate, runs it with the threshold in your hand.

The headline is read at 0.5, the threshold written down before the run, and on this pilot: 24.0% less per correct, none lost. A cut chosen after the run, at the best-looking threshold, is a look and not a result; it would make the saving the sweep's, not the gate's. Retrieval answered seventeen of the forty in chapter four; the gate's job is to know which before the agent is paid to walk. The tally under the plate is chapter four's cost per correct, with the gate and without; the gate's forty calls are on the bill. A perfect gate would have chosen the same here, so this run cannot tell the two apart. The gate asks chapter three's question the other way round: is what is in the room enough?

Three

Asked to find structure without names, the same model failed: the edge of a fast call.

E2, the links. Same pilot: “does this part satisfy this requirement?” over 672 pairs, three repeats. AUC again, as pairs: names kept 0.791, 79 times in 100; scrambled 0.528, 53; under word matching 0.547, 55; a coin, 50. Refuted as registered: the rule written before the run called the prediction refuted if the scrambled score fell below word matching's, and 0.528 is under 0.547. Whether scrambling is a fair test is open.

The named score was the names; without them the model guesses at a link the parser read from the syntax. Refuted means the prediction we wrote down failed, not the harness: the parser already reads that link, so the division holds, structure from the meta model, closed calls over it, reasoning within what they admit. Our forecast, and no source we checked makes it: meta-modeling is chapter one's move one level up, typed judgments about the models themselves.

with the gate: 17 to retrieval, $0.32, all right; 23 walked, $2.43, 17 right; forty gate calls, $0.01. $2.77 over 34 is $0.081 per correct; alone, $0.107: 24.0% less, none lost.

Fig. 7 Chapter five's walk with a gate in front of it. The fast model's one closed question decides whether what retrieval fetched answers or the agent walks; the tally is cost per correct answer against the agent alone, and the two cards are our pilot's two questions, each stamped with its verdict, labelled one pilot on this deck.