Skip to content
andrew.dunn.dev

sysml-bench

Source Results Dataset Pre-registration

sysml-bench asks a narrow question with a real answer behind it: does giving a tool-augmented LLM more tools actually help it answer questions about an engineering model. The model in question is SysML v2, the OMG’s textual systems-modeling language for describing requirements, structure, and behavior in system designs.

The harness drives models over a real SysML v2 corpus through CLI tool sets exposed via MCP, backed by nomograph/sysml, a knowledge-graph CLI I also built, and scores every answer against typed, published schemas rather than free-text grading.

Four tool sets are compared, from about 250 schema tokens (search plus read_file) up to about 1500 tokens (the full set, which adds a planner and a stats tool). The answer turns out to be task-dependent rather than uniform: graph-traversal tools help layer-classification tasks and hurt discovery tasks, and no single tool set wins across every category.

Tool augmentation is usually validated on retrieval or code tasks (ReAct, Toolformer). I wanted the same question asked of a structured domain model instead of prose, and I wanted the answer treated as provisional until it had been tested on a second, independent corpus rather than reported once and left alone.

claimsFED IN6 corpora88 typed tasks4 tool setsHARNESSthe runtool sets over MCPN runs per cellGRADINGtyped scorerper field, againsta published schemaCOMES OUTresultsrepo, HuggingFaceexport + defectsGATEpre-registrationhypothesis, direction and snapshot frozenbefore any Apollo score existed
Corpora, tasks and tool sets go in over MCP, a typed scorer grades every answer field, and a result only stands if the pre-registration allowed it.
answersscoresscores ingatescorporasix public SysML v2 modelsEve, 19 files, primaryApollo 11, 28 files, replicationtask suite88 tasks, 8 categoriesa published answer schemaon every taskfour tool setssearch + read_file, ~250schema tokens, up to ~1500with planner and stats addedthe runtool sets over MCP, backed by nomograph/sysmlone condition per model and tool set, N runs eachtyped scorerper field: Bool, Float, Str, StrContains,ListStr by F1; task = mean of its fields,condition = mean over runspublished resultsresults repo and a HuggingFace export,gated by validate_export()defects logged, not filteredpre-registrationhypothesis, direction, and model snapshotcommitted before any Apollo scorefrozen analysisone-sided permutation test, Hedges g,BCa bootstrap CI at alpha = 0.025
Three input sets feed one run over MCP, the typed scorer turns answers into per-field scores, the results publish with their defects logged, and the frozen analysis only draws the conclusion the pre-registration committed to first.

Highlights

  • Scoring is per-field and typed, not an LLM judge by default: Bool, Float within tolerance, Str with qualified-name suffix matching, StrContains, and a set-based ListStr scored by F1 threshold. Task score is the mean of field scores, condition score the mean of task scores over N runs. See eval/scoring.py.
  • 88 tasks across 8 categories run against the primary corpus, a 19-file Eve Online Mining Frigate SysML v2 design, plus five more public corpora, including the 28-file Airbus Apollo 11 model, for scaling and cross-corpus generalization.
  • Mean discovery score falls from 0.880 at 19 files to 0.423 at 95 files, and neither graph tools nor vector search recovers it: the primary corpus is small enough that scaling behavior is a real open question, not a settled one.
  • Sonnet scores 0.880 on discovery tasks against 0.529 for the best OpenAI model tested (gpt-4o-mini), a 35-point gap, though gpt-4o-mini is also roughly 87 times cheaper to run per task.
  • The Apollo 11 replication was pre-registered before any Apollo score existed: hypothesis, direction, model snapshot, and the permutation-test analysis script were committed and frozen at a650aa2 (MR !27), with inference gated at a one-sided permutation test, Hedges g, and a BCa bootstrap CI at α = 0.025.
  • Between freezing the pre-registration and running it, the pinned model snapshot was retired from the API outright. The registered deviation process required disclosing that in the document, switching to the nearest available snapshot, and rerunning the entire corpus-1 baseline on the new snapshot to bound drift, all before any Apollo score was observed.
  • Cutting the HuggingFace dataset export past 0.1.1 surfaced real defects, not polish: a schema mismatch broke load_dataset() for every caller, 175 task rows exported an empty ground-truth object because the loader read the wrong field, a task_id collision across two suites silently attached the wrong answer key to half of one suite, and duplicate ingestion had inflated the results split by 960 rows. All fixed and gated by a validate_export() check before 0.1.2 shipped.
  • Both repos disclose data quality rather than filter it: the known-defects log documents that about 4.5% of the published 11,986-row export carries a harness or provider failure recorded as score 0.0 (treat as missing, not as a model failure, before averaging), and CI’s pytest stage explicitly treats zero collected tests as success, which is what this repo currently has. Correctness rests on the frozen analysis script, validated to reproduce the corpus-1 numbers exactly, not on a conventional test suite.
  • MIT licensed.

I don’t know why a newer, more capable snapshot would erase the guidance effect rather than preserve it.

The pre-registration doesn’t resolve that either; it only establishes that the effect is tied to a specific snapshot rather than to the task. That’s one open question sitting on top of the harness, and the frozen analysis script is already in place for whoever runs the next confirmatory test.