sysml-bench
sysml-bench asks a narrow question with a real answer behind it: does giving a tool-augmented LLM more tools actually help it answer questions about an engineering model. The model in question is SysML v2, the OMG’s textual systems-modeling language for describing requirements, structure, and behavior in system designs.
The harness drives models over a real SysML v2 corpus through CLI tool sets exposed via MCP, backed by nomograph/sysml, a knowledge-graph CLI I also built, and scores every answer against typed, published schemas rather than free-text grading.
Four tool sets are compared, from about 250 schema tokens (search plus read_file) up to about 1500 tokens (the full set, which adds a planner and a stats tool). The answer turns out to be task-dependent rather than uniform: graph-traversal tools help layer-classification tasks and hurt discovery tasks, and no single tool set wins across every category.
Tool augmentation is usually validated on retrieval or code tasks (ReAct, Toolformer). I wanted the same question asked of a structured domain model instead of prose, and I wanted the answer treated as provisional until it had been tested on a second, independent corpus rather than reported once and left alone.
Highlights
- Scoring is per-field and typed, not an LLM judge by default:
Bool,Floatwithin tolerance,Strwith qualified-name suffix matching,StrContains, and a set-basedListStrscored by F1 threshold. Task score is the mean of field scores, condition score the mean of task scores over N runs. Seeeval/scoring.py. - 88 tasks across 8 categories run against the primary corpus, a 19-file Eve Online Mining Frigate SysML v2 design, plus five more public corpora, including the 28-file Airbus Apollo 11 model, for scaling and cross-corpus generalization.
- Mean discovery score falls from 0.880 at 19 files to 0.423 at 95 files, and neither graph tools nor vector search recovers it: the primary corpus is small enough that scaling behavior is a real open question, not a settled one.
- Sonnet scores 0.880 on discovery tasks against 0.529 for the best OpenAI model tested (gpt-4o-mini), a 35-point gap, though gpt-4o-mini is also roughly 87 times cheaper to run per task.
- The Apollo 11 replication was pre-registered before any Apollo score existed: hypothesis, direction, model snapshot, and the permutation-test analysis script were committed and frozen at
a650aa2(MR !27), with inference gated at a one-sided permutation test, Hedges g, and a BCa bootstrap CI at α = 0.025. - Between freezing the pre-registration and running it, the pinned model snapshot was retired from the API outright. The registered deviation process required disclosing that in the document, switching to the nearest available snapshot, and rerunning the entire corpus-1 baseline on the new snapshot to bound drift, all before any Apollo score was observed.
- Cutting the HuggingFace dataset export past 0.1.1 surfaced real defects, not polish: a schema mismatch broke
load_dataset()for every caller, 175 task rows exported an empty ground-truth object because the loader read the wrong field, atask_idcollision across two suites silently attached the wrong answer key to half of one suite, and duplicate ingestion had inflated the results split by 960 rows. All fixed and gated by avalidate_export()check before 0.1.2 shipped. - Both repos disclose data quality rather than filter it: the known-defects log documents that about 4.5% of the published 11,986-row export carries a harness or provider failure recorded as score 0.0 (treat as missing, not as a model failure, before averaging), and CI’s pytest stage explicitly treats zero collected tests as success, which is what this repo currently has. Correctness rests on the frozen analysis script, validated to reproduce the corpus-1 numbers exactly, not on a conventional test suite.
- MIT licensed.
I don’t know why a newer, more capable snapshot would erase the guidance effect rather than preserve it.
The pre-registration doesn’t resolve that either; it only establishes that the effect is tied to a specific snapshot rather than to the task. That’s one open question sitting on top of the harness, and the frozen analysis script is already in place for whoever runs the next confirmatory test.