Skip to content
andrew.dunn.dev

sysml

Source crates.io

sysml is a CLI-native knowledge graph toolkit for SysML v2, the OMG systems-engineering modeling language. It’s a single Rust binary that runs in two modes, a clap CLI for humans and CI, and a stdio MCP server for agents, both delegating to the same sysml-core domain logic so neither mode duplicates the other’s behavior. Parsing comes from tree-sitter-sysml, a grammar I also maintain; sysml builds a queryable knowledge graph on top of it and exposes it through commands like search, trace, check, and query.

ONE RUST BINARYSYSML V2model files.sysml sourcesGRAMMARtree-sitter-sysmlits own repoCOREsysml-coreknowledge graphHUMANS, CIclap CLI~50 tokensAGENTSMCP server~15,000 tokens
One Rust binary parses model files into sysml-core, then serves that one graph as a 50-token CLI call or an MCP server costing an estimated 15,000.

The design follows GitLab’s Global Knowledge Graph architecture (tree-sitter parsing into a property graph, then agentic traversal over it), but the working thesis is that CLI beats MCP as the transport: a trace call costs on the order of 50 tokens against an estimated 15,000 for an MCP tool’s schema load. That 15,000 is an estimate, and nothing has measured it. The companion benchmark harness doesn’t settle it either, because it varies tool sets rather than transports and exposes every condition over MCP; its four tool sets bill between about 250 and 1500 schema tokens, an order of magnitude under my estimate, which is a reason to distrust the estimate before the thesis. The comparison I actually want is still unrun.

Highlights

  • The 9 coverage tests in crates/sysml-core/src/graph.rs, documented in CONTRIBUTING.md, parse the test corpora on every cargo test --workspace run and fail by name when tree-sitter-sysml’s grammar adds a node type the walker doesn’t know about yet, rather than silently dropping it.
  • Those tests caught a real bug: query and trace had hardcoded skip_kinds = ["import", "member"], hiding 848 of 1454 relationships (58%) from every graph command until Phase 9.5 removed the filter from query and made it opt-in on trace via --include-structural.
  • Test fixtures are two real SysML v2 models, not synthetic ones: an LLM-generated Eve Online mining frigate (19 files, BSD-3-Clause) and Airbus’s Apollo 11 CoSMA mission model (28 files, MPL-2.0), which sysml-bench’s own README calls the largest public real-world SysML v2 model.
  • cargo-deny gates the pipeline with audit_allow_failure: false. Three transitive RUSTSEC advisories are waived in deny.toml with a one-line reason each (an unmaintained paste crate, an unsound rand::rng() path this code doesn’t use, a bincode 1.3.x pin inside hnsw_rs), and nothing else is allowed through.
  • The mcp feature wasn’t covered by a plain cargo clippy --workspace, so a broken build under --features mcp shipped once before the pipeline moved to an all-features gate; one commit fixed the break, a second closed the gap in CI so it can’t happen quietly again.
  • CI is a fail-closed automerge gate, not just a green pipeline: osv-verdict and cargo-verdict each report a check, and automerge-gate blocks the merge unless both are present and passing, which is the signal Renovate’s automatic merges read.
  • 10 MCP tools, 0 prompts (an early CHANGELOG draft claimed 15 tools and 4 prompts before a docs pass corrected it). Five CLI surfaces, plan, skill, --format, search --layer, trace --trace-format, have no MCP equivalent yet, tracked as a known gap in CONTRIBUTING.md.
  • Ships two reusable GitLab CI components for consuming projects: a model validation gate (syntax check plus a completeness badge) and an MR model diff comparing the base branch’s index against the head branch’s.

Vector search behind fastembed is written but deferred: structural signals alone carry the corpora sysml ships against, and past that range the benchmark says vector search doesn’t rescue discovery either, so it’s unclear it earns its weight. The scaling wall in the Aside above is the question I’d want answered first.