Skip to content
andrew.dunn.dev

kebnf

Source crates.io

The OMG publishes the concrete syntax of SysML v2 and KerML as KeBNF text, not as a ready ANTLR4 or tree-sitter grammar, and KeBNF isn’t consumable by either parser generator directly: it encodes metamodel-binding rules neither one understands. kebnf reads the KeBNF source files and emits both targets from the same AST, covering the full 640 rules across both specifications. It’s a single Rust binary with its own KeBNF parser written on chumsky.

KeBNF annotates its rules for metamodel binding: type annotations, property assignments, boolean flags, cross-references, and inline semantic actions that tell an OMG-conformant tool how to build an abstract syntax tree. None of that has syntactic meaning to a parser generator, so kebnf strips it out during conversion and records it in a mapping.json file for any downstream tool that still needs the traceability.

kebnf’s tree-sitter output is a second attempt at a grammar problem I’d already solved once by hand. tree-sitter-sysml, the grammar sysml actually uses to parse SysML v2 models, was brute forced empirically over hundreds of iterations rather than derived from the KeBNF specification, and its first commit landed six minutes before kebnf’s own, on the same day. kebnf asks whether the same grammar can come out mechanically instead, straight off the OMG’s own KeBNF text, rather than hundreds of hand-tuned iterations. It hasn’t replaced anything yet: sysml still parses with the hand-tuned grammar, and the benchmark harness sits on top of that. kebnf’s own tree-sitter emitter is checked only against its own corpus, on a track running in parallel.

640 rulesrecordsOMG SPECKeBNF sourceKerML, SysML v2ONE ASTkebnfchumsky parserGRAMMARANTLR4CI runs TestRigGRAMMARtree-sitter199 LR conflictsSIDE OUTPUTmapping.jsonmetamodel binding
kebnf parses the OMG's KeBNF text for KerML and SysML v2 into one AST, then emits an ANTLR4 grammar and a tree-sitter grammar from it and records the metamodel binding it strips to mapping.json.
recordsONE KEBNF AST, TWO PARSER GENERATORSOMG SPECKeBNF sourceKerML and SysML v2 concrete syntaxpublished by the OMG—fetch-spec pulls upstreamONE ASTkebnfone Rust binary, a chumsky parser, one ASTmetamodel binding stripped hereSIDE OUTPUTmapping.jsontypes, properties, flags, cross-refs,semantic actions, kept for downstream toolsGRAMMARANTLR4emitted straight from the ASTCI compiles it, TestRig parses a SysML v2 fileGRAMMARtree-sittershared prefix keywords inlined into everydefinition and usage rule, so GLR resolves199 LR conflicts (335+ naive), corpus 159/192PER-RULE CONVERSION, 640 RULES247 direct, 353 strip-and-convert37 best-effort, 3 flagged for manual reviewvalidated to different depths; neither has meta large real-world model yet
Both grammars come off one parsed AST, so the 640 rules convert once, but the two backends are validated to different depths and neither has met a large real model yet.

The interesting failure came early: a mechanical, AST-walking emitter I wrote first produced a grammar that compiled to a valid parser.c but timed out parsing a one-line file. KeBNF factors shared prefix keywords out into their own rules for metamodel binding, and tree-sitter needs that disambiguating keyword inlined early to resolve its GLR tables: a rule factored out is a keyword tree-sitter can’t see in time. Inlining every prefix back into its own definition or usage rule took parse time from a timeout to 0.15ms and cut LR conflicts from 335+ down to 199, all of them ordinary two-way conflicts instead of the mega-conflicts naive conversion produced. Full numbers and the remaining known gaps live in TREE-SITTER-FINDINGS.md and AMBIGUITY-RESOLUTIONS.md; neither backend has touched a large real-world model yet.