kebnf
The OMG publishes the concrete syntax of SysML v2 and KerML as KeBNF text, not as a ready ANTLR4 or tree-sitter grammar, and KeBNF isn’t consumable by either parser generator directly: it encodes metamodel-binding rules neither one understands. kebnf reads the KeBNF source files and emits both targets from the same AST, covering the full 640 rules across both specifications. It’s a single Rust binary with its own KeBNF parser written on chumsky.
KeBNF annotates its rules for metamodel binding: type annotations, property assignments, boolean flags, cross-references, and inline semantic actions that tell an OMG-conformant tool how to build an abstract syntax tree. None of that has syntactic meaning to a parser generator, so kebnf strips it out during conversion and records it in a mapping.json file for any downstream tool that still needs the traceability.
kebnf’s tree-sitter output is a second attempt at a grammar problem I’d already solved once by hand. tree-sitter-sysml, the grammar sysml actually uses to parse SysML v2 models, was brute forced empirically over hundreds of iterations rather than derived from the KeBNF specification, and its first commit landed six minutes before kebnf’s own, on the same day. kebnf asks whether the same grammar can come out mechanically instead, straight off the OMG’s own KeBNF text, rather than hundreds of hand-tuned iterations. It hasn’t replaced anything yet: sysml still parses with the hand-tuned grammar, and the benchmark harness sits on top of that. kebnf’s own tree-sitter emitter is checked only against its own corpus, on a track running in parallel.
The interesting failure came early: a mechanical, AST-walking emitter I wrote first produced a grammar that compiled to a valid parser.c but timed out parsing a one-line file. KeBNF factors shared prefix keywords out into their own rules for metamodel binding, and tree-sitter needs that disambiguating keyword inlined early to resolve its GLR tables: a rule factored out is a keyword tree-sitter can’t see in time. Inlining every prefix back into its own definition or usage rule took parse time from a timeout to 0.15ms and cut LR conflicts from 335+ down to 199, all of them ordinary two-way conflicts instead of the mega-conflicts naive conversion produced. Full numbers and the remaining known gaps live in TREE-SITTER-FINDINGS.md and AMBIGUITY-RESOLUTIONS.md; neither backend has touched a large real-world model yet.