#benchmarking
2 entries
projects · March 2026
sysml-bench
An evaluation harness asking whether giving a tool-augmented LLM more tools actually helps it answer questions about a SysML v2 engineering model. 88 tasks scored per field against published schemas, run across four tool sets and six public corpora, with a pre-registered replication whose frozen hypothesis did not survive the independent corpus.

writing · March 2026
Meeting Notes RAG
A RAG pipeline for searching customer meeting notes: embedding, vector search, and synthesis without a SaaS intermediary.