Skip to content
andrew.dunn.dev

jig

Source crates.io

jig points a real agent at a real CLI and counts what happens. Each subject tool declares a battery in an agent-shape.toml (TOML): a fixture to set up, tuning and holdout tasks, success criteria, and a rubric. jig spawns the runtime (today Claude Code headless, claude -p --output-format stream-json) against the fixture, captures every bash command the agent tries, and scores the transcript with an LLM-as-judge. The metric I actually care about is first-try command success: whether the agent’s first attempt was a real subcommand that worked, or whether it invented one, or gave up and fell back to grep/jq against the fixture files instead of the tool.

spawnsinvokestranscriptverdictjig check guards the rubricagainst drift in the CLI —helpBATTERYagent-shape.tomltasks, criteria, rubricRUNTIMEheadless agentevery bash command loggedSUBJECTthe CLI under testreal command or inventedJUDGELLM judgefour-point scaleLEDGERJSONL trialsappend-only, resumable
A battery declared in agent-shape.toml spawns a headless agent at the subject CLI, and an LLM judge scores the transcript into an append-only JSONL ledger.
batterytranscriptverdictthe agent invokes it—helprubric driftoffline, no tokensHOW ONE TRIAL GETS SCOREDdouble_score judges twiceand reports the deltabatteryagent-shape.tomlfixture, tuning, holdouttasks, criteria, rubricrunspawns the runtimeheadless, captures everybash command it triesLLM judge1.0 real command, first try0.5 invented one first0.25 raw shell, not the tool0.0 did not completejig checkrubric top_level againstthe binary —help, inboth directionssubject CLIthe tool under testits —help is theground truthtrial ledgerappend-only JSONL, one(trial, verdict) per linea killed run resumesoffline viewsjig render and jig compare, reading saved JSON reportsno agent or judge tokens
One trial runs from battery to ledger: a headless agent works the subject CLI, an LLM judge marks the transcript on four points, the verdict appends to resumable JSONL, and jig check watches the rubric for drift against the CLI --help.

I built this to answer a narrower question than “is my CLI good”: when an agent reaches for a command that doesn’t exist, is that the agent’s fault or the surface’s? jig runs the same battery against itself, which is a useful forcing function; its own agent-shape.toml only exercises the offline subcommands (check, render, compare), because run and rejudge need Claude in the loop and cost real tokens.

Highlights

  • Judge rubric is a 4-point scale, not pass/fail: 1.0 for a first command that was real and worked, 0.5 for completing but reaching for a nonexistent command first, 0.25 for falling back to raw shell instead of the tool, 0.0 for not completing.
  • jig check --binary <bin> cross-references the rubric’s [commands].top_level against the binary’s --help output in both directions, catching rubric drift before it produces phantom “invented command” scores. Per the README, rubric drift is the dominant source of measurement error, and the fix is landing subcommand changes and rubric updates in the same commit.
  • Optional double_score runs the judge twice per trial for inter-rater reliability, reporting the absolute score delta. CLAUDE.md records typical IRR delta as 0.05 to 0.30 at n=5 and says not to trust anything under 0.20 without n>=20 or an effect-size test like Cliff’s delta.
  • Trial checkpointing is append-only JSONL, one completed (trial, verdict) pair per line, so a killed run resumes instead of re-spending judge and agent tokens. jig render and jig compare then work entirely offline against saved JSON reports.
  • Tuning and holdout batteries are separate fields in the schema from v1, but the holdout set ships empty by design: it’s meant to be written by an author who hasn’t seen the tuning tasks, and nobody has done that yet.
  • CI is a fail-closed chain built on nomograph/pipeline components: OSV and cargo verdicts plus a blocking secret scan feed an automerge gate that requires all three present and passing before GitLab’s native automerge fires.
  • House style bans em dashes everywhere in the shipped surface, source comments included, on the grounds that they’re an LLM tell; #![deny(warnings, clippy::all)] at the crate root means no silent #[allow(...)] either.
  • MIT licensed, published as nomograph-jig on crates.io.

The runner only knows how to spawn Claude right now; the README already frames it as runtime-agnostic pending a configurable spawn command, so GPT, Gemini, or a local model would slot in without a rewrite. Whether the scores it produces actually predict anything about real-world CLI usability is still a single n-of-1 self-test, not a validated claim.