#testing
3 entries
projects · April 2026
jig
An agent-shape harness that points a real agent at a real CLI and counts what happens: a battery declared in agent-shape.toml, a headless runtime spawned against the fixture, and an LLM judge scoring the transcript on a 4-point scale whose top mark is a first command that was real and worked. A companion check guards the rubric against drift in the tool's own help output.

writing · April 2026
The Hunt for Leverage
On test frameworks that tell agents what to do next, and the design patterns that make LLM-generated tests actually correct.

writing · March 2026
On Building a Complicated Foot Gun
Deploying a home server with immutable infrastructure tooling: what the pipeline tested, what it didn't, and what broke anyway.