Skip to content
andrew.dunn.dev

Harness Explorer

Source

There’s a growing ecosystem around evaluating what LLMs produce. Promptfoo tests model outputs against assertions. Langfuse traces runtime behavior. DSPy optimizes prompts programmatically. But as far as I can tell, nobody does static analysis of the instruction files themselves: the AGENTS.md, CLAUDE.md, and skills files that increasingly define how agents behave in a codebase. he is an attempt to find out whether that’s possible and whether it tells you anything useful.

INPUTharness filesAGENTS.md, skillsGRAPHknowledge graphrefs followedSCORINGdetectorsoverlap, conflictOUTPUTranked fixeshe recommend
Harness Explorer reads the instruction files an agent obeys, follows their references into one graph, scores that graph, and ranks the edits worth making.

The harness itself grew out of building Immutable Base. The self inventory covers the broader shift in what skills matter when implementation commoditizes, and the AI Harnesses Are For Everyone essay is the longer treatment of why harnesses matter and how to build one. he is the attempt to measure one instead of just building it.

The current hypothesis is that the most useful output is a ranked list of leverage points: where should I invest my next hour of harness work? he builds a knowledge graph out of the instruction files and scores it for redundancy, contradiction, and information density, then he recommend ranks the highest-leverage fixes by estimated token savings; he lint runs the same detectors with CI-friendly exit codes. It ships as a single Go binary, model and dashboard embedded via go:embed.

reference80% redundancycontradictionKNOWLEDGE GRAPH16 files, 24,562 tokensROOT FILEAGENTS.md1,036 tokensDOCframework.md1,726 tokensDOCinstance.md1,326 tokensSKILLvoice.md2,841 tokensSKILLdesign-system1,891 tokensSKILLaccessibility983 tokensSKILLdesign-critique983 tokens
Following references across the 16 instruction files behind this site, he finds an 80 percent overlap between two docs and a contradiction between two skills.

Running he against this site’s own harness (16 instruction files, about 24,500 tokens) produces the dashboard below.

Harness Explorer dashboard showing metrics, quality signals, and ranked recommendations for the andrew.dunn.dev harness

OpenAI published how they built a product with zero manually-written code, and their fix for the same failure mode, a thin root file pointing into a structured docs/ tree, is the same progressive disclosure shape he scores for. They also run agents that open fix-up PRs automatically; he stops at detection, in keeping with the case for single-purpose, composable tools made elsewhere on this site.

No calibration dataset exists for instruction file quality: nobody knows what a good redundancy score looks like, or whether static deontic strength predicts behavioral compliance. The theory is that tokens spent on redundant rules, contradictions between files, and low-information prose are invisible costs, and that surfacing them beats not surfacing them; that might be wrong, and it’s testable, just not yet tested. The contradiction detector’s thresholds (polarity at 0.6, scope at 0.7) were tuned against my own harness files until false positives disappeared and real contradictions still triggered. That’s one data point, not a validation study.