kit
kit resolved developer tool versions across multiple git-based registries, generated mise configuration from the result, and verified what it installed. A tool definition was a small TOML file: name, source (GitHub, GitLab, npm, crates.io, a direct URL, rustup, or Homebrew), version, per-platform asset names, a checksum location, and a signature method where one existed. A registry was a git repo of those files under tools/. Multiple registries could be configured on one machine at once, and when two defined the same tool, declaration order picked the winner; the loser was logged as shadowed, not silently dropped.
Every tool also carried a trust tier, own, high, or low, and the tier picked how hard a bump had to work to land. own (tools I built and published myself) required cosign keyless verification with an anchored certificate identity. high accepted a GitHub attestation or a checksum. low accepted a checksum alone. A registry’s _meta.toml decided which tiers, and which size of version bump, could auto-merge without a human looking at it first. None of that was inferred from a tool’s popularity or source; tier was a judgment call recorded once, in a TOML field, and kit enforced whatever it said.
The automation side was a three-pipeline CI architecture that tracked upstream releases without anyone polling GitHub tags by hand, and it is where most of the interesting bugs turned up.
Highlights
- Three pipelines, in sequence: sense (always exits 0, only classifies risk by bump size, tier, and checksum status), evaluate (deterministic rules approve clean patches and reject checksum mismatches outright; an LLM only sees edge cases: missing checksums, major bumps, download failures, and never sees release notes), apply (partitions updates into an
auto_merge_groupand areview_groupper registry policy and opens a separate MR for each), andverify-registryas the hard merge gate re-checking every checksum on the MR. - A
contains()match on release asset names once matchedkit-darwin-arm64.bundleahead ofkit-darwin-arm64itself, because the bundle’s link happened to appear first in GitLab’s JSON array. kit downloaded and hashed the cosign bundle instead of the binary, producing a spurious checksum mismatch on everyown-tier release. The fix matched the exact asset name first and fell back to a URL-suffix match, explicitly excluding.bundle,.sha256, and.sbomvariants. - Verifying npm-sourced tools meant fetching a package’s
dist.integritySRI hash and hashing the tarball; verifying yq meant a parser for its own multi-hash checksums file, which lists one column per algorithm alongside a separate index file naming the column order. That work shipped in 0.14.0 and broke the very next auto-merge MR:ToolDef::validatedid not recognize the new"npm"checksum key and rejected every npm tool definition it had just correctly verified. 0.14.1 fixed the validator the next day. - A checksum mismatch on my own machine, unreproducible on a clean VM running the same container image, produced a day of retry logic: re-download both the binary and the checksum file, verify content length, log diagnostics for a CDN edge-caching theory. All of it got reverted the next day back to “download, hash, compare, report,” on the reasoning that retry belongs in the CI job’s
retry:config, not in application code that would otherwise mask a real signal. - A jig tuning battery (n=50) scored an LLM operating kit’s own CLI: 0.430 mean, 94% task completion, but a wide per-task spread.
estate-overviewscored near 1.0 for opus;pin-version,release-pin, andregistry-precedenceall sat around 0.2 to 0.25 for both opus and sonnet. Help text at startup and akit skillsubcommand did not guarantee an agent reached for the right subcommand. - The
brewsource backend delegated install to mise’s ownbrew:scheme, so CI never needed a local Homebrew install; kit tracked new formula versions through the formulae.brew.sh API and emittedbrew:<formula>@<version>into the generated mise config. - A CI job failed any
.rsfile over 500 lines; ten pre-existing offenders were named in.file-size-waiverrather than split under deadline pressure, and any new file that crossed the line failed outright. - Published to crates.io as
nomograph-kit; MIT licensed.
The tier model generalized past kit itself: verification strength is a per-artifact judgment call worth stating explicitly, not one to infer from a tool’s popularity. The agent-shape numbers were a separate finding, that a CLI’s discoverability to an LLM operator does not track with how well it is documented for a human, and that it is worth measuring on its own rather than assumed.
Building a verification layer rather than trusting an install path directly is the same shape immutable applies one level down, at the operating system: pin it, verify it, track what moved.