synthesist
Synthesist is a specification graph manager for AI-augmented development: a CLI that records the process governing a body of code (what was planned, what was agreed, what phase the work is in) as a graph of specs and tasks, separate from the code itself. It starts from the observation that technically correct, agent-authored pull requests still get rejected when they violate scope, architecture, or process expectations the agent never had visibility into. Synthesist’s answer is to make that missing context a queryable, first-class thing instead of something an agent has to infer from a diff. The name borrows the crew role aboard the Theseus in Peter Watts’ Blindsight, whose job was not expertise but coherence.
Every piece of workflow state is a claim: a typed, timestamped assertion appended to a per-asserter JSON-LD log under claims/, read back through a disposable redb-backed index called gamma. Nothing is overwritten. A field update is a new claim that supersedes the prior one, so history survives per field, and concurrent sessions merge by taking the union of every asserter’s log, with no merge step to run. It’s a Rust workspace: the claim crate is the vocabulary-agnostic substrate, synthesist on top defines the workflow vocabulary and the CLI.
Every fact lives in the per-asserter logs, so gamma can be thrown away and rebuilt from them, and two sessions writing at once converge by union rather than by a merge step.
Highlights
- The v3 claim substrate measured 378x smaller than v2’s Automerge store on a real single-asserter corpus: 143 claims that cost 34 MB as
.amcchange files cost 92 KB as JSON-LD, a gap proposal 002 traces to Automerge’s per-change vector-clock envelope, not payload size. - The workflow state machine (ORIENT, PLAN, AGREE, EXECUTE/REFLECT, REPLAN, REPORT) is enforced algorithmically:
phase setrejects invalid transitions, and AGREE forbids every write, so a plan snapshot has to be pinned before the agent presents it for approval. - A path-traversal regression suite closes three untrusted-input vectors into the asserter-to-directory mapping (
--session,$USER, and an import payload’sprov:wasAttributedTo). Two are rejected outright; the third is neutralized by collapsing unsafe characters rather than silently dropped. - The v2-to-v3 migration shipped broken in rc.1 (issue #11: a compacted production estate was silently misreported as fresh), then was rebuilt as real infrastructure and validated end to end against a live 9,530-claim estate with zero skipped claims and zero dangling supersedes.
agent-shape.tomldrives jig’s runtime-in-the-loop battery to score whether the CLI surface is shaped so agents reach for real commands. The v5.1.0 baseline (50 trials across two models) completed 96% of tasks but invented 80 non-existent commands and fell back to raw SQL 7 times.- A tuning pass meant to improve that shape instead surfaced a measurement problem: rejudging the identical before/after transcripts moved the aggregate score delta from -0.13 to +0.00, while individual task/model cells still swung by up to 0.4. That reads as single-judge, n=5 noise the study doesn’t yet trust.
- CI runs a fail-closed automerge gate requiring OSV, cargo, and secret-scan verdicts all present and passing. The secret verdict replaces GitLab’s upstream template because that template’s gitleaks invocation exits 0 even when it finds a secret, so it could never actually fail a pipeline.
- MIT licensed, with signed release binaries for macOS ARM64 and Linux amd64/arm64.
The observation layer this used to carry (stakeholders, dispositions, signals) now belongs to a separate, private companion tool; synthesist kept the workflow graph. The origin story and the disposition-graph idea that came before the claim substrate are in Get Lamp. Whether the agent-shape numbers hold up under more trials and more than one judge is still open.