On Harnesses · Larder · Claude Code
Larder
We all wrestle with dinner. This lesson builds a small assistant for it, using one agent and a folder. The shape you learn here transfers to most of the other things you keep doing by hand.
Sunday morning, after you build it:
Five minutes. One plan. One shopping list. The rest of the lesson is how to get to that half-page on Sunday.
What you'll need
A short, honest list.
This lesson assumes you can create a folder and open a terminal window. That is it. You do not need to know how to code.
Everything below is free or low-cost. The only required piece of software isClaude Code, an agent you run from the terminal. You install it once and point it at a folder. That folder becomes the project.
The only other tool the lesson uses ispandoc, a free, open-source document converter. Claude Code will install it for you when the time comes.
Dedicated apps exist for this (Paprika, Mealime, Plan to Eat, and others). If you just want a meal planner, use one of those. The reason this lesson uses meal planning is that almost everyone has wrestled with it, and it is a good worked example: a small problem you can check the answers to. Some of it is judgment, some of it is bookkeeping, and a useful answer has to do both. The method transfers to anything with that shape.
- One sitting to build
- Five minutes a week to plan
- Whatever your Claude plan costs
- A weekly plan that respects the pantry and the household
- A shopping list grouped by aisle, ready to send to your phone
- A harness that gets sharper every week you use it
- A framework you install
- A product you log into
- A guide to writing perfect prompts
It's a small set of plain files in a folder, edited over time, that an agent reads at the start of every session. That's it.
On a Mac it's the app called Terminal. On Windows, search for Terminal or PowerShell. On Linux you already know. The window looks like this, give or take a color scheme:
You type commands after the prompt, Claude answers in the same window, and everything lives in files in that folder. Nothing else to learn about terminals for this lesson.
Visit theinstall guideand come back. Pick the plan that fits how much you expect to use it.
Two kinds of work
Both of them land on the same Sunday.
Planning a week of dinners is two kinds of work in the same task. The first is unstructured: reading recipe pages written for humans, sensing which dinner fits the night, noticing constraints a rule never captures. The second is structured: subtracting the pantry from the week's ingredients, scaling servings, grouping a list by aisle. Same half-hour on Sunday. Completely different shapes of thinking.
Until recently we did all of the synthesis ourselves, and we wrote or bought software for the bookkeeping. The translation between the two was always a human. What changed is that an LLM can now do the synthesis AND, when a script would actually help, write the script for you. The same tool handles both halves of the job.
- Extract ingredients from a recipe blog
- Judge whether a recipe fits a tired Tuesday
- Notice when a new plan contradicts an old preference
- Rewrite a step a child can follow
- Subtract pantry from the week's ingredient list
- Scale a recipe from four servings to six
- Group a shopping list by grocery aisle
- Fire a reminder at 10am on Sunday
An LLM is cheap at the unstructured work. A script is cheap at the structured work. You are the one who decides which constraints matter and when the system is wrong. A harness is the small set of files in a folder that gives each piece of a task to the tool suited to it and leaves the judgment with you.
Most of this lesson is the files that make the routing possible. Before we add any of them, though, a word on what the routing actually feels like. Nothing about working with an agent starts with a file. It starts with a conversation.
Your first session
An empty folder and a fresh agent.
You run claude in an empty folder. The agent reads the folder, finds nothing, and waits. Most people freeze here. We've been told a good prompt is a well-phrased paragraph, and the paragraph refuses to come. That is fine. A first session is a short conversation, not a prompt. The exchanges in this lesson are composites, assembled from real sessions to show one thing at a time.
You answered two questions. The second answer ("the folder") changed the shape of the project, and Claude adjusted. At the end, Claude offered to write the agreement down. Every file in this lesson comes out of a moment like that one.
Every claude call opens a session against the current folder. Claude reads what is there. Nothing from a previous session carries over except the files. That is why this conversation has to end in a file, not a memory.
The first file captures who the project is for and how the agent should behave. That's the next file.
Identity
The first file is not code. It is a briefing.
The "yes" at the end of your first session kicks off the first real file. Claude drafts CLAUDE.md, the identity file: a short markdown file at the top of the folder that captures the conventions you just agreed to. It's the first thing Claude reads at the start of every future session.
You review the draft, tweak a sentence or two, and that becomes the identity the agent reads forever after. You do not write this file from scratch yourself. You shape it by correcting the draft.
Identity tells the agent which judgments are in scope and which are not. "You are my meal-planning helper" keeps it from offering to build a full application. "Ask about dietary constraints once per session" keeps it from asking every turn. Limits like these are what let you trust what comes back.
When you're shaping CLAUDE.md, tell Claude what you want the project to feel like before you tell it what it should do. Tone shapes behavior more than you'd expect. "I want this to feel like my friend who is good at planning dinner" gives the agent a posture. "You are a meal-planning assistant" reads like a job title, and the answers come back stiffer.
Identity alone is not enough. The agent now knows who it is, but the fridge is still empty and the recipes are still scattered across the websites you bookmarked this year. The next thing the agent needs is context.
Context
What the agent sees when it starts.
Context is what the agent sees when a session starts. For this project, context is a handful of files on your disk. No app, no database. Just files.
Two files matter most at the start: a folder for recipes and a text file for what is in the fridge.
pantry.md is a plain markdown list that lives next to the recipes. Keeping it accurate is a real problem: groceries arrive, leftovers get eaten, things expire, and nobody updates the file. Meal-planning apps run into this. We are not going to solve it. We are going to let the agent help you keep it close to the truth by asking for an update at the start of every run, and make the memory file, preferences.md, catch what you forgot. For now, the file just exists.
Context is whatever the agent reads before it acts. Structuring that context well is the actual skill the lesson is teaching. Flat markdown files are not the only shape this can take, they are the shape chosen here because they make the structure visible. You can open the folder, see what the agent sees, and fix it in a text editor. The point is that the context is legible and shaped by you, not that it lives in one format or another.
Files are enough for a fridge. They are not enough for a workflow. The agent needs to know how to go from recipes plus pantry to a plan plus a shopping list. That is a skill.
Skills
A workflow written as a document, not as code.
A skill is a workflow written as a document. Not a script, not a config file. A short piece of prose that explains to the agent how to do one specific job, step by step, with moments where it should stop and ask you before continuing.
The skill lives at a specific path so Claude Code can find it: .claude/skills/plan-the-week/SKILL.md.
Code forces you to enumerate every case before it runs. A document written in English lets the LLM reason about the cases you did not name. That is why prose beats code here: you do not have to name every case. Your job shifts from "anticipate every branch" to "describe the shape of the problem and the judgment you want applied." The LLM handles the rest.
This is specifically what LLMs are good at that nothing else is. A plain-prose skill puts that capability to work instead of working around it.
When you ask Claude to draft a skill, walk it through one real example of the workflow first. Tell it what you would do on a specific Sunday, step by step, as if coaching a friend. Then ask it to generalize that into a skill. LLMs template from examples much better than from abstract descriptions. You get a tighter first draft if you give one.
A skill tells the agent how to work. It does not tell it what to care about. The next file is the agent's memory: what lands on the table, what does not, what we already decided last month.
Memory
Taste under version control.
After the first plan, you will notice things. The kids picked around the cauliflower. Saturday turned out to be a night nobody wanted to cook on, and the weeknights that worked were the ones with a single pan. These are facts about your household, not about recipes. They belong somewhere durable.
That place is preferences.md. A plain markdown file the agent reads at the start of every run and proposes updates to at the end of every week.
Memory here is a markdown file you can open and edit yourself. Keep it in the folder, edit it, copy it into your next project. The agent gets better because the file gets better.
So far, every piece of the harness has been a file full of prose. The next piece is different. It is an existing tool the agent calls on your behalf.
Hands
A tool the agent runs instead of working it out.
An LLM is very good at reading a recipe blog. It handles the story about the author's grandmother, the ad in the middle, the ingredient box, and the comments at the bottom, and it returns a clean ingredient list. Parsing messy prose is exactly what LLMs do well.
The catch is cost. Every session that re-reads the same blogs pays for those parsing tokens again. Recipes do not change week to week. Parsing the same page fifty times is waste. A hand is an existing tool that does the mechanical parse once and leaves a clean file the agent can read for free from then on.
You usually do not have to write this tool. There is a well-known one called pandoc that converts HTML (and most document formats) to clean markdown. You do not install it yourself either. You ask Claude, Claude runs the install with your approval, and the tool ends up scoped to this project so your system stays clean.
mise use pandoc@latest? it goes in this folder only, not on your system.curl -sL URL | pandoc -f html -t markdown > recipes/SLUG.md.| Per recipe (ingested once) | Per weekly read (50 recipes) | |
|---|---|---|
| Claude extracts from blog HTML | ~8K tokens | ~50K tokens (if cached locally) |
pandoc converts, then Claude reads | ~0 tokens (deterministic) | ~50K tokens |
pandoc is deterministic, so the parse never changes between runs. It is free, so you can re-ingest a recipe any time without worrying about tokens. And every token Claude doesn't spend on HTML parsing is a token it spends on the part you actually want its attention on: whether the week hangs together.
An LLM is cheap to ask, expensive to ask a hundred times. A hand is a piece of software that does a mechanical thing once and stores the result where the LLM can read it for free from then on, so the model looks it up instead of working it out. The LLM's real attention stays on the parts that vary (taste, constraints, judgment). A plain tool does the boring part ahead of time.
Most of the hands you want already exist. A few examples from the standard toolbox:
pandoc: document conversion (the one used here)jq: filter and transform JSONexiftool: read and edit photo metadataparallel: run many copies of a command at onceffmpeg: anything to do with audio and video
Adding a hand is not writing code. It is installing an existing tool and telling Claude inCLAUDE.md when to reach for it.
Before you approve an install, ask Claude what the tool does and why it's the right one. If the answer feels vague, that is a signal to dig in, not a signal to stop. "Can you show me what the output looks like on one recipe before we make this the default?" is a fair question and usually a productive one.
The agent reads files, writes files, and now runs a tool on your behalf. What it still needs is a moment to stop and ask. That is the next piece.
Gates
Where your judgment enters the loop.
A gate is a place where the agent stops and asks. The skill has two: after the weekly proposal (before it commits the plan) and before writing the next round of preferences (before it changes the memory file).
Gates exist because the agent moves faster than you can check it, so the skill stops and waits for you.
monday weeknight-pasta-primavera
tuesday sheet-pan-chicken-and-broccoli
wednesday red-lentil-soup
thursday tacos
friday takeout
saturday frozen-pizza
approve to continue, or tell me what to change.
Never let the shopping list write before the week is approved. A wrong plan is cheap to fix. A wrong shopping list is not.
A gate catches more than typos. It catches the first time Claude proposes something it shouldn't know is wrong. With an empty preferences.md the agent has nothing to check a proposal against, so the gate is the only filter.
Taste corrections are the easy ones. Gates also catch the two awkward failure modes every LLM has: proposing something that doesn't exist, and missing something that does.
LLMs don't announce when they're filling in a blank with something plausible. They sound confident either way. The gate's job is to be the place where "wait, is that real?" can happen. This is one of the most important habits you can build.
You will notice these three corrections went three different places. "No cauliflower" is a fact about the household, so it goes in preferences.md. "Don't propose recipes that don't exist" is a rule about how the agent works, so it goes in CLAUDE.md. "Ask before assuming ambiguous pantry state" is a new instruction to pause, so it goes inCLAUDE.md too. A useful rule of thumb: corrections about what you like go in memory, corrections abouthow the agent should behave go in identity.
The fourth gate type is the one readers rarely see written down. Sometimes Claude stops and says it doesn't know.
When Claude says it doesn't know, it is asking to be taught. The fix isn't always to answer the immediate question. Often it's to add a rule that turns the question into a pattern: here's what we do when we hit this kind of thing.
The whole system is here now. Files, skills, memory, hands, gates. How do you actually start it on a Sunday?
Invocation
One command, many steps.
You open the folder in Claude Code on Sunday morning. You type one thing.
The command is short because the harness carries the context. The agent does not need you to re-explain the project, the family, the rituals, the pantry, or the recipes. It reads the files. It has been reading them since session one.
When you run claude in the folder, a session opens and CLAUDE.md is read automatically. Other files (the skill, pantry, preferences, recipes) are read when the skill tells the agent to open them. When a session gets long, or the conversation drifts somewhere you didn't intend, close it and start a new one. The files carry you across. A fresh session with a good harness is almost always sharper than a long session fighting its own history.
The week goes well. Some weeks do not. That is where the most important habit lives.
The loop
The files get better every week you use them.
A week where the plan mostly held, and one dinner ran long enough that you gave up and ordered pizza. That small specific failure is the thing worth writing down.
A harness is not a tool you install once and then use. It is a garden you keep. Every week adds a line topreferences.md. Every new dinner you like ends up as a recipe. Every awkward night becomes a rule that keeps the next Thursday from going the same way. The work of tending looks like this.
Every file the skill opens is read on every run. Every sentence you add to preferences.md, every recipe you let pandoc convert, every tweak to the skill becomes free context for the next invocation. You pay the thinking cost once. The agent gets the benefit on every run after.
That asymmetry is the compounding. One good retrospective improves every week after it. One new sentence in the skill shifts the tone of every proposal. The more you tend, the less the agent has to guess.
The loop is about memory. But memory is not the only file that changes over time. The skill changes. The identity changes. Here is what that looks like.
Evolving the harness
The files move too, not just the memory.
The retrospective in the last chapter made it sound like only preferences.md changes. In practice, every file in the harness gets edited as you learn what the current one is missing. Three edits, and what pushed on them.
A rule the gate found. Chapter 7 ended with a correction: never propose a recipe that isn't already a file in recipes/. That is a rule about how the agent should behave, not a taste fact, so it belongs in CLAUDE.md and not in preferences. Here is the edit.
A question the skill never asked. The plan never asked what the week ahead looked like, so it could not tell a quiet Tuesday from one you were already out of time on. That isn't a taste problem. The skill is missing a question.
A file that got long. preferences.md grows until lines start getting missed. You ask Claude to reorganize the file into sections (allergies, weeknight rules, rituals, what lands, what doesn't). Nothing in the content changes. Structure makes it easier to see what the agent is working from.
None of these edits were planned when you set up the project. Each one came out of a gate moment, a retrospective, or a week that felt off. The files got edited because something specific pushed on them. This is the actual rhythm: use it, notice what's missing, update the file that should have caught it, repeat.
Tone and judgment corrections →preferences.md. Rules about what the agent should and shouldn't do → CLAUDE.md. Changes to the workflow itself (phases, gates, questions) →SKILL.md. If you're not sure, ask Claude which file it thinks a correction belongs in. Its answer is usually right, and when it's wrong, the disagreement teaches you both something.
Now look at what it feels like once this has been running a while.
Compound
What the files look like once you have used them a while.
Keep using it and preferences.md gets longer, recipes/ gets fuller, and the skill picks up a phase or two. A Sunday plan takes about five minutes.
None of the files change shape. The harness you built at the start is the same harness running now. The files got fuller. That was the point.
Four markdown files, one folder of recipes, and one existing CLI the agent runs for you. No application. The harness gives each kind of work to the thing that does it best, and you get Sunday afternoons back.
The chapters named several pieces (identity, context, skills, memory, hands, gates). They all collapse into two things.
You write markdown to shape the agent's context.CLAUDE.md, the skill, preferences, and the pantry are all plain text you can read and edit.You let the agent call tools.Sometimes that's a CLI like pandoc. Sometimes it's the file system itself.
Every other problem that used to look like "I need an app" starts looking like "I need a folder with the right files and a tool or two the agent can reach." That's what this lesson was actually teaching.
The harness got sharper as you used it. You did too. Describing a workflow so Claude can run it is something you had to learn. Reading Claude's draft and spotting a missing phase is something you had to learn. Knowing which file a correction goes in is something you had to learn. None of that is meal-planning knowledge. It transfers.
Three cases where reaching for this harness is the wrong move. First, when the problem is fully deterministic (tax math, date arithmetic, currency conversion): a five-line script is cheaper, faster, and never hallucinates. Second, when the output has to be correct in ways you can't verify at a glance (a medication schedule, a legal filing): the gate can't catch what you can't read. Third, when the state changes faster than the memory file can track (a moving project with a dozen people editing in real time): the harness is built for slow, considered compounding, not for live coordination.
Knowing where the harness doesn't belong is part of knowing where it does.
- Clone
lesson-larder. - Skim the files in this order:
CLAUDE.md→.claude/skills/plan-the-week/SKILL.md→pantry.md→preferences.md. Reading them in order shows why they're shaped the way they are. - Pick something you actually want done this week that isn't meal planning. Something small.
- Rename the files to fit, rewrite the words, keep the structure. Run it. Let the gate catch the first thing that's wrong. Write that correction down.
The next time you catch yourself saying "I wish there were an app for this," try a folder instead.