I stopped negotiating with Claude

August 30, 2026

I'm halfway through a feature and my context window runs out.

What I used to do was ask the dying session to write me a prompt for the next one. "Sum up where we left off so I can paste it tomorrow."

It works. Sort of.

The problem is that summary gets written by the same session that's about to die, with no guidance about what to include. And with no guidance, it includes something different every time. It remembers the last file we touched and forgets the decision we made three hours earlier, the one we'd been sticking to all session and that never made it into writing anywhere.

The next day I'd paste the prompt and the negotiation would start. Claude would ask about something we'd already settled. I'd clarify. It'd build something close to what I wanted. I'd correct it. It'd read three files it didn't need to read. Three or four rounds before we actually got to work.

And I didn't always notice what had gone missing. Sometimes it hit me two hours later, when something had been built differently than what we'd agreed on the day before.

The two commands

I split it at both ends.

/handoff is the close. Instead of asking for a free-form summary, I tell it exactly what has to be written down and in which files. Since I've had it I can stop halfway through a spec without worrying: I run /handoff and that's it. If the feature is ready for a PR, I run /handoff and /pr.

/cold-start is the open, and inside it's just a short list. What to read and in what order: repo context, the spec with which features are done and which one is in flight, the two sections of PROGRESS.md that matter, git state, and what evidence went stale.

The part that pays off most is the other half: telling it what not to read. When PROGRESS.md gets long, the old stuff rotates out into a separate file. That file is there in case I ever need it, not to be read on every start: if Claude opens it, it eats half my context window before I write a line. And without an explicit instruction, it opens it. Exploring makes sense when you don't know where the information lives, and here I do.

The output format lives in there too, so the answer just shows up: five lines with where it stopped, what the exact next action is, and what evidence went stale. Then it closes by proposing where to start.

Where it stopped: F17 merged (PR #88), phase 3 is at 19/28 passing, nothing in
flight, clean tree. Branch feat/serie-a-serie-f18 already exists with 0 commits
of its own.

Exact next action: F18 is test_first, so it starts with the red, not the code.
Add the "two out of four sets logged doesn't count the exercise as logged" case
to lib/home.test.ts and run vitest expecting EXIT 1.

Stale evidence: 119 rows, none from F17 or F18. They get re-run, not re-hashed.

Should I start F18 with the red?

Look at the second line. It reminds me that this feature starts with the failing test, not with the implementation. I decided that once, calmly, and now I can't skip it even when I'm in a hurry and itching to jump straight into the code.

Why it works

Saving tokens isn't about prompt length, it's about precision. And precision is exactly what commands give you, along with consistency.

An imprecise prompt opens a conversation. Claude asks, I clarify, it builds something close, I correct. Every one of those rounds drags the entire prior context along with it: you're not paying for the prompt, you're paying for the whole conversation again, on every turn. Three rounds of back-and-forth cost more than the longest instruction you could possibly write.

So where does the precision come from? From three things the command has and the free-form summary doesn't.

It names the sources. I don't ask it to remember where we left off. I tell it which files to pull the state from and in what order. It doesn't rebuild the state from memory, it reads it.

It says what stays out. Without that Claude explores, and exploring costs context window.

It pins the output. I tell it exactly what format I want. Without that, every start hands me a different summary and I end up asking for whatever was missing.

The summary my old session left me had none of this. It decided on its own what mattered, and it decided at the worst possible moment, with the context window already full.

A command is still a prompt, it's just that nobody improvises it halfway through a feature. Since it's precise, there's nothing to negotiate: the first turn is already work.

How I spot something that could be a command

After these two I started seeing the pattern elsewhere.

The simplest signal: if you're writing the same prompt for the third time, it's already a command you haven't written yet.

Then I added two questions.

Does the right answer depend on my judgment or on the state of the project? "Should I split this module in two?" is judgment, and I'm never going to pin that down. "What went stale in PROGRESS.md?" is state, and it has exactly one right answer.

Can I list out the steps? If I sit down and write what has to happen, in order, it's already solved. If I get stuck listing them, I haven't understood what I'm asking for yet, and turning it into a command just freezes my own confusion.

The limit

Commands are for how the harness works, not for what gets built.

Project decisions stay out. Whether this feature ships, how we model this, whether it's worth refactoring now. That's where I want judgment, and judgment doesn't fit in a list of steps.

What does fit is the part that repeats the same way every time around: closing, opening, keeping the files current, checking what went stale. That's not the work, it's the ceremony around the work. And the ceremony is worth having written down.