← Back to Tutorials
TutorialFor: Software Engineers, AI Engineers, ML Engineers, Platform Engineers

How to Reduce Claude Code Token Usage: A Measured Setup

Shrink what Claude Code re-reads on every turn, stop tool output from flooding the context, and prove the saving with real numbers.

#tutorial#intermediate#claude-code#token-optimization#prompt-caching#cost-optimization#hooks

What you'll build: a Claude Code setup that costs less per task

This tutorial shows you how to reduce Claude Code token usage with a measured, twelve-step project setup. Here is the proof first: the same bug fix, run headless by Claude Code on the same small repository, before and after the setup in this tutorial:

text
--- before: cold cacheturns=7  cost=USD 0.1256input processed=168,321 (cache write 21,554 + cache read 146,757 + uncached 10)input per turn=24,045  output=1,004tests: PASS--- after: cold cacheturns=6  cost=USD 0.0552input processed=96,296 (cache write 7,386 + cache read 88,902 + uncached 8)input per turn=16,049  output=786tests: PASS

Both runs fixed the bug. The second read 43% fewer input tokens and cost 56% less. One pair of runs is noisy, so later you will also see averages over repeated runs, and they tell the same story with a smaller gap.

By the end of this tutorial you will have a token-lean Claude Code project setup and a script that measures what any task costs. The setup has five parts: a short CLAUDE.md, path-scoped rules, Read deny rules for build output, a hook that shrinks test output, and a default effort level. You will also know which popular token-saving tips stopped working, and what replaced them.

This is an intermediate tutorial. It assumes you use Claude Code every day and know what CLAUDE.md, slash commands, and hooks are. It does not assume you know how the context window, the prompt cache, or tool output limits work. Allow about 60 minutes, plus roughly USD 0.50 of model usage for the measurement runs.

Verified against Claude Code 2.1.284, Node.js 20.20.0, npm 10.8.2, Python 3.13.14, and Git 2.51.0.windows.1 (Git Bash on Windows 11) on 2026-09-29, with Sonnet 5.5 as the model. Every claude command in this tutorial needs a logged-in account, so those blocks were not machine-verified: the /context checks, the measure.sh runs, the deny and effort checks, the live hook call in Step 8, the one-time trust dialog in Step 1, and the Step 12 session commands. I ran them by hand, and every transcript and number shown for them is a real capture from those runs. The Claude Code installer was not run either, because it changes the machine. All other blocks were executed or structurally checked on a clean rebuild.

Prerequisites

Level: intermediate. You should be comfortable in a terminal, and able to read JSON and short Python and Bash scripts.

RequirementVersion usedWhy
Claude Code2.1.284The status line's cache fields need 2.1.251 or later, and the hook and settings features used here are current in 2.1.284
Node.js and npm20.20.0 and 10.8.2The sample project runs its tests with the built-in node:test runner
Python3.13.14, callable as python3Every script and hook in this tutorial is standard-library Python
Git2.51.0The measurement script uses Git to put the bug back after each run. On Windows, Git Bash also runs the hooks
A Claude accountPro, Max, Team, or an API keyEvery claude -p run uses your plan's usage limits, or is billed on an API key. total_cost_usd is a list-price estimate either way. The whole tutorial used about USD 0.50 at list price

Install Claude Code with the native installer if you do not have it yet. It does not need Node.js:

bash
curl -fsSL https://claude.ai/install.sh | bash

On Windows, run irm https://claude.ai/install.ps1 | iex in PowerShell instead. Then confirm the whole toolchain in one command:

bash
claude --version && node --version && npm --version && python3 --version && git --version

Expected output (your patch versions can be newer):

text
2.1.284 (Claude Code)v20.20.010.8.2Python 3.13.14git version 2.51.0.windows.1

On Windows, run every command in this tutorial from Git Bash, not PowerShell. Two details matter there, and the troubleshooting section shows the exact errors they cause: set MSYS_NO_PATHCONV=1 before any claude -p "/..." command, and make sure python3 runs a real Python, not the Microsoft Store shortcut. The version check above is the test: if its python3 --version line is missing or prints a message about the Microsoft Store instead of Python 3.x, fix that first: install Python 3.13 from python.org, turn off the python.exe and python3.exe entries under Settings > Apps > Advanced app settings > App execution aliases, and open a new Git Bash window. Do this before Step 1, because every hook in this tutorial calls python3 from a non-interactive shell where a Git Bash alias does not apply.

How Claude Code spends tokens: context, cache, and tool output

Read this section before you change anything. Every later step pulls one of the levers described here, and the measurements only make sense once you know what they count.

Every turn re-sends the whole conversation

A Claude Code session is a loop. You send a prompt, and the model answers with either text or a tool call. Claude Code runs the tool and sends the result back, and the model answers again. Each of those round trips is one API request, and each request carries the entire context again: the system prompt, the tool definitions, your CLAUDE.md files, and every message and tool result so far.

d2
direction: right
you: "You: one prompt" {
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}
turn1: "Request 1\nprefix + prompt" {
  style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
}
turn2: "Request 2\nprefix + prompt\n+ tool result 1" {
  style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
}
turn3: "Request 3\nprefix + prompt\n+ tool results 1-2" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}
done: "Answer" {
  style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
}
you -> turn1 -> turn2 -> turn3 -> done

So what you pay for a task is roughly the size of the context multiplied by the number of requests. You can shrink either one. This tutorial shrinks both.

What sits in the context before you type anything

The fixed part of every request is called the prefix. /context shows it, and here is the real breakdown for the sample repository you will build in Step 1, before any cleanup:

d2
grid-rows: 1
grid-gap: 0
sys: "System\n2.4k" {
  width: 90
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}
tools: "Built-in tools\n12.9k" {
  width: 320
  style: {fill: "#7B68EE"; stroke: "#5A48C8"; font-color: "#FFFFFF"}
}
mcp: "MCP\n1.4k" {
  width: 40
  style: {fill: "#955D37"; stroke: "#6E4225"; font-color: "#FFFFFF"}
}
mem: "CLAUDE.md\n7k" {
  width: 175
  style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}
}
skills: "Skills\n2.1k" {
  width: 55
  style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
}

Most of that prefix belongs to Claude Code itself, and you cannot remove it. The red block is yours. In this sample, CLAUDE.md is 27% of the fixed context on every single request.

Two rows in real /context output are marked (deferred). Those are tool definitions that are not in the context: Claude Code loads only their names and fetches the full definition when a tool is needed. That is why removing MCP servers saves much less than it used to, as the table of outdated advice below shows.

The prompt cache decides what a token costs

Re-sending the prefix on every request would be very expensive without the prompt cache. The API stores the start of each request, and the next request that starts with exactly the same bytes reads that part from the cache at a steep discount. The cache is a prefix match. The first byte that differs ends the cached part, and everything after it is processed again at full price.

d2
direction: right
l1: "Layer 1\nsystem prompt + tools" {
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}
l2: "Layer 2\nCLAUDE.md + rules without paths:" {
  style: {fill: "#7B68EE"; stroke: "#5A48C8"; font-color: "#FFFFFF"}
}
l3: "Layer 3\nconversation so far" {
  style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
}
new: "New turn\n(always full price)" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}
l1 -> l2 -> l3 -> new
breaks1: "Invalidates from layer 1:\nswitching model, adding an MCP server\nwith tool search off, fast mode on" {
  style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}
}
breaks3: "Rewrites layer 3:\n/compact" {
  style: {fill: "#FFA07A"; stroke: "#D9774F"; font-color: "#2C2C2A"}
}
breaks1 -> l1: {style.stroke: "#B03A2E"}
breaks3 -> l3: {style.stroke: "#D9774F"}

Two terms in that diagram: tool search is the deferred loading of tool definitions described above, and it is on by default; fast mode is an optional faster, more expensive output mode you turn on with /fast. If you have not changed either, the red box only applies when you switch models: with tool search on, adding an MCP server does not break the cache.

Prices for Sonnet 5.5, the model used in this tutorial, per million tokens:

Token typePriceRelative to input
Uncached inputUSD 2.001x
Cache write, 5-minute lifetimeUSD 2.501.25x
Cache write, 1-hour lifetimeUSD 4.002x
Cache readUSD 0.200.1x
OutputUSD 10.005x

The cache-read discount depends on the model. It is 0.1x on most models, but 0.05x on Opus 5.5 and 0.025x on Fable 5.1, so check the pricing page for the model you use. You can check the formula against the "before" run above. That run was on a subscription, whose main conversation uses the 1-hour cache (next subsection), so writes cost USD 4.00. On an API key they would cost USD 2.50. 21,554 written at USD 4.00, plus 146,757 read at USD 0.20, plus 1,004 output tokens at USD 10.00, is USD 0.1256. The cache write is 69% of that bill. Every token you trim from the prefix is one fewer token in that write, and one fewer in every read after it.

The cache has a lifetime, and it is not always one hour

A cached prefix expires if nothing reads it for a while. The lifetime depends on how you pay:

d2
direction: down
sub: "Subscription (Pro, Max, Team)" {
  style: {fill: "#FFFFFF"; stroke: "#2C6FB0"; font-color: "#2C2C2A"}
  main: "Main conversation\n1-hour cache" {
    style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
  }
  other: "Subagents, compaction, titles\n5-minute cache" {
    style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
  }
}
api: "API key, usage credits, cloud providers" {
  style: {fill: "#FFFFFF"; stroke: "#B03A2E"; font-color: "#2C2C2A"}
  all: "Everything\n5-minute cache" {
    style: {fill: "#FFA07A"; stroke: "#D9774F"; font-color: "#2C2C2A"}
  }
}

On an API key, a coffee break longer than five minutes means the next prompt writes the whole prefix again. API users can set "promptCacheTtl": "1h" in their user settings, ~/.claude/settings.json. If you often pause for more than five minutes, do it: a 1-hour write costs 2x input but pays off after two reads.

Claude Code already caps tool output, but not at zero

A single tool result can be larger than your whole CLAUDE.md. Claude Code has built-in limits:

d2
direction: right
bash: "Bash tool output" {
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}
ok: "Command succeeded" {
  style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
}
fail: "Command failed\n(non-zero exit)" {
  style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
}
inline30: "Inline up to ~30,000 chars,\nthen saved to a file\n+ 2,000-char preview" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}
inline10: "Inline up to ~10,000 chars,\nhead and tail excerpt" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}
bash -> ok -> inline30
bash -> fail -> inline10

Those caps stop the worst floods. But 10,000 characters of repeated stack traces are still about 3,000 tokens that stay in the conversation and get re-read on every later turn. Step 7 and Step 8 bring that down to a few hundred.

Token-saving advice that no longer works

Claude Code changes almost every week, and a lot of popular token advice has not kept up. Drop these claims. Each one was checked against the current docs and the changelog:

Old adviceWhat is true in 2.1.284Do this instead
Add a .claudeignore fileIt was never a Claude Code feature. The file is silently ignored (issue #92816)Read(...) rules in permissions.deny (Step 6)
Say "think", "think hard", "think harder" to control the thinking budgetPlain text since the effort system arrived. Only ultrathink is recognized, and it adds an instruction without changing the effort sent to the API/effort or effortLevel (Step 10)
Cap cost with MAX_THINKING_TOKENSIgnored on Opus 5.5, Sonnet 5.5, and Fable, which set their own thinking budgetLower the effort level
Remove MCP servers to reclaim contextTool definitions are deferred by default. Only names load at startStill prefer a CLI (gh, aws) over an MCP server, but expect a small gain
Split CLAUDE.md with @path imports to save tokensImported files load at launch, in fullPath-scoped rules, or plain files Claude reads on demand (Step 4)
respectGitignore keeps Claude out of ignored filesIt only filters the @ file pickerRead(...) deny rules
Stay under 200K tokens because long context costs doubleClaude 4.6 and later models bill 1M context at standard ratesKeep context small anyway: every token is re-read every turn
A PostToolUse hook can trim any tool outputTrue for successful calls since 2.1.121, but a failed command fires PostToolUseFailure, which cannot replace outputRewrite the command before it runs (Step 8)
Use # to add a memory quicklyRemoved in 2.0.70Ask Claude to edit CLAUDE.md, or run /memory
/cost is the cost commandAn alias of /usage since 2.1.118/usage
opusplan is a free cost winEach plan-mode toggle switches model, which starts a new cacheUse it only for long planning phases
d2
direction: right
old: "Tips from 2025" {
  style: {fill: "#FFFFFF"; stroke: "#B03A2E"; font-color: "#2C2C2A"}
  a: ".claudeignore" {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
  b: "think harder" {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
  c: "MAX_THINKING_TOKENS" {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
  d: "@imports" {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
}
new: "What works in 2.1.284" {
  style: {fill: "#FFFFFF"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
  a: "permissions.deny Read()" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
  b: "/effort, effortLevel" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
  c: "lower effort" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
  d: ".claude/rules with paths:" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
}
old.a -> new.a
old.b -> new.b
old.c -> new.c
old.d -> new.d

Step 1: Create a sample repo with the usual token sinks

Goal: generate a small Node.js project that wastes tokens the way real projects do.

Why this step: you need a repository where you can change one thing at a time and measure the effect, and your own repository has too many moving parts for a first pass. This one has a real bug, a noisy test suite, a large committed build file, and a CLAUDE.md that grew into a wiki: 336 lines of API reference, deploy runbook, and changelog. The script also writes a .gitattributes file that keeps line endings as LF, so Git on Windows does not print conversion warnings.

Create a working folder anywhere, and save this as make_demo.py inside it:

python
"""Generate cartsvc: a small Node project with the usual token sinks."""import pathlibimport randomimport sysroot = pathlib.Path(sys.argv[1] if len(sys.argv) > 1 else "cartsvc")random.seed(7)def write(rel, text):    path = root / rel    path.parent.mkdir(parents=True, exist_ok=True)    path.write_text(text, encoding="utf-8", newline="\n")# Keep LF endings everywhere, so Git on Windows does not warn or convert.write(".gitattributes", "* text=auto eol=lf\n")write("package.json", """{  "name": "cartsvc",  "version": "1.4.0",  "type": "module",  "scripts": { "test": "node --test --test-reporter=spec test/" }}""")# The bug: a percent discount is subtracted as a flat amount.write("src/pricing.js", """export function subtotal(items) {  return items.reduce((sum, i) => sum + i.price * i.qty, 0);}export function applyDiscount(amount, discount) {  if (discount.kind === "flat") return Math.max(0, amount - discount.value);  if (discount.kind === "percent") return Math.max(0, amount - discount.value);  throw new Error(`unknown discount kind: ${discount.kind}`);}export function tax(amount, rate) {  return Math.round(amount * rate * 100) / 100;}""")# 240 table-driven tests; the 40 percent-discount cases fail.cases = []for n in range(240):    price, qty = random.randint(5, 400), random.randint(1, 9)    kind = "percent" if n % 6 == 0 else "flat"    value = random.choice([5, 10, 15, 20, 25])    base = price * qty    want = base * (1 - value / 100) if kind == "percent" else max(0, base - value)    cases.append(f"  [{price}, {qty}, '{kind}', {value}, {round(want, 2)}],")write("test/pricing.test.js", "\n".join([    "import { test } from 'node:test';",    "import assert from 'node:assert/strict';",    "import { subtotal, applyDiscount } from '../src/pricing.js';",    "",    "const cases = [",    *cases,    "];",    "",    "for (const [price, qty, kind, value, want] of cases) {",    "  test(`${kind} ${value} on ${qty} x ${price}`, () => {",    "    const base = subtotal([{ price, qty }]);",    "    const got = applyDiscount(base, { kind, value });",    "    assert.equal(Math.round(got * 100) / 100, want);",    "  });",    "}",    "",]))# A committed minified bundle: every grep for a function name lands here too.chunk = ('function applyDiscount(a,d){if(d.kind==="flat")'         'return Math.max(0,a-d.value);if(d.kind==="percent")'         'return Math.max(0,a-d.value)}function subtotal(i){'         'return i.reduce((s,x)=>s+x.price*x.qty,0)}')write("dist/cartsvc.min.js", ";".join([chunk] * 1500) + "\n")# A CLAUDE.md that grew the way they do: runbooks, API docs, history.lines = ["# cartsvc", "", "Pricing service for the checkout flow.", "",         "## Commands", "", "- Test: `npm test`", "- Node 20+", ""]lines += ["## HTTP API reference", ""]for n in range(60):    lines += [f"### GET /v1/carts/{{id}}/lines/{n}",              f"Returns line {n} of the cart with price, qty, discount and tax fields.",              "Errors: 404 when the cart does not exist, 409 when it is locked.", ""]lines += ["## Deploy runbook", ""]lines += [f"{n}. Step {n} of the blue-green deploy: check the dashboard, "          "then flip the load balancer weight by 10 percent." for n in range(1, 41)]lines += ["", "## Changelog", ""]lines += [f"- 1.{n // 10}.{n % 10}: fixed rounding in tax() for region {n}."          for n in range(40)]write("CLAUDE.md", "\n".join(lines) + "\n")print(f"created {root} with {len(list(root.rglob('*.*')))} files")

Run it from the working folder, then commit the result so Git can restore the bug later:

bash
python3 make_demo.py cartsvccd cartsvcgit init -q && git add . && git commit -qm "cartsvc baseline"npm test 2>&1 | grep -E "^ℹ (tests|pass|fail)"

Expected output:

text
created cartsvc with 6 filesℹ tests 240ℹ pass 200ℹ fail 40

What just happened: you now have a Git repository called cartsvc with one real bug. applyDiscount subtracts a percent discount as if it were a flat amount, so all 40 percent cases fail. The full test output is about 75,000 characters, and the committed bundle in dist/ is 286,500 characters on a single line. Stay inside cartsvc/ from here on, because every later command assumes it is your working directory.

One more action before Step 2: run claude once in cartsvc/, accept the workspace trust dialog, and type /exit. Claude Code ignores a project's permissions.allow rules and its status line in a folder you have not trusted. Deny rules and hooks still apply there. Step 8 depends on the allow rules, and Step 9 depends on the status line.

Step 2: Measure your fixed context with /context

Goal: record how many tokens the prefix costs before you change anything.

Why this step: /context is the only measurement in this tutorial that gives the same number every time. Model runs vary. The prefix does not, and every request carries it before any work begins, which makes it your most reliable before-and-after number.

The --setting-sources project,local flag loads only this repository's settings (.claude/settings.json, plus a personal .claude/settings.local.json if you create one) and ignores your personal ~/.claude setup, such as your plugins and user-level CLAUDE.md. That keeps the comparison about this project. MSYS_NO_PATHCONV=1 stops Git Bash on Windows from turning /context into a file path, and it does nothing on macOS or Linux.

Run it:

bash
MSYS_NO_PATHCONV=1 claude -p "/context" --setting-sources project,local \  < /dev/null | grep -E "Tokens:|Memory files|Free space"

Expected output:

text
**Tokens:** 25.8k / 1m (3%)| Memory files | 7k | 0.7% || Free space | 941.2k | 94.1% |

What just happened: Claude Code printed its context report without calling the model, and grep kept three rows. Memory files is your CLAUDE.md: 7,000 tokens on every request. Your Tokens: total can differ from mine, because it depends on your model and on any connectors your account adds. Memory files should match exactly, since it depends only on the file you just generated. Drop the grep to see every row, including the (deferred) tool rows described in the mental model.

Step 3: Measure what a real task costs with claude -p

Goal: write a script that runs one fixed task headless and prints what it cost.

Why this step: the prefix is only half of the cost. The other half is how many requests the task needs and how much tool output they carry. Headless mode (claude -p) with --output-format json returns the exact token counts and the list-price cost of a run, so you can compare setups on the same task. It is the same pattern used to test a Claude Code setup, pointed at cost instead of correctness.

Save this as measure.sh in cartsvc/:

bash
#!/usr/bin/env bash# measure.sh: run one fixed task headless, print what it cost, restore the bug.set -uo pipefailTASK="The test suite is failing. Run npm test, find the root cause, and fix itin src/. Do not edit the tests. Confirm the suite passes."claude -p "$TASK" --model sonnet --output-format json \  --setting-sources project,local \  --allowedTools "Bash" "Read" "Edit" "Grep" "Glob" \  < /dev/null > last-run.jsonpython3 - <<'PY'import jsonrun = json.load(open("last-run.json", encoding="utf-8"))u = run["usage"]fresh = u["input_tokens"]write, read = u["cache_creation_input_tokens"], u["cache_read_input_tokens"]print(f"turns={run['num_turns']}  cost=USD {run['total_cost_usd']:.4f}")print(f"input processed={fresh + write + read:,}"      f" (cache write {write:,} + cache read {read:,} + uncached {fresh:,})")print(f"input per turn={(fresh + write + read) // run['num_turns']:,}"      f"  output={u['output_tokens']:,}")PYnpm test > /dev/null 2>&1 && echo "tests: PASS" || echo "tests: FAIL"git checkout -- src/   # put the bug back so the next run starts equal

--allowedTools "Bash" lets Claude run any shell command without asking. That is fine inside this throwaway repository, but do not copy the flag into a script that runs in a real one.

Run it twice in a row:

bash
bash measure.shbash measure.sh

Expected output (yours will differ; see below):

text
turns=7  cost=USD 0.1256input processed=168,321 (cache write 21,554 + cache read 146,757 + uncached 10)input per turn=24,045  output=1,004tests: PASSturns=4  cost=USD 0.0380input processed=123,727 (cache write 2,065 + cache read 121,654 + uncached 8)input per turn=30,931  output=536tests: PASS

What just happened: Claude fixed the bug twice, and the script restored it after each run. Look at the two runs side by side:

d2
direction: right
run1: "Run 1: cold cache" {
  style: {fill: "#FFFFFF"; stroke: "#B03A2E"; font-color: "#2C2C2A"}
  w: "cache write 21,554" {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
  c: "USD 0.1256" {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
}
run2: "Run 2: warm cache, same prefix" {
  style: {fill: "#FFFFFF"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
  w: "cache write 2,065" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
  c: "USD 0.0380" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
}
run1 -> run2: "prefix already cached"

The second run found the first run's prefix still in the cache and wrote almost nothing new, so the same task cost 70% less. That is the prompt cache from the mental model, working across sessions in the same directory.

You will need two lessons from this in Step 11. First, compare costs only between runs in the same cache state: cold against cold, or warm against warm. Second, the number of turns changes from run to run (4 and 7 here), so a single run is an illustration, not a benchmark. The input processed and input per turn lines are less sensitive to the cache than the dollar cost, which makes them easier to compare.

Step 4: Move reference material out of CLAUDE.md

Goal: move the API reference into a path-scoped rule and the runbook into a plain document.

Why this step: everything in CLAUDE.md is loaded into every request, whether or not the task needs it. Fixing a discount bug does not need 60 endpoint descriptions or a 40-step deploy runbook. Claude Code has two cheaper places for material like this:

d2
direction: right
always: "CLAUDE.md\nloaded at launch,\nevery request" {
  style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}
}
scoped: ".claude/rules/*.md with paths:\nloaded only when Claude reads\na matching file" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}
ondemand: "docs/*.md\nread only when a task\nasks for it" {
  style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
}
always -> scoped: "API reference\n(needed for src/routes/)"
always -> ondemand: "deploy runbook\n(needed rarely)"

A rule file with a paths: list in its frontmatter loads only when Claude reads a file that matches one of the patterns. A rule file without paths: loads at launch, exactly like CLAUDE.md, so it saves nothing. A normal Markdown file loads only when Claude decides to read it, and the new CLAUDE.md will tell it to do that only for deploys. The changelog goes entirely, because git log already holds it. Do not reach for @path imports here: an imported file loads at launch like the rest of CLAUDE.md, so it helps organization but saves nothing.

This one-off script moves the two sections, from cartsvc/:

bash
python3 - <<'EOF'import pathlibtext = pathlib.Path("CLAUDE.md").read_text(encoding="utf-8")api = text[text.index("## HTTP API reference"):text.index("## Deploy runbook")]runbook = text[text.index("## Deploy runbook"):text.index("## Changelog")]rules = pathlib.Path(".claude/rules")rules.mkdir(parents=True, exist_ok=True)scope = '---\npaths:\n  - "src/routes/**"\n---\n\n'(rules / "http-api.md").write_text(scope + api, encoding="utf-8")pathlib.Path("docs").mkdir(exist_ok=True)pathlib.Path("docs/deploy-runbook.md").write_text(runbook, encoding="utf-8")print("moved", len(api) + len(runbook), "of", len(text), "characters")EOF

cartsvc has no src/routes/ folder. It stands in for the part of a real service that the API reference belongs to, so in this tutorial the rule never loads, which is exactly the point.

Check it: look at the two new files. In the expected output below, the first line is what the script printed, and the rest comes from these two commands:

bash
wc -l .claude/rules/http-api.md docs/deploy-runbook.mdhead -4 .claude/rules/http-api.md

Expected output:

text
moved 14445 of 16464 characters  247 .claude/rules/http-api.md   43 docs/deploy-runbook.md  290 total---paths:  - "src/routes/**"---

What just happened: 88% of the old CLAUDE.md now lives in two files that are not part of the prefix, and the rule file starts with the paths: frontmatter that scopes it to src/routes/. CLAUDE.md itself has not changed yet, so the prefix is still 7k tokens until Step 5 rewrites it.

Step 5: Rewrite CLAUDE.md to what Claude needs on every turn

Goal: replace CLAUDE.md with a short file that holds only instructions every task needs.

Why this step: Anthropic's guidance is to keep each CLAUDE.md under 200 lines, and to test each line by asking whether Claude would make mistakes without it. For a small service the honest answer is under twenty lines: how to run the tests, which folder never to touch, and where the rest of the material lives. Those pointers are what make the cut safe. Claude can still find the runbook when a task needs it, without paying for the runbook on every other task.

The file also gets a # Compact instructions section. When the conversation is compacted (Step 12), Claude Code uses these instructions to decide what the summary keeps.

Overwrite CLAUDE.md in cartsvc/:

bash
cat > CLAUDE.md <<'EOF'# cartsvcPricing service for the checkout flow. Node 20+, ES modules, no dependencies.## Commands- Test: `npm test` (node:test, 240 table-driven cases in test/pricing.test.js)## Layout- src/pricing.js: subtotal, applyDiscount, tax- dist/: committed build output. Never read or edit it; change src/ instead.## Where the rest lives- HTTP API reference: .claude/rules/http-api.md (loads when you work in src/routes/)- Deploy runbook: docs/deploy-runbook.md. Read it only when asked to deploy.- Release history: `git log`, not this file.# Compact instructionsKeep failing test names, files edited, and the current hypothesis. Drop raw test output.EOF

Run it: measure the prefix again with the same command as Step 2:

bash
MSYS_NO_PATHCONV=1 claude -p "/context" --setting-sources project,local \  < /dev/null | grep -E "Tokens:|Memory files|Free space"

Expected output:

text
**Tokens:** 19.1k / 1m (2%)| Memory files | 272 | 0.0% || Free space | 947.9k | 94.8% |

What just happened: Memory files dropped from 7,000 tokens to 272, and the whole prefix from 25.8k to 19.1k, on every request of every session in this repository. The rule file is not counted. That confirms it stays unloaded while nothing in src/routes/ has been read.

Step 6: Block build output with Read deny rules, not .claudeignore

Goal: make it impossible for Claude to read dist/.

Why this step: a pointer in CLAUDE.md is only advice, and the model can decide to ignore it. That is the advisory versus enforced boundary, and a deny rule sits on the enforced side. The first time Claude reads the 286,500-character bundle, it receives this error and then tries again with a smaller slice:

File content (147008 tokens) exceeds maximum allowed tokens (25000). Use offset and limit parameters to read specific portions of the file, or search for specific content instead of reading the whole file.

That is at least one wasted request, and on a file with normal line breaks the slice itself can be 25,000 tokens. Many guides tell you to list such folders in .claudeignore, but Claude Code never read that file. The supported mechanism is a Read(...) rule in permissions.deny. Claude Code enforces it for the Read tool, Grep, Glob, @ mentions, and the Bash file commands it recognizes (cat, head, tail, sed). A Read deny also blocks Edit and Write on the same path.

d2
direction: right
rule: "Read(./dist/**)\nin permissions.deny" {
  style: {fill: "#7B68EE"; stroke: "#5A48C8"; font-color: "#FFFFFF"}
}
covered: "Blocked" {
  style: {fill: "#FFFFFF"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
  a: "Read, Edit, Write" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
  b: "Grep, Glob" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
  c: "@ mentions" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
  d: "cat, head, tail, sed" {style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
}
gap: "Not blocked" {
  style: {fill: "#FFFFFF"; stroke: "#B03A2E"; font-color: "#2C2C2A"}
  a: "grep -r pattern ." {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
  b: "scripts that open\nfiles themselves" {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
}
rule -> covered
rule -> gap

Create .claude/settings.json in cartsvc/:

json
{  "permissions": {    "deny": ["Read(./dist/**)"]  }}

In your own repositories, add the other usual suspects to the same list: Read(./node_modules/**), Read(./build/**), Read(./coverage/**), lockfiles such as Read(./package-lock.json), and Read(**/*.map). A path that starts with ./ is relative to the directory you start Claude Code in, which is the repository root here. A path-scoped deny like these does not change the tool definitions, so it does not break the prompt cache.

Run it: ask Claude to ignore CLAUDE.md and read the bundle anyway, then print which tool calls were denied:

bash
claude -p "Ignore the CLAUDE.md rule once: read dist/cartsvc.min.js." \  --setting-sources project,local --output-format json < /dev/null > deny-check.jsonpython3 -c "import json; r = json.load(open('deny-check.json')); \print([d['tool_name'] for d in r['permission_denials']])"

Expected output:

text
['Read']

What just happened: Claude tried the read, and the settings stopped it before any file content reached the model. The tool result it received was the one-line error File is in a directory that is denied by your permission settings. In my run, Claude then explained that the block came from settings and did not try to work around it. The instruction in CLAUDE.md stops most attempts. The deny rule stops the rest.

Step 7: Write a filter that shrinks test output

Goal: write a small script that reduces test-runner output to the failures and the summary.

Why this step: when tests fail, the Bash tool keeps up to about 10,000 characters of output in the conversation, and most of it is the same stack trace repeated for 40 failures. To find the bug, Claude needs one or two examples, the file and line where the assertion failed, and the counts. Keeping only those lines is plain text processing, so you can build and test the filter without Claude at all.

Create the folder first, with mkdir -p .claude/hooks, then save this as .claude/hooks/test_filter.py:

python
"""Read test-runner output on stdin; print only failures and the summary."""import pathlibimport reimport sysFAIL_START = re.compile(r"^\s*(✖|not ok|FAILED|FAIL\b|--- FAIL)")PASS_LINE = re.compile(r"^\s*(✔|ok |PASSED|--- PASS)")SUMMARY = re.compile(r"^\s*(ℹ |=+ .*(passed|failed)|Tests?:|# (pass|fail))")RUNTIME_FRAME = re.compile(r"^\s+at .*(node:|node_modules)")NOISE = re.compile(r"^\s*(generatedMessage|code|operator):|^\s*\}$")# file:///home/you/cartsvc/test/x.js -> test/x.js (long paths cost tokens too)HERE = re.compile("file:///?" + re.escape(pathlib.Path.cwd().as_posix() + "/"), re.I)MIN_CHARS = 4000   # short output passes through untouchedMAX_FAILURES = 3   # full detail for the first N failures only# Test runners print UTF-8 (✖, ℹ); Windows Python would decode it as cp1252.sys.stdin.reconfigure(encoding="utf-8", errors="replace")sys.stdout.reconfigure(encoding="utf-8")text = sys.stdin.read()if len(text) < MIN_CHARS:    sys.stdout.write(text)    sys.exit(0)def keep(line):    return line.strip() and not RUNTIME_FRAME.match(line) and not NOISE.match(line)failures, current, total_failed = [], None, 0for line in text.splitlines():    if line.strip().startswith("✖ failing tests"):        break  # node's spec reporter repeats every failure in a recap    if FAIL_START.match(line):        total_failed += 1        current = [line] if total_failed <= MAX_FAILURES else None        if current:            failures.append(current)    elif PASS_LINE.match(line):        current = None    elif current is not None and keep(line):        current.append(HERE.sub("", line))print(f"[test_filter: {len(text):,} chars reduced to failures + summary]")for block in failures:    print("\n".join(block))if total_failed > len(failures):    print(f"... {total_failed - len(failures)} more failing tests omitted")print("\n".join(line for line in text.splitlines() if SUMMARY.match(line)))

The patterns cover Node's test runner, TAP output, pytest, and go test. The first line of output always says that filtering happened, so Claude knows it is looking at a summary and can run the tests unfiltered if it needs more.

Run it:

bash
npm test 2>&1 | python3 .claude/hooks/test_filter.py

Expected output (timings and the character count change a little on every run):

text
[test_filter: 75,031 chars reduced to failures + summary]✖ percent 20 on 3 x 170 (3.4198ms)  AssertionError [ERR_ASSERTION]: Expected values to be strictly equal:  490 !== 408      at TestContext.<anonymous> (test/pricing.test.js:252:12)    actual: 490,    expected: 408,✖ percent 25 on 2 x 128 (0.2243ms)  AssertionError [ERR_ASSERTION]: Expected values to be strictly equal:  231 !== 192      at TestContext.<anonymous> (test/pricing.test.js:252:12)    actual: 231,    expected: 192,✖ percent 5 on 9 x 78 (0.3861ms)  AssertionError [ERR_ASSERTION]: Expected values to be strictly equal:  697 !== 666.9      at TestContext.<anonymous> (test/pricing.test.js:252:12)    actual: 697,    expected: 666.9,... 37 more failing tests omittedℹ tests 240ℹ suites 0ℹ pass 200ℹ fail 40ℹ cancelled 0ℹ skipped 0ℹ todo 0ℹ duration_ms 310.1433

What just happened: about 75,000 characters became about 900, and nothing Claude needs to find the bug was lost: three examples of the failure, the test file and line, and the counts. For now the filter only runs when you pipe output into it yourself. Step 8 connects it to Claude Code.

Step 8: Route test commands through the filter with a PreToolUse hook

Goal: make Claude Code pipe every bare test command through the filter automatically.

Why this step: the obvious design is a PostToolUse hook that replaces the tool output after the command runs. The difference between the two events matters here, and hooks as policy-as-code covers it in depth. PostToolUse can replace output through updatedToolOutput, but only for commands that succeed. A test run with failures exits with code 1, so Claude Code fires PostToolUseFailure instead, and that event can only add context. It cannot replace anything. The large output you want to trim is exactly the output this design never sees.

A PreToolUse hook avoids the problem by rewriting the command before it runs. Anthropic's cost guide shows the same pattern with a one-line grep filter. The version here also keeps the exit code, keeps a few complete failures, and passes short output through unchanged. Its updatedInput field replaces the tool's input, and Claude Code then checks your permission rules against the rewritten command. That has one visible effect: Bash(npm test:*) on its own no longer covers the command, because the rewrite turns it into a chain of three parts. Claude Code checks each part of a chain separately, so you need one rule per part, and the settings below have all three: npm test, set -o pipefail, and the python3 call to the filter. Without them, Claude asks for approval every time, and a headless run is denied. set -o pipefail is what preserves the test runner's exit code, so Claude still sees Exit code 1 when tests fail.

d2
direction: right
claude: "Claude wants:\nnpm test" {
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}
hook: "PreToolUse hook\nroute_tests.py" {
  style: {fill: "#7B68EE"; stroke: "#5A48C8"; font-color: "#FFFFFF"}
}
bash: "Bash runs:\nset -o pipefail;\nnpm test 2>&1 | test_filter.py" {
  style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
}
result: "Claude sees:\nExit code 1\n+ ~900 chars" {
  style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
}
post: "PostToolUse\n(success only)" {
  style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}
}
claude -> hook: "tool_input"
hook -> bash: "updatedInput"
bash -> result
bash -> post: "never fires\non exit 1" {style: {stroke: "#B03A2E"; stroke-dash: 4}}

The hook leaves alone any command that already contains a pipe, a redirect, or a chain (|, >, ;, &). Claude often pipes test output through head or grep by itself, and a command it has already shaped should run as written. On Windows, Claude Code also has a PowerShell tool. A test run through that tool does not match the Bash matcher and is not filtered, because the rewrite uses Bash syntax.

Save this as .claude/hooks/route_tests.py:

python
"""PreToolUse hook: pipe bare test-runner commands through test_filter.py."""import jsonimport pathlibimport reimport sysRUNNER = re.compile(    r"^\s*(npm (run )?test|node --test|pytest|go test|npx (jest|vitest))\b")# Absolute path: $CLAUDE_PROJECT_DIR exists for hooks, not inside the Bash tool.FILTER = (pathlib.Path(__file__).parent / "test_filter.py").as_posix()event = json.load(sys.stdin)tool_input = event.get("tool_input", {})command = tool_input.get("command", "")# Leave anything Claude already shaped (pipes, redirects, chains) alone.if not RUNNER.match(command) or any(c in command for c in "|>;&"):    sys.exit(0)tool_input["command"] = f'set -o pipefail; {command} 2>&1 | python3 "{FILTER}"'print(json.dumps({"hookSpecificOutput": {    "hookEventName": "PreToolUse",    "updatedInput": tool_input,}}))

Then register it. Replace .claude/settings.json with this version, which keeps the deny rule from Step 6 and adds the three allow rules:

json
{  "permissions": {    "allow": [      "Bash(npm test:*)",      "Bash(set -o pipefail)",      "Bash(python3 *test_filter.py*)"    ],    "deny": ["Read(./dist/**)"]  },  "hooks": {    "PreToolUse": [      {        "matcher": "Bash",        "hooks": [          {            "type": "command",            "command": "python3 \"$CLAUDE_PROJECT_DIR/.claude/hooks/route_tests.py\""          }        ]      }    ]  }}

Run it: feed the hook the same JSON that Claude Code would send, once for a bare command and once for a command Claude already piped:

bash
echo '{"tool_name":"Bash","tool_input":{"command":"npm test"}}' \  | python3 .claude/hooks/route_tests.py | grep -o 'set -o pipefail; npm test 2>&1'echo '{"tool_name":"Bash","tool_input":{"command":"npm test | tail -5"}}' \  | python3 .claude/hooks/route_tests.py | wc -c

Expected output:

text
set -o pipefail; npm test 2>&10

What just happened: the hook rewrote the bare npm test and printed nothing for the piped one, which means "run it unchanged". To see the rewrite in a real session, ask Claude to run the bare command, then search the session transcript for the filter's header line. Claude Code stores each session as a JSONL file under ~/.claude/projects/, named by its session ID:

bash
claude -p "Use the Bash tool to run exactly: npm test (no pipes). How many fail?" \  --setting-sources project,local --output-format json < /dev/null > hook-check.jsonSID=$(python3 -c "import json; print(json.load(open('hook-check.json'))['session_id'])")grep -oh '\[test_filter: [0-9,]* chars reduced to failures + summary\]' \  ~/.claude/projects/*/"$SID".jsonl | sort -u

Expected output (the character count varies a little):

text
[test_filter: 75,671 chars reduced to failures + summary]

If grep prints nothing, the hook did not fire. Check three troubleshooting rows in this order: the untrusted-workspace warning, a permission_denials entry for the rewritten chain, and the can't open file path error. In my run, the whole tool result Claude received was 894 characters: the Exit code 1 line that Claude Code adds, the filter's output, and nothing else. Without the hook, the same call returned 10,040 characters. Claude answered that 40 of 240 tests fail, all of them percent-discount cases.

Step 9: Show context and cache state in the status line

Goal: see the context size, the cache state, and the session cost at the bottom of the terminal while you work.

Why this step: everything so far happens in the background, and you cannot debug what you cannot see. The status line puts the cost in front of you during the session, while you can still act on it: /compact when the context grows, or a warning when the cache has gone cold after a break. It runs locally and uses no tokens. After each assistant message, Claude Code sends your script a JSON document on standard input and shows whatever the script prints.

d2
direction: right
cc: "Claude Code\nafter each message" {
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}
script: "~/.claude/statusline.py\n(local, no API call)" {
  style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
}
bar: "Sonnet 5.5 | ctx 2% (19.1k of 1M)\n| cache warm 91% | USD 0.06" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}
cc -> script: "JSON on stdin:\ncontext_window,\nprompt_cache, cost"
script -> bar: "stdout"

A status line is a personal preference, so it belongs in your user settings rather than in the repository. Save this as ~/.claude/statusline.py:

python
"""Status line: context fill, prompt-cache state, and session cost."""import jsonimport syss = json.load(sys.stdin)ctx = s.get("context_window") or {}cache = s.get("prompt_cache") or {}used = ctx.get("total_input_tokens") or 0size = ctx.get("context_window_size") or 200_000pct = ctx.get("used_percentage")ctx_part = f"ctx {pct:.0f}%" if pct is not None else "ctx --"window = f"{size // 1_000_000}M" if size >= 1_000_000 else f"{size // 1000}k"ctx_part += f" ({used / 1000:.1f}k of {window})"if not cache:    cache_part = "cache --"elif cache.get("warm"):    cache_part = f"cache warm {cache.get('hit_ratio') or 0:.0%}"else:    cache_part = "cache COLD"miss = (cache.get("last_miss_cause") or {}).get("causes")if miss:    cache_part += f" (last miss: {','.join(miss)})"cost = (s.get("cost") or {}).get("total_cost_usd") or 0model = (s.get("model") or {}).get("display_name", "?")print(f"{model} | {ctx_part} | {cache_part} | ${cost:.2f}")

Every field is read with a fallback, because several are missing or null early in a session: prompt_cache appears only after the first response, and used_percentage can be null before it. last_miss_cause names why the last request missed the cache, for example tools_changed after an MCP server was added in a session with tool search turned off.

Then add the statusLine key to ~/.claude/settings.json, next to any keys already in that file:

json
{  "statusLine": {    "type": "command",    "command": "python3 ~/.claude/statusline.py"  }}

Run it: test the script with a sample of the JSON Claude Code sends, then with an almost empty one:

bash
python3 ~/.claude/statusline.py <<'EOF'{  "model": {"display_name": "Sonnet 5.5"},  "context_window": {"total_input_tokens": 19100, "context_window_size": 1000000,                     "used_percentage": 1.9},  "prompt_cache": {"warm": true, "hit_ratio": 0.91},  "cost": {"total_cost_usd": 0.0613}}EOFecho '{"model": {"display_name": "Sonnet 5.5"}}' | python3 ~/.claude/statusline.py

Expected output:

text
Sonnet 5.5 | ctx 2% (19.1k of 1M) | cache warm 91% | $0.06Sonnet 5.5 | ctx -- (0.0k of 200k) | cache -- | $0.00

What just happened: the script handled a full payload and a nearly empty one, so it will not go blank at the start of a session. To see it live, start an interactive claude session in cartsvc/ and send one message. The line appears below the input box. If you skipped the trust step at the end of Step 1, accept the trust prompt now, or the line stays blank.

Step 10: Set a default effort level for the repository

Goal: make routine tasks in this repository run at low effort by default.

Why this step: effort controls how much the model reasons before it answers. Reasoning tokens are billed as output, and more reasoning often means more tool calls. On current models, effort is the control that replaced the "think harder" keywords and MAX_THINKING_TOKENS. Effort is not the only model-side lever. Picking the model matters more: Anthropic's cost guide recommends Sonnet for most coding work and Opus for hard design and multi-step reasoning. Make that choice when the session starts, not in the middle, because a switch starts a new cache (Step 12). The default on Sonnet 5.5 and Opus 5.5 is medium, which a small service with a good test suite rarely needs for everyday fixes. When a task is hard, raise it for one session with /effort high, or start a session with claude --effort high. The same --effort flag works in headless claude -p runs.

effortLevel in project settings applies to every model, and it takes precedence over the level saved in your user settings. Add it to .claude/settings.json, which now reads:

json
{  "effortLevel": "low",  "permissions": {    "allow": [      "Bash(npm test:*)",      "Bash(set -o pipefail)",      "Bash(python3 *test_filter.py*)"    ],    "deny": ["Read(./dist/**)"]  },  "hooks": {    "PreToolUse": [      {        "matcher": "Bash",        "hooks": [          {            "type": "command",            "command": "python3 \"$CLAUDE_PROJECT_DIR/.claude/hooks/route_tests.py\""          }        ]      }    ]  }}

Run it: make one short headless call, then look up the effort level recorded in its session transcript:

bash
claude -p "Reply with the word ok." --setting-sources project,local \  --output-format json < /dev/null > effort-check.jsonSID=$(python3 -c "import json; print(json.load(open('effort-check.json'))['session_id'])")grep -oh '"effort":"[a-z]*"' ~/.claude/projects/*/"$SID".jsonl | sort | uniq -c

Expected output:

text
      1 "effort":"low"

What just happened: Claude Code records the effort of each response in the session transcript under ~/.claude/projects/, and this one says low. Without the setting, the same check prints "effort":"medium". Anthropic warns that the transcript format is internal and can change between versions, so treat this as a quick manual check and do not build a script on it.

Changing the effort level in the middle of a session breaks the prompt cache on most models, but not on Opus 5.5, Sonnet 5.5, or Fable 5.1. Setting it in the file, before the session starts, is safe on every model.

Step 11: Re-measure token usage and compare costs

Goal: run the same task on the lean setup and compare it with Step 3.

Why this step: a setup change is only worth keeping if the numbers move. You changed four things: the prefix, the file access, the test output, and the effort. measure.sh still runs the same task with the same model, so the comparison is fair as long as you respect the cache lesson from Step 3.

Run it twice, exactly as in Step 3:

bash
bash measure.shbash measure.sh

Expected output:

text
turns=6  cost=USD 0.0552input processed=96,296 (cache write 7,386 + cache read 88,902 + uncached 8)input per turn=16,049  output=786tests: PASSturns=6  cost=USD 0.0330input processed=95,586 (cache write 1,599 + cache read 93,979 + uncached 8)input per turn=15,931  output=781tests: PASS

The first run is cold again, because the new CLAUDE.md and settings changed the prefix. You can see that in its own numbers: a cache write of several thousand tokens, like the first run in Step 3, instead of the one or two thousand of a warm run. Two details explain what you may see. First, a cold write is smaller than the /context total: the tool definitions and most of the system prompt sit at the very start of the prefix and are the same in every session in this folder, so they are usually cached already. Second, my Step 11 runs were captured with no other call on the lean setup before them. If you ran the checks in Steps 6, 8, and 10 within the last hour, they already cached most of the lean prefix, and your first run will look warmer than mine. Compare its input processed line with mine, which does not depend on the cache. Compare it with the first run of Step 3, and compare the warm second runs with each other:

Cold runWarm runInput per turn
Before (Step 3)USD 0.1256USD 0.038024,045 to 30,931
After (Step 11)USD 0.0552USD 0.033016,049 to 15,931

Single runs vary because the number of turns varies. So I also ran the task three or more times per setup, each in a fresh copy of the repository so every run started with a cold cache:

SetupRunsMean costMean cache writeMean cache read
Original repository3USD 0.094516,749107,348
Lean setup, default effort6USD 0.06119,44575,501
Lean setup, low effort (3 runs with --effort low, 3 with effortLevel)6USD 0.05337,65686,641
d2
grid-rows: 3
grid-gap: 8
a: "Original: USD 0.0945 per fix" {
  width: 540
  style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}
}
b: "Lean, default effort: USD 0.0611 (-35%)" {
  width: 349
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}
c: "Lean + low effort: USD 0.0533 (-44%)" {
  width: 305
  style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
}

What just happened: the cold cost of a fix fell by 44% on average, and every run still fixed the bug. According to the transcripts, most of the saving is the smaller prefix: about 7,000 fewer tokens written to the cache once and then read on every turn. Effort added a smaller, less consistent gain on a task this easy. The test filter and the deny rule barely moved these averages, because in this task Claude piped npm test through head by itself and never opened dist/. They are insurance for the sessions where it does not, and when they trigger the difference is large: 10,040 characters against 894 for the test output in Step 8, and a blocked 147,008-token read in Step 6.

Step 12: Compact a long session with /compact and instructions

Goal: shrink the conversation part of the context with /compact, and know when to use its three cheaper alternatives instead.

Why this step: the setup fixes the prefix, but in a long session the conversation becomes the largest part of the context. Every message and tool result stays there and is re-read on every turn until you remove it. These are the commands that remove it, and what each one costs:

d2
direction: down
q: "What do you need?" {
  shape: diamond
  style: {fill: "#7B68EE"; stroke: "#5A48C8"; font-color: "#FFFFFF"}
}
new: "/clear\nnew task, fresh context\n(costs nothing)" {
  style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
}
cont: "/compact <what to keep>\nsame task, long history\n(one summarizing request)" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}
back: "/rewind\nlast steps went wrong\n(reuses the cached prefix)" {
  style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
}
side: "/btw <question>\nquick side question\n(answer never enters history)" {
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}
q -> new: "different task"
q -> cont: "same task,\ncontext is big"
q -> back: "wrong turn"
q -> side: "side question"
  • /clear starts a new conversation. It is free, and after it, the next request re-reads only the prefix. Use /rename first if you may want to come back with /resume.
  • /compact <instructions> replaces the history with a summary. The summary itself costs one request that reads the whole conversation, which is cheap while the cache is warm and expensive after it expires. Your # Compact instructions section from Step 5 is applied automatically.
  • /rewind goes back to an earlier point. That point is already a cached prefix, so the next request reads it from the cache.
  • /btw <question> answers a side question in a separate request whose answer never enters the conversation history.

Two habits matter as much as the commands. First, do not switch models in the middle of a session: each model has its own cache, so the next request writes the whole context again. Anthropic's own example is that 100,000 tokens into an Opus session, it is cheaper to let Opus answer an easy question than to switch to Haiku for it. Second, delegate broad searches to a subagent, which works in its own context and returns only a summary. The subagent's own startup context is not free, and Claude Code Subagent Token Cost: The Preamble Never Stops measures exactly what it costs.

Run it: in an interactive session where you have already worked on a task for a while (a new session answers /compact with Not enough messages to compact.), check the Messages row with /context, compact with an instruction, and check it again:

text
/context/compact Keep only the root cause and the one-line fix/context

You can do the same thing headless on the last measure.sh session. --resume continues that session, and each command runs as a separate call:

bash
SID=$(python3 -c "import json; print(json.load(open('last-run.json'))['session_id'])")MSYS_NO_PATHCONV=1 claude -p "/context" --resume "$SID" --fork-session \  --setting-sources project,local < /dev/null | grep -E "Tokens:|Messages"MSYS_NO_PATHCONV=1 claude -p "/compact Keep only the root cause and the one-line fix" \  --resume "$SID" --setting-sources project,local < /dev/nullMSYS_NO_PATHCONV=1 claude -p "/context" --resume "$SID" \  --setting-sources project,local < /dev/null | grep -E "Tokens:|Messages"

Expected output:

text
**Tokens:** 25.3k / 1m (3%)| Messages | 6.1k | 0.6% |**Tokens:** 22.7k / 1m (2%)| Messages | 3.7k | 0.4% |

What just happened: the conversation part of the context dropped from 6,100 tokens to 3,700, and every later turn in that session re-reads the smaller version. The headless /compact call prints nothing, so the four lines above are all you see. The first /context used --fork-session so that it could read the session without changing it. The last one needs no fork, because by then the compaction has already happened and reading it changes nothing. On a six-turn session the gain is small. On a three-hour session with hundreds of thousands of tokens of history, it is the largest lever you have.

Troubleshooting Claude Code token, hook, and context errors

Every error below was hit while building this tutorial, or is quoted from Claude Code's error reference.

Symptom (verbatim)Root causeFix
Python was not found; run without arguments to install from the Microsoft Store, or disable this shortcut from Settings > Manage App Execution Aliases.On Windows, python3 is resolving to the Store shortcut, not an installed PythonInstall Python 3.13 (from python.org or the Microsoft Store). If a python.org install still hits the shortcut, turn off the python.exe and python3.exe entries under Settings > Apps > Advanced app settings > App execution aliases, then open a new Git Bash window
claude -p "/context" answers a prompt instead of printing the report, because the prompt it received was C:/Program Files/Git/contextGit Bash on Windows converts arguments that start with / into Windows pathsPrefix the command with MSYS_NO_PATHCONV=1
Warning: no stdin data received in 3s, proceeding without it. If piping from a slow command, redirect stdin explicitly: < /dev/null to skip, or wait longer.claude -p waits for piped input when standard input is not a terminalAdd < /dev/null to headless calls, as every command in this tutorial does
python3.exe: can't open file 'C:\\.claude\\hooks\\test_filter.py': [Errno 2] No such file or directoryThe rewritten command used $CLAUDE_PROJECT_DIR, which is set for hook processes but not inside the Bash tool, so it expanded to nothingBuild an absolute path in the hook, as route_tests.py does with pathlib.Path(__file__)
The filter prints only [test_filter: 76,328 chars reduced to failures + summary] and nothing elseOn Windows, Python decodes piped input as cp1252, so ✖ and ℹ never matchKeep the two reconfigure(encoding="utf-8") lines in test_filter.py
Ignoring 3 permissions.allow entries from .claude/settings.json: this workspace has not been trusted.You never accepted the trust dialog in this folder, so its allow rules are skippedRun claude interactively once in the folder, accept the dialog, and /exit
permission_denials lists a Bash call whose command starts with `set -o pipefail; npm test 2>&1python3`The hook rewrote the command, and no allow rule matches every part of the new chain
Test output is still about 10,000 characters although a PostToolUse hook is configuredA failing command fires PostToolUseFailure, which cannot replace outputRewrite the command in a PreToolUse hook (Step 8)
File content (147008 tokens) exceeds maximum allowed tokens (25000). Use offset and limit parameters to read specific portions of the file, or search for specific content instead of reading the whole file.Claude tried to read a huge single-line fileDeny the path with a Read(...) rule (Step 6)
File is in a directory that is denied by your permission settings.The deny rule is workingIf you really need the file, remove the rule from .claude/settings.json. That changes it for everyone who uses the repository. An allow rule in another settings file cannot override it, because deny rules always win
The status line stays blank; the debug log shows Status line command skipped: workspace trust not acceptedNew folders start untrustedStart claude interactively once in the folder and accept the trust prompt
Context limit reached · /compact or /clear to continueThe conversation filled the window/compact with instructions, or /clear and restart the task with a sharper prompt
Not enough messages to compact./compact in a new or almost empty sessionNothing to fix. There is nothing to compact yet
You've hit your session limit · resets 3:45pmThe plan's 5-hour usage window is used upWait for the reset. To see it coming, extend statusline.py to print rate_limits.five_hour.used_percentage, which Claude Code sends to Pro and Max subscribers

How the token-lean setup fits together

This is the setup you built, placed on the request that Claude Code sends every turn:

d2
direction: right
repo: "cartsvc/" {
  style: {fill: "#FFFFFF"; stroke: "#2C6FB0"; font-color: "#2C2C2A"}
  claudemd: "CLAUDE.md\n18 lines, 272 tokens" {
    style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}
  }
  rules: ".claude/rules/http-api.md\npaths: src/routes/**" {
    style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
  }
  docs: "docs/deploy-runbook.md\nread on demand" {
    style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}
  }
  settings: ".claude/settings.json" {
    style: {fill: "#7B68EE"; stroke: "#5A48C8"; font-color: "#FFFFFF"}
  }
  hooks: ".claude/hooks/\nroute_tests.py + test_filter.py" {
    style: {fill: "#7B68EE"; stroke: "#5A48C8"; font-color: "#FFFFFF"}
  }
  dist: "dist/ (denied)" {
    style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}
  }
}
request: "Each request" {
  style: {fill: "#FFFFFF"; stroke: "#2C6FB0"; font-color: "#2C2C2A"}
  prefix: "Prefix: system + tools\n+ CLAUDE.md" {
    style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
  }
  history: "History: messages +\nfiltered tool results" {
    style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
  }
}
user: "~/.claude/statusline.py\nshows ctx, cache, cost" {
  style: {fill: "#FFA07A"; stroke: "#D9774F"; font-color: "#2C2C2A"}
}
repo.claudemd -> request.prefix: "every turn"
repo.rules -> request.history: "only when\nsrc/routes/ is read"
repo.settings -> repo.hooks: "PreToolUse"
repo.settings -> repo.dist: "Read deny"
repo.hooks -> request.history: "~900 chars\ninstead of ~10,000"
request -> user: "JSON after\neach message"

The files on the left of the repository box change what is in the prefix. You pay for the prefix on every turn, so this is the largest and most reliable saving. The settings and hooks change what a single tool call can add to the history, and they save the most in exactly the sessions that would otherwise go wrong. The effort level in settings.json sets how much the model reasons, and the status line shows the result while you work.

The complete token-lean setup

The final state of the repository, not counting .git/:

text
cartsvc/├── .claude/│   ├── hooks/│   │   ├── route_tests.py│   │   └── test_filter.py│   ├── rules/│   │   └── http-api.md│   └── settings.json├── .gitattributes├── CLAUDE.md├── deny-check.json      (written by Step 6)├── dist/│   └── cartsvc.min.js├── docs/│   └── deploy-runbook.md├── effort-check.json    (written by Step 10)├── hook-check.json      (written by Step 8)├── last-run.json        (written by measure.sh)├── measure.sh├── package.json├── src/│   └── pricing.js└── test/    └── pricing.test.js~/.claude/├── settings.json        (adds the statusLine key)└── statusline.py

The files you wrote by hand, in their final form:

CLAUDE.md (Step 5):

markdown
# cartsvcPricing service for the checkout flow. Node 20+, ES modules, no dependencies.## Commands- Test: `npm test` (node:test, 240 table-driven cases in test/pricing.test.js)## Layout- src/pricing.js: subtotal, applyDiscount, tax- dist/: committed build output. Never read or edit it; change src/ instead.## Where the rest lives- HTTP API reference: .claude/rules/http-api.md (loads when you work in src/routes/)- Deploy runbook: docs/deploy-runbook.md. Read it only when asked to deploy.- Release history: `git log`, not this file.# Compact instructionsKeep failing test names, files edited, and the current hypothesis. Drop raw test output.

.claude/settings.json (Step 10):

json
{  "effortLevel": "low",  "permissions": {    "allow": [      "Bash(npm test:*)",      "Bash(set -o pipefail)",      "Bash(python3 *test_filter.py*)"    ],    "deny": ["Read(./dist/**)"]  },  "hooks": {    "PreToolUse": [      {        "matcher": "Bash",        "hooks": [          {            "type": "command",            "command": "python3 \"$CLAUDE_PROJECT_DIR/.claude/hooks/route_tests.py\""          }        ]      }    ]  }}

.claude/hooks/test_filter.py (Step 7):

python
"""Read test-runner output on stdin; print only failures and the summary."""import pathlibimport reimport sysFAIL_START = re.compile(r"^\s*(✖|not ok|FAILED|FAIL\b|--- FAIL)")PASS_LINE = re.compile(r"^\s*(✔|ok |PASSED|--- PASS)")SUMMARY = re.compile(r"^\s*(ℹ |=+ .*(passed|failed)|Tests?:|# (pass|fail))")RUNTIME_FRAME = re.compile(r"^\s+at .*(node:|node_modules)")NOISE = re.compile(r"^\s*(generatedMessage|code|operator):|^\s*\}$")# file:///home/you/cartsvc/test/x.js -> test/x.js (long paths cost tokens too)HERE = re.compile("file:///?" + re.escape(pathlib.Path.cwd().as_posix() + "/"), re.I)MIN_CHARS = 4000   # short output passes through untouchedMAX_FAILURES = 3   # full detail for the first N failures only# Test runners print UTF-8 (✖, ℹ); Windows Python would decode it as cp1252.sys.stdin.reconfigure(encoding="utf-8", errors="replace")sys.stdout.reconfigure(encoding="utf-8")text = sys.stdin.read()if len(text) < MIN_CHARS:    sys.stdout.write(text)    sys.exit(0)def keep(line):    return line.strip() and not RUNTIME_FRAME.match(line) and not NOISE.match(line)failures, current, total_failed = [], None, 0for line in text.splitlines():    if line.strip().startswith("✖ failing tests"):        break  # node's spec reporter repeats every failure in a recap    if FAIL_START.match(line):        total_failed += 1        current = [line] if total_failed <= MAX_FAILURES else None        if current:            failures.append(current)    elif PASS_LINE.match(line):        current = None    elif current is not None and keep(line):        current.append(HERE.sub("", line))print(f"[test_filter: {len(text):,} chars reduced to failures + summary]")for block in failures:    print("\n".join(block))if total_failed > len(failures):    print(f"... {total_failed - len(failures)} more failing tests omitted")print("\n".join(line for line in text.splitlines() if SUMMARY.match(line)))

.claude/hooks/route_tests.py (Step 8):

python
"""PreToolUse hook: pipe bare test-runner commands through test_filter.py."""import jsonimport pathlibimport reimport sysRUNNER = re.compile(    r"^\s*(npm (run )?test|node --test|pytest|go test|npx (jest|vitest))\b")# Absolute path: $CLAUDE_PROJECT_DIR exists for hooks, not inside the Bash tool.FILTER = (pathlib.Path(__file__).parent / "test_filter.py").as_posix()event = json.load(sys.stdin)tool_input = event.get("tool_input", {})command = tool_input.get("command", "")# Leave anything Claude already shaped (pipes, redirects, chains) alone.if not RUNNER.match(command) or any(c in command for c in "|>;&"):    sys.exit(0)tool_input["command"] = f'set -o pipefail; {command} 2>&1 | python3 "{FILTER}"'print(json.dumps({"hookSpecificOutput": {    "hookEventName": "PreToolUse",    "updatedInput": tool_input,}}))

~/.claude/statusline.py (Step 9):

python
"""Status line: context fill, prompt-cache state, and session cost."""import jsonimport syss = json.load(sys.stdin)ctx = s.get("context_window") or {}cache = s.get("prompt_cache") or {}used = ctx.get("total_input_tokens") or 0size = ctx.get("context_window_size") or 200_000pct = ctx.get("used_percentage")ctx_part = f"ctx {pct:.0f}%" if pct is not None else "ctx --"window = f"{size // 1_000_000}M" if size >= 1_000_000 else f"{size // 1000}k"ctx_part += f" ({used / 1000:.1f}k of {window})"if not cache:    cache_part = "cache --"elif cache.get("warm"):    cache_part = f"cache warm {cache.get('hit_ratio') or 0:.0%}"else:    cache_part = "cache COLD"miss = (cache.get("last_miss_cause") or {}).get("causes")if miss:    cache_part += f" (last miss: {','.join(miss)})"cost = (s.get("cost") or {}).get("total_cost_usd") or 0model = (s.get("model") or {}).get("display_name", "?")print(f"{model} | {ctx_part} | {cache_part} | ${cost:.2f}")

The statusLine key in ~/.claude/settings.json (Step 9):

json
{  "statusLine": {    "type": "command",    "command": "python3 ~/.claude/statusline.py"  }}

.claude/rules/http-api.md and docs/deploy-runbook.md were generated by the script in Step 4, and measure.sh is from Step 3. Commit everything under .claude/ except settings.local.json, so the whole team gets the same lean prefix and the same hooks.

To apply this to a real repository, work in the same order: measure with /context and a fixed claude -p task, shrink CLAUDE.md, deny the build and vendor folders, filter your noisiest command, set an effort level, and measure again.

Next steps to cut Claude Code costs further

  1. Filter your second-noisiest command. Test output is rarely the only flood. Find the commands whose output is largest in your own transcripts (build logs, terraform plan, linters), and add a runner pattern and a filter for each one. Test every filter with a pipe first, as in Step 7, before you connect it.
  2. Put the cost on a dashboard. Set CLAUDE_CODE_ENABLE_TELEMETRY=1 and export the claude_code.token.usage and claude_code.cost.usage metrics over OpenTelemetry. That gives you per-developer and per-model numbers over weeks, instead of one measure.sh run.
  3. Run a cheaper model for search-heavy subagents. Define a project subagent with model: haiku for codebase exploration, and measure it with the same measure.sh pattern. The subagent cost analysis tells you how large its fixed startup cost is before you decide.

References


Agentic AI

Follow for more technical deep dives on AI/ML systems, production engineering, and building real-world applications:


Books by Ranjan Kumar

The 7 GenAI Architectures cover

The 7 GenAI Architectures

Building Real-World Agentic AI Systems with LangGraph cover

Building Real-World Agentic AI Systems

The ChatML Handbook cover

The ChatML Handbook

The Chat Templates Handbook cover

The Chat Templates Handbook

Comments