Skip to content
Let’s talk

AI engineering

Claude Code tokens: you pay for the area, not the height

Every message re-sends everything before it. Picture a session as a staircase and every saving becomes one of four moves you can work out on a napkin.

  • 10 min read
  • Prices and sources checked
What one Claude Code request carriesEvery message re-sends the whole context. The repeated part is cheap only while the cache holds.What one Claude Code request carriesEvery message re-sends the whole context. The repeated part is cheap only while the cache holds.The request Claude Code sendsSystem prompt + tool definitionsCLAUDE.md, memory, skill listConversation so far: messages,file reads, command outputYour new messageRe-read from cache$0.20 per milliontokens on Opus 5.5New tokens: 5-minute cache write, $5 per millionYou typea message1Claude API2One request3Reply and tool results join the conversation(output, including thinking: $20 per million)4Next message:the whole stack is sent again
Fig. 01 — What one Claude Code request carries: everything before your new message is sent again, and is cheap only while it stays cached.

Key takeaways

  • Claude Code re-sends the whole conversation on every request, so a session's input is the sum of every request's context: the area under a staircase, not the size of your last message.
  • Four moves shrink the area: lower the floor (what loads every time), shorten the risers (what each turn adds), reset the run (clear between tasks, compact at breaks), and keep the repeats on the cache.
  • In a long session the risers dominate, because their cost grows with the square of the session length. Resetting halfway through cut our worked example by 42%; trimming the floor saved 6%.
  • Cache reads cost a fraction of fresh input ($0.20 vs $4 per million tokens on Opus 5.5), but only while the start of the conversation stays identical and the cache is warm.
  • Never save tokens at the cost of a finished task: a failed attempt and a retry cost more than any trim.

A one-line question at the end of a long day can use more tokens than the first hour of work. Nothing is broken when that happens: it is how the tool is built. In Anthropic’s words, Claude Code makes a new API request for every message, and because the model remembers nothing between requests, it “re-sends the full context: the system prompt, your project context, every prior message and tool result, and your new message” (Claude Code docs). Every time Claude uses a tool, another request goes out carrying the same history (cost guide).

So the useful question is not “how big is this prompt?” It is “how much did this session carry, request after request?” That has a shape, and once you can see the shape, every saving becomes obvious.

The staircase model

Draw one bar for every request in a session. Its height is everything sent with that request. Three things make up each bar:

  • The floor. What loads on every request before you type a word: Claude Code’s system prompt and tool definitions, your CLAUDE.md files, auto memory, and the one-line descriptions of your skills and subagents.
  • The riser. What one turn adds: your message, Claude’s reply, the files it read, the output of the commands it ran.
  • The run. How many requests go by before something resets the stack.

Because every request repeats everything before it, the bars climb like a staircase. The input you pay for is not the last bar. It is the area of all of them.

The token staircaseEach bar is one request. Every bar carries the floor, everything re-read from earlier turns, and the new tokens of that turn. The bars climb until a compaction drops them back to the floor plus a short summary. The total area of the bars is the input you pay for.Floor — loads on every requestRe-read from earlier turnsNew this turn/compactback to the floor + a summaryfloorrun: requests before a resetrequests →You pay for the area of every bar,not the height of the last one.riser: what a turn adds
Fig. 02 — Each bar is one request. The floor repeats every time; earlier turns are re-read; each turn adds its own step on top. A compaction drops the stack back to the floor plus a summary.

That picture gives a formula you can do on a napkin. For a session of n requests with a floor of F tokens and an average riser of r tokens:

Input processed ≈ n × F + r × n × (n + 1) ÷ 2

The first term grows in a straight line with the length of the session. The second grows with its square. Double the length of a session and the riser part roughly quadruples. That one fact explains why long, never-cleared sessions are expensive, and Anthropic’s own cost guide names exactly that, together with leaving the most expensive model on by default, as the usual cause of high spend (cost guide).

A worked example

Take a long working session and make three round assumptions (they are ours, for illustration; run /context to see your own floor):

  • a floor of 20,000 tokens,
  • each request adds 4,000 tokens (a file read and some output),
  • 60 requests before you stop.

The area is 60 × 20,000 + 4,000 × 60 × 61 ÷ 2 = 1,200,000 + 7,320,000 = 8.52 million input tokens. The last request alone is only 260,000 tokens. The session processed more than thirty times that.

Now try each move on the same session:

Change Input processed Saving
Baseline: one 60-request session 8,520,000 —
Lower the floor from 20k to 12k 8,040,000 6%
Shorten the riser from 4k to 2.5k 5,775,000 32%
Reset halfway: two 30-request sessions 4,920,000 42%
All three together 3,045,000 64%

Two lessons hide in that table. In long sessions, the riser and the run matter far more than the floor. In short sessions it flips: over 10 requests the same floor is nearly half the area. So trim the floor if you start many short sessions (and every subagent starts a fresh one), and watch the risers and resets if you work in long ones.

What does it cost? On the Claude API, with Opus 5.5 at $4 per million fresh input tokens, $5 for a five-minute cache write and $0.20 for a cache read (pricing), the baseline session costs about $2.95 in input if the cache stays warm throughout: 8.26 million tokens read from cache ($1.65) plus 260,000 tokens written ($1.30). The same 8.52 million tokens processed with no cache at all would be about $34. Output is billed on top at $20 per million on Opus 5.5, and thinking counts as output. On a Claude subscription you pay in usage limits rather than dollars, but the shape is identical.

Move 1: lower the floor

The floor repeats on every request, in every session and every subagent. Look at it once, in a fresh session, with /context: it draws the window as a grid, lists the memory files that loaded and suggests what to trim (commands).

  • CLAUDE.md loads in full, every time. Anthropic suggests keeping each file under 200 lines, and notes that @ imports do not reduce the cost, because imported files also load at launch (memory docs). The test from the best-practices page is a good one: “Would removing this cause Claude to make mistakes? If not, cut it.” Block-level HTML comments are stripped before injection, so notes for humans can stay.
  • Move procedures into skills. A skill costs only its name and description until it is used; with disable-model-invocation: true, even the description stays out until you call it yourself (skills). /skills sorts by token cost when you press t.
  • Prefer a CLI to an MCP server where one exists. MCP tool definitions are already deferred by default, but Anthropic still calls tools like gh, aws and gcloud more context-efficient, because they add no per-tool listing at all. Switch off servers you are not using in /mcp (MCP).

Move 2: shorten the risers

Every token a turn adds rides along on every later request. A 30,000-token log read at request 10 of a 60-request session is paid for about fifty times.

  • Ask narrowly. The cost guide contrasts “improve this codebase”, which triggers broad scanning, with “add input validation to the login function in auth.ts”, which needs a handful of reads (cost guide).
  • Quiet the loud commands. Test runners and build tools print far more than Claude needs. The official example is a hook that pipes test output through a filter so only failures come back, which the docs say turns “tens of thousands of tokens” into hundreds. Shell output over about 30,000 characters already comes back as a file path and a short preview (tools reference), but well before that limit it is cheaper to ask for less.
  • Mention a file once. An @ mention puts the whole file into the conversation; mentioning it again later attaches a second copy.
  • Explore in a side staircase. A subagent works in its own context window and hands back only its final answer (subagents). Anthropic’s engineering team describes subagents that use tens of thousands of tokens and return a summary of often 1,000 to 2,000 (context engineering).
Exploring in a subagentLeft: when the main session explores itself, every file it reads stays in context and is re-read on every later turn. Right: a subagent reads the same files in its own window, which is thrown away; only a short summary joins the main session.Exploring in the main sessionevery file read stays, and is re-read on every later turnreading filesExploring in a subagentits window is thrown away; only the summary comes backsubagent window (discarded)summary
Fig. 03 — The same exploration, two ways. On the left, every file read stays in the main session and is re-read on every later turn. On the right, a subagent climbs its own staircase, which is then thrown away; only the summary joins the main session.

A subagent is not free: it pays for its own requests, so use it for verbose, self-contained work (searching, reading logs, running a test suite), not for quick edits that need the conversation. It can also run on a cheaper model with model: haiku in its definition.

Move 3: reset the run

The run is where the square term lives, so this move pays the most.

  • /clear between unrelated tasks. It starts a fresh conversation and costs nothing. Name the old one with /rename first and you can /resume it later. The best-practices page adds a rule worth keeping: after two failed corrections, clear and write a better prompt, because “a clean session with a better prompt almost always outperforms a long session with accumulated corrections” (best practices).
  • /compact at natural breaks, with a focus. /compact focus on the auth bug fix keeps what matters; a # Compact instructions section in CLAUDE.md sets the default. Compaction re-injects CLAUDE.md and up to five recently edited files, and summarises the rest (context window). It is one request that reads the whole conversation, so it is cheap only while the cache is warm.
  • /rewind instead of arguing. If Claude went down a wrong path, rewinding cuts the conversation back to an earlier point that is already cached, rather than piling corrections on top.
  • Set your own ceiling. On the 1M-token models, automatic compaction waits until about 967,000 tokens by default (model configuration). That is a very tall staircase. /autocompact 200k resets it at a height you would actually choose to work at.

Move 4: keep the repeats on the cache

The re-read part of every bar is where caching earns its keep. Claude Code caches automatically, and a cache read costs a fraction of fresh input: 0.05× on Opus 5.5 and 0.1× on Sonnet 5.5 and Haiku 4.5 (pricing).

Model Fresh input 5-min cache write 1-hour cache write Cache read Output
Opus 5.5 $4 $5 $8 $0.20 $20
Sonnet 5.5 $2 $2.50 $4 $0.20 $10
Haiku 4.5 $1 $1.25 $2 $0.10 $5

Prices per million tokens on the Claude API, checked 4 October 2026.

Two conditions decide whether you get that price:

  • The start of the conversation must stay identical. The cache is an exact prefix match, and “a change anywhere in the prefix recomputes everything after it” (prompt caching). Switching models mid-session, turning on fast mode, compacting and upgrading Claude Code all rebuild it; on most models so does changing effort. Anthropic’s advice is to pick your model and effort level at the top of a session. A counter-intuitive consequence: switching a long session to a cheaper model for one easy question can cost more than staying put. Hand that question to a subagent instead.
  • The cache must still be warm. On a Claude subscription the main conversation’s cache lasts an hour; on an API key it lasts five minutes. After a longer pause, the next message reprocesses everything at full price. Hence the official rule of thumb: compact before a break, not after one. /usage shows your cache hit rate and the likely cause of the last miss.

When saving tokens costs more

Smaller context usually means better work, not just cheaper work. Anthropic’s engineers describe “context rot”: as the token count grows, the model’s ability to recall what is in the window declines, so the goal is “the smallest set of high-signal tokens” that gets the job done (context engineering).

But the trade-off is real. Anthropic’s own post on Opus 5.5 costs warns that “every way to spend fewer tokens can also cost you a finished task”, and that a retry costs more than the savings (Claude blog). The best-practices page agrees that when you are deep in one hard problem, the history is worth keeping. And on model choice, Anthropic’s pages lean differently: the cost guide puts most coding on Sonnet, while the newer Opus 5.5 post treats Opus 5.5 as the everyday model and suggests raising effort before switching. Both agree on the cheap part: run searches and log reading on smaller models.

The 15-minute token audit

Our checklist for a team adopting Claude Code. Do it once, then keep the habits.

  1. Measure the floor. Open a fresh session and run /context. Write the number down.
  2. Cut CLAUDE.md with the “would removing this cause a mistake?” test, aiming under 200 lines; move procedures into skills.
  3. Prune skills and servers. /skills (press t) and /mcp; switch off what you do not use and prefer CLIs.
  4. Quiet your loudest command. Add a filter hook or quiet flags for the test runner or build tool you run most.
  5. Make a cheap explorer. A subagent with model: haiku for searching and reading, so exploration climbs its own staircase.
  6. Set a ceiling with /autocompact at a height you would choose.
  7. Keep four habits: model and effort at the start, /clear between tasks, /compact before a break, /usage once a week to check the cache line.

Do that, and the staircase stays low, the cache stays warm, and the model works with the context it needs instead of everything it has ever seen.

Questions people ask

Does Claude Code send my whole conversation every time?

Yes. Each message is a new API request and the model keeps no memory between requests, so Claude Code re-sends the system prompt, your project context, every earlier message and tool result, and your new message. Tool calls add further requests that carry the same history.

Is /compact free?

No. /compact is one request that reads the conversation it summarises, so compacting a large context is itself a large request. It is cheap while the cache is warm and expensive after a long break. /clear, which starts a fresh conversation, costs nothing.

Does the 1M-token context window cost more per token?

No. On Claude 4.6 and later models the full 1M window is billed at the standard rate. But a bigger window lets the staircase climb higher before anything resets it, and every token in it is re-sent on each request.

Do MCP servers fill my context?

Less than they used to. Claude Code defers MCP tool definitions by default and loads only tool names and server instructions at the start. Tool output is a different matter: Claude Code warns when one MCP result passes 10,000 tokens and caps it at 25,000 by default.

Are thinking tokens billed?

Yes, as output tokens, even when the thinking is collapsed or hidden. On Opus 5.5, Sonnet 5.5 and the Fable models thinking cannot be switched off, so the lever is the effort level (/effort).

Should I use Sonnet or Opus to save money?

Anthropic's pages differ in emphasis. The cost guide says Sonnet handles most coding and costs less, with Opus kept for complex reasoning; a newer Anthropic post treats Opus 5.5 as the daily driver and suggests raising effort before changing models. Both agree on putting searches and log reading on cheaper models.

What does Claude Code cost on average?

Anthropic's cost guide says that across enterprise deployments the average is around $13 per developer per active day and $150–250 per developer per month, with 90% of users below $30 per active day.

Why did one small question use so much?

Because it carried the whole session with it. A one-line question at the end of a long day re-sends everything before it, and if the cache has expired since your last message, all of it is processed at the full input price.

Sources

  1. Claude Code docs — Manage costs effectively · checked 4 Oct 2026
  2. Claude Code docs — How Claude Code uses prompt caching · checked 4 Oct 2026
  3. Claude Code docs — Explore the context window · checked 4 Oct 2026
  4. Claude Code docs — Model configuration · checked 4 Oct 2026
  5. Claude Code docs — How Claude Code works · checked 4 Oct 2026
  6. Claude Code docs — Best practices for Claude Code · checked 4 Oct 2026
  7. Claude Code docs — How Claude remembers your project (CLAUDE.md) · checked 4 Oct 2026
  8. Claude Code docs — Create custom subagents · checked 4 Oct 2026
  9. Claude Code docs — Extend Claude with skills · checked 4 Oct 2026
  10. Claude Code docs — Connect Claude Code to tools via MCP · checked 4 Oct 2026
  11. Claude Code docs — Tools reference · checked 4 Oct 2026
  12. Claude Code docs — Commands · checked 4 Oct 2026
  13. Claude API docs — Pricing · checked 4 Oct 2026
  14. Claude API docs — Context windows · checked 4 Oct 2026
  15. Anthropic Engineering — Effective context engineering for AI agents · 29 Sep 2025
  16. Claude blog — What a task costs on Opus 5.5 · 25 Sep 2026

How we research: every price in this article was checked on the vendor's own page, and every claim links to where it came from, as of 4 October 2026. Prices change — confirm them before you buy.

More insights

All insights
How an AI assistant answers from your documentsNothing is retrained: it looks up your files on every question, then answers only from what it found.How an AI assistant answers from your documentsNothing is retrained: it looks up your files on every question, then answers only from what it found.Customeror staffYour assistantpermissions · ruleslogs · handoffSearch yourdocumentsPolicies, helparticles, dataAI modelvia APIA persontakes over1Asks a question2Question3Best passages + page numbers4Question + passages + rules:cite, or say “I don't know”5Answer + citation“Returns policy, p.2”No good passage? Hand off.

AI · 15 min read

AI chatbot trained on your own data: a plain guide for small businesses

You are not really training it, and that is good news. The assistant looks up your documents every time it answers, which changes how you should think about privacy, wrong answers and cost.

Rent, build, or something in between?Four questions, in order. Most small firms should stop before the last one.Rent, build, or something in between?Four questions, in order. Most small firms should stop before the last one.A process notool does wellCommodity tool?email, CRM, helpdeskProcess provenand stable?Owner + 5 years ofupkeep budgeted?5-year build costbelow 5-year rent?NoYesYesYesRent itright-size seats,cap renewal risesNoProve it firstspreadsheet orno-codeNoMiddle pathconnect tools or athin custom layerYesBuild itand own the codeand accountsNo: rent, or a thin layer

Software costs · 14 min read

Custom software vs off-the-shelf: what each really costs over five years

A renewal notice on one side of the desk, an agency quote on the other. Here is how to compare them fairly: real prices, upkeep counted in hours, and a worksheet that will sometimes tell you not to build.

Shared working hours with IndiaA developer in India working 11:00–20:00 IST, against a 09:00–17:30 client day (summer time).Shared working hours with IndiaA developer in India working 11:00–20:00 IST, against a 09:00–17:30 client day (summer time).09:0012:0015:0018:0021:0000:0003:0006:00ISTIndia developer11:00–20:00 ISTLondon (BST)09:00–17:30 = 13:30–22:00 IST6.5 h sharedNew York (EDT)09:00–17:30 = 18:30–03:00 IST1.5 h sharedLos Angeles (PDT)09:00–17:30 = 21:30–06:00 ISTno overlapIndia does not change its clocks (UTC+5:30). In UK and US winter, every client time above shifts one hour later in IST:London overlap becomes 5.5 h, New York 0.5 h.

Hiring · 15 min read

Hire dedicated developers in India: what it really costs, and how to stay in control

A big saving against US pay, a small one against a UK salaried hire. The real risks sit in the contract and the accounts, and you can fix both before day one.

We use cookies to improve your browsing experience.

Read our Privacy Policy