BLOG

Why Claude Code burns through tokens and how to find the cause

4 min read • September 2026
On this page
  1. Why one message can cost a whole conversation
  2. The usual culprits
  3. Find your own culprit first
  4. Then fix only what is big

Claude Code re-sends the whole conversation with every request. A long session, a cache miss after a break, a large MCP result or a loop multiplies the cost. Run /usage to see which source takes the most, then fix only that one.

Most advice about Claude Code tokens is a list of habits: clear more often, pick a cheaper model, trim CLAUDE.md. All of it is correct, and most of it will not help you much, because only one or two of those causes are behind your own usage.

So start by measuring. Then fix the thing that is actually big.

Why one message can cost a whole conversation

Claude Code sends the full conversation with every request. Each time Claude uses a tool, the result comes back as another request that carries the same history plus the new output.

The Claude Code docs put the consequence plainly: a one-line question in a session that has been open all day still draws usage for the whole conversation. Prompt caching makes re-reading that history cheaper, but not free.

That one mechanism is behind most of the list below. Anything that makes the context bigger, or makes it get re-read more often, multiplies.

The usual culprits

1. A session that never ends

Unrelated tasks in one long session mean every new question carries all the old ones. The fix is /clear between tasks. It costs nothing, while /compact is itself a large request because it has to read what it summarizes.

2. Cache misses after a break

The prompt cache lives for an hour on a subscription and five minutes on an API key by default. The first message after a longer break reprocesses the full context at the uncached rate. A big session you come back to after lunch pays for itself twice.

3. MCP tool results

MCP tool definitions are deferred by default, so an idle server costs little. The results are another matter: a server that returns a large page, a full issue list or a whole database row puts it all into context, and it stays there for every following request. The docs suggest using CLI tools like gh where they exist and disabling servers you are not using with /mcp.

4. Memory files that grew

CLAUDE.md and memory files load at session start and ride along on every model call. A file with detailed instructions for a workflow you run once a week costs tokens in every session, all week. The docs recommend keeping CLAUDE.md under 200 lines and moving specialized instructions into skills, which load only when used.

5. Subagents, workflows and teammates

Each subagent sends its own requests on top of the main conversation. That is good for keeping verbose output out of your context, but the subagent's tokens still count against your plan. A workflow that fans out to ten agents is ten conversations.

6. Loops and scheduled tasks

A /loop or scheduled task fires on its interval even while you are away, and each run sends the full context of its session. Leave one running overnight and it can outspend your whole working day.

A real example from one of our own nights: a session was left waiting for edits to a document, checking every 30 minutes. No edits came. By morning it had woken up 61 times, re-read about 650K tokens of context each time, and used 39.8 million tokens on nothing. Nothing in the terminal looked wrong. It only stood out in a per-session breakdown.

7. Thinking and model choice

Extended thinking tokens are billed as output. Lowering the effort level with /effort on simple tasks, or using a smaller model for subagents, cuts this directly.

Find your own culprit first

Three ways to see which of these applies to you:

/usage splits recent usage between skills, subagents, plugins and MCP servers, and flags long context or cache misses when one accounts for 10% or more. Start here. It is built in and takes one command.

/context shows what fills the current window, which answers "why does a fresh session already start heavy".

SkillKeeper keeps the breakdown live in the menu bar and goes one level deeper. It lists tokens by skill, subagent type, built-in tool, MCP server, slash command, memory file, plugin and session, for the last 5 hours, day, week or month. The All view puts every source in one sorted list, so the biggest one is at the top, whatever kind it is. Hovering a row shows its share on the chart, so a loop that spikes every hour is easy to spot.

Then fix only what is big

Once you know the top source, the fix is usually obvious:

Top source What to try
One long session /clear between tasks, /rename first so you can /resume later
An MCP server Use its CLI instead, or disable it when you are not using it
A memory file Cut it down, move workflow-specific parts into skills
A subagent type Give it a smaller model, narrow its prompt
A loop Stretch the interval, or stop it when you are done
A skill Check that it fires only when it should. A skill with a vague description can load on requests it has nothing to do with. See why a skill fires or never fires

Check the numbers again after a day. If the top source is still the same one, the fix did not work. If something new is on top, you are done with the old one.

For a comparison of the ways to watch usage over time, see how to check Claude Code token usage.

About SkillKeeper