BLOG

Claude Code context window full: what happens and what to do

4 min read • October 2026
On this page
  1. What is in the context before you type
  2. When Claude Code compacts
  3. What survives compaction
  4. What to do before the automatic pass
  5. See which chat is heavy, across all of them
  6. Frequently asked questions
  7. Related

When the context window fills, Claude Code compacts the conversation on its own, so the session keeps going. What costs you usage is the stretch before that point: every request re-reads the whole context, and a long session carries it for hundreds of requests.

This post covers what is in the window before you type, when the automatic compaction runs, what it keeps, and the commands that keep a long session cheap.

What is in the context before you type

Claude Code loads these into context at the start of a session, before your first message (context window docs):

  • The system prompt and environment info
  • CLAUDE.md files and auto memory (the first 200 lines or 25 KB of MEMORY.md)
  • MCP tool names, with full schemas deferred until a task needs them
  • One-line descriptions of your skills, except skills with disable-model-invocation: true

After that, everything adds up: each file Claude reads, each command output, path-scoped rules that load next to a matching file, and context that hooks add. A subagent works in its own window, and only its final summary comes back into yours.

When Claude Code compacts

The docs say Claude Code "compacts automatically as you approach the limit, so a full context window doesn't end your session". Where the limit sits depends on the model (auto-compact thresholds):

Model Automatic compaction runs at
Sonnet 5, Haiku 5.5, Fable models, Opus 4.7 and later on the Anthropic API (native 1M window) About 967,000 tokens by default
Sonnet 4.6 and Opus 4.6 without extended context The 200,000-token boundary
Any 1M model with CLAUDE_CODE_DISABLE_1M_CONTEXT=1 The 200,000-token boundary

On a 1M window a session can grow close to a million tokens before anything is compacted, and each reply re-reads that context. A larger window is not a smaller bill. The prompt cache makes repeated reads cheaper but not free, and it expires: after an hour on a subscription, after five minutes with an API key (prompt caching docs). The first request after a break pays the full rate for the whole context again.

What survives compaction

Compaction replaces the conversation with a structured summary. The docs list what comes back:

  • Re-injected from disk: the project-root CLAUDE.md, unscoped rules, auto memory and the plan from plan mode
  • Re-read: up to five files Claude read or edited, most recently modified first; a file over 5,000 tokens comes back as a path reference
  • Re-injected with caps: invoked skill bodies, up to 5,000 tokens each and 25,000 tokens in total, oldest dropped first
  • Summarized away: path-scoped rules, nested CLAUDE.md files and anything hooks added, until Claude touches their trigger file again

If a rule must survive compaction, drop its paths: field or move it to the project-root CLAUDE.md.

What to do before the automatic pass

The docs give five moves:

  1. Compact with a focus. /compact focus on the auth bug fix keeps what you choose, instead of what the automatic pass guesses matters.
  2. Compact part of the conversation. /rewind, pick a message, then Summarize from here or Summarize up to here.
  3. Compact earlier. /autocompact 500k sets how full the window gets before the automatic pass runs. The range is 100,000 to 1,000,000 tokens, and /autocompact auto returns to the default.
  4. Clear between tasks. /clear when you switch to unrelated work: old conversation "crowds out the files you need next and costs tokens on every message".
  5. Delegate large reads. Send research to a subagent so the file contents stay in its window, not yours.

To see where one session stands right now, run /context: a live breakdown by category, with suggestions.

See which chat is heavy, across all of them

/context describes the session you are in. To find the heavy chat among many, you need the tokens per session over a day or a week.

SkillKeeper reads your local Claude Code logs and lists chats by tokens. A chat row opens to show its context against the /compact line, and the menu bar icon turns yellow when a chat's context keeps growing. On Windows and Linux the same view runs inside Claude Code as the SkillKeeper mod.

Frequently asked questions

What happens when the Claude Code context window is full?

Claude Code compacts the conversation automatically. It replaces the history with a structured summary, reloads CLAUDE.md and memory from disk, and re-reads the files you touched most recently. The session continues.

When does Claude Code auto-compact?

On models with a native 1M window, at about 967,000 tokens by default. On models that run at 200,000 tokens, at that boundary. /autocompact sets a different point between 100,000 and 1,000,000 tokens.

Does a bigger context window use more of my limit?

A bigger window lets a session carry more context, and every request re-reads what is in it. Cached reads are cheaper than fresh ones, but the cache expires, so a long session after a break pays the full rate once.

What is the difference between /compact and /clear?

/compact summarizes the conversation and keeps the thread of the work. /clear empties it. Use /clear when the next task is unrelated.

Does compaction lose my CLAUDE.md?

No. The project-root CLAUDE.md, unscoped rules and auto memory are re-injected from disk. Path-scoped rules and nested CLAUDE.md files are summarized away until Claude reads a file that triggers them again.

About SkillKeeper