BLOG

Claude Code MCP token usage: what servers cost and how to cut it

5 min read • October 2026
On this page
  1. What an MCP server adds at session start
  2. When tool definitions load upfront anyway
  3. The cost that adds up: tool results
  4. How to see what your MCP servers cost
  5. How to cut MCP cost
  6. Frequently asked questions
  7. Related

By default, an MCP server costs Claude Code very little at session start: only its tool names and server instructions load, and the full tool definitions wait until Claude needs them. The cost that adds up is on the other side: large tool results, and setups where tool search is off and every definition loads upfront.

This post covers what loads, when the default changes, the limits on tool output, and how to see what each server costs you.

What an MCP server adds at session start

With tool search, which is on by default, Claude Code defers MCP tool definitions. According to the MCP docs, "only tool names and server instructions load at session start, so adding more MCP servers has minimal impact on your context window". Claude Code sets no fixed cap on tools per server. The practical limit is your context window budget.

Two details to know:

  • Descriptions are truncated. Claude Code cuts each tool description and each server's instructions at 2,048 characters by default. CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH changes the limit, from v2.1.280.
  • One server can opt out. Set alwaysLoad: true on a server in .mcp.json and all its tools load at session start. The docs recommend it for a small number of tools Claude needs on every turn, because "each upfront tool consumes context that would otherwise be available for your conversation".

When tool definitions load upfront anyway

Tool search needs a model that supports tool_reference blocks: Claude Sonnet 4.5, Haiku 4.5, Opus 4.5 and later. Claude Code loads MCP tools upfront, with no deferral, in these cases:

  • ENABLE_TOOL_SEARCH=false is set
  • ANTHROPIC_BASE_URL points to a non-first-party host, such as a proxy or gateway, unless you set ENABLE_TOOL_SEARCH explicitly
  • The model runs on Google Cloud's Agent Platform and is older than the Claude 4.5 generation
  • The deployment is Microsoft Foundry hosted on Azure, which rejects tool search server-side
  • CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS is set, which keeps tool search off

ENABLE_TOOL_SEARCH takes five kinds of value:

Value Behavior
unset All MCP tools deferred and loaded on demand, with the fallbacks above
true All MCP tools deferred, also through a proxy, which must forward tool_reference blocks: requests fail on proxies that do not
auto Loads tools upfront while their definitions total under 10% of the context window, defers all of them above that
auto:N The same threshold mode with your own percentage, 0-100, for example auto:5
false All MCP tools loaded upfront

If you run Claude Code through a gateway and the MCP cost looks high, this is the first place to look.

The cost that adds up: tool results

A deferred definition is loaded only when needed. A tool result is paid on the request that reads it, then stays in the context until a compaction or /clear, re-read by every later request. The docs set these limits on MCP output:

  • Claude Code warns when a single MCP tool output passes 10,000 tokens.
  • The default maximum is 25,000 tokens, set with MAX_MCP_OUTPUT_TOKENS, for tools that do not declare their own limit.
  • A text result over the limit, or longer than 50,000 characters, is saved to a file in the session's tool-results folder under ~/.claude/projects/. The conversation gets the file path instead, and Claude reads the file when it needs the content.

A server that returns whole pages, query results or logs is the one to watch. Ask it for less per call, or raise the limit on purpose: export MAX_MCP_OUTPUT_TOKENS=50000.

How to see what your MCP servers cost

  • /mcp shows which servers are connected and their status.
  • /context breaks the current window down by category, with suggestions. It describes the session you are in.
  • Across sessions: SkillKeeper reads your local Claude Code logs. Its MCP tab lists each server with its calls and the tokens of the model calls that read its results, split evenly between the tools in a call. A pinned row shows how much of your usage had nothing to do with MCP. The same tab runs inside Claude Code on Windows and Linux through the SkillKeeper mod.

How to cut MCP cost

  1. Keep tool search on. If it is off, find out why: a gateway, an old model or ENABLE_TOOL_SEARCH=false.
  2. Use alwaysLoad sparingly. Only for the few tools Claude needs every turn.
  3. Find the server with big results. Sort your MCP servers by tokens over a week and open the top one.
  4. Narrow what the tool returns. Ask for fewer rows or fields per call.
  5. Remove servers you do not use. A server you never call still lists its tool names and instructions at every session start.

Frequently asked questions

Do MCP servers use tokens when I am not calling them?

A little. Tool names and server instructions load at session start. With tool search on, the full definitions stay deferred until Claude needs a tool.

How many MCP servers is too many?

Claude Code sets no fixed cap per server. The limit is your context window budget. With tool search on, more servers add names and instructions, not full definitions.

What is MAX_MCP_OUTPUT_TOKENS?

The ceiling on a single MCP tool result, 25,000 tokens by default. Claude Code warns above 10,000 tokens, and a result over the ceiling is saved to a file instead of entering the conversation.

How do I stop Claude Code loading all MCP tool definitions?

Leave tool search on, which is the default. If you run through a proxy or gateway, Claude Code turns it off; set ENABLE_TOOL_SEARCH=true or auto only when the gateway forwards tool_reference blocks, because requests fail on proxies that do not.

Why is my MCP context usage high?

Check three things: whether tool search is off (a gateway, an older model or ENABLE_TOOL_SEARCH=false), whether a server uses alwaysLoad, and whether one server returns very large results.

About SkillKeeper