cobrain
Create your brainGo to your brain
Blog · 10 October 2026 · 7 min

How to reduce token usage in Claude Code (measured)

What enters Claude Code's context each turn, the official commands that shrink it (/context, /clear, /compact), MCP tool search, CLAUDE.md size and our numbers.

"Usage limit reached" in the middle of a refactor is annoying the first time and expensive by the third. Most of the tokens Claude Code spends aren't your words. They're what travels with them: instructions, history, files, tool results. To spend less, you need to know what's in there.

This article goes through what enters Claude Code's context on every turn, the official commands and settings that keep it small, and then what we measured on our own work, with a method you can repeat. Everything about Claude Code comes from code.claude.com and support.claude.com, read on 10 October 2026.

What Claude Code sends on every turn

Before you type anything, the context already holds the system prompt, environment details (working directory, platform, git status), your CLAUDE.md files and anything they import with @path, the first 200 lines or 25 KB of auto memory, a one-line description of each skill, and the names of your MCP tools with each server's instructions.

Then the session grows. Every file Claude reads, every command output and every tool result stays in the conversation, and Claude Code sends the full conversation with every request. When Claude uses a tool, that's one more request carrying the whole lot again. Anthropic's support page lists what gets resent each turn: the conversation so far, your CLAUDE.md, the files Claude has read, and your new prompt.

Prompt caching makes this cheaper. Claude Code re-reads the history at the cached token rate, not the full one. But the cache expires: after an hour on a subscription, after five minutes once you're on usage credits or an API key. The first message after a longer break reprocesses everything. So a one-line question in a session that has been open all day still pays for the whole day.

The commands that keep context small

/context shows what's taking space right now, as a grid by category, with suggestions. Start here. It's the only way to see your real numbers instead of guessing.

/usage (alias /cost) shows your plan usage bars. On Pro, Max, Team and Enterprise it also breaks down what counted: skills, subagents, plugins and each MCP server, as a share of the total.

/clear starts a new conversation with empty context. Use it between unrelated tasks, because old history costs tokens on every message that follows. It costs nothing. Run /rename first if you want to come back to the session with /resume.

/compact summarizes the conversation and carries on, and you can tell it what to keep: /compact focus on the auth bug fix. The docs point out the catch. To summarize, it reads the whole conversation, so on a large context it's a large request. If you don't need continuity, /clear is cheaper.

/model switches model. Anthropic's advice: Sonnet for most coding, Opus for hard architecture and multi-step reasoning, Haiku for simple subagent tasks. Switching mid-session has a price, because each model has its own cache and the next request reads the entire history uncached.

/effort lowers reasoning effort for simple work. Thinking tokens are billed as output. On most models, changing effort mid-session also invalidates the cache. On Opus 5.5, Sonnet 5.5, Haiku 5.5 and Fable 5.1, with a subscription or an API key, it doesn't. Thinking can't be turned off on the 5.5 models or on Fable.

Two habits from the same pages: send verbose work, like test runs or log digging, to a subagent, so only its summary comes back; and point to files by path instead of pasting them.

Each MCP tool has a definition: name, description, input schema. Without tool search, all of them load at the start of every session. Claude Code defers them by default: only tool names and server instructions enter context, and Claude loads a tool's full definition when it needs it.

ENABLE_TOOL_SEARCH changes this. With auto, definitions load upfront while they total less than 10% of the context window and are deferred once they reach 10%. auto:N sets your own percentage, and false loads everything upfront. Claude Code also cuts each tool description and each server's instructions at 2,048 characters by default.

A connector still isn't free. Names and instructions sit in context from the first turn, and every definition Claude loads, along with every result it gets back, joins the history that is resent. Disable the servers you aren't using in /mcp. Where a good command-line tool exists (gh, aws), the docs say the CLI uses less context.

When CLAUDE.md is too long

CLAUDE.md loads in full at the start of every session, and again after /clear or /compact. Anthropic's target is under 200 lines per file, because longer files take more context and get followed less closely. Instructions for a single workflow, like PR reviews or database migrations, belong in a skill, which loads only when used.

Lines are a poor measure, though. When we wrote CLAUDE.md vs AGENTS.md in September, the CLAUDE.md in Cobrain's own repository had 51 lines and weighed 43 KB, because each line is a paragraph. On 10 October it has 75 lines and weighs 65 KB. Well under 200 lines, and all of it in context on every turn. We haven't fixed this for ourselves yet. Check the bytes, not the lines.

What we measured

We wanted Claude's own numbers, not estimates. Claude Code keeps a transcript of every session in ~/.claude/projects/<project>/*.jsonl. Each assistant response records its usage: input_tokens, cache_creation_input_tokens and cache_read_input_tokens, which add up to that request's full input. When exactly one tool result sits between two consecutive responses, its size in tokens is about the second input, minus the first input, minus the first response's output. Put that next to the length of the result's text and you get Claude's characters per token, for free. The leftover error is a few tokens of wrapping around the tool result. We did this over 258 real conversations and cross-checked with OpenAI's tokenizer.

On our notes, which are Italian plus Markdown, Claude counts about 2 characters per token, around 1.58 times what OpenAI's tokenizer gives on the same text.

Then we took a real project from our brain: Cobrain itself, 134 notes. Handing Claude Code the files that project marks as always needed (brief, index, decisions) in full at the start of each chat comes to about 69,400 tokens. Cobrain's opening summary for the same project is 4,700 tokens. Connecting Cobrain also adds a fixed cost of about 5,200 tokens, the definitions of its tools. We measured that one with /context, with tool search off so every definition loads, comparing the total with Cobrain enabled and disabled. With tool search, less of that enters at the start, but we count all of it.

So the first message is 69,400 tokens against roughly 9,900, about 7 times less. Over a whole working chat (35 real chats on the same project, every note read included), the median is 2.9 times less. If you count that the model re-reads its context on every response, it's 3.6 times less.

In 4 of the 35 chats Cobrain delivered more than the "without" case. Those chats opened a lot of notes, notes Claude Code would have needed anyway. The details are in context costs.

Where Cobrain fits

Cobrain is a Markdown memory, in real folders, that Claude Code reads and writes over MCP. Instead of a CLAUDE.md that grows with every decision, you keep two lines in it: load the project with start_session, and write new decisions to the brain. Claude Code then searches and opens only the notes it needs, and the same notes work from the Claude app, ChatGPT and Codex. The Claude Code guide takes a few minutes, and the free plan is enough to try it.

Quick answers

How do I reduce token usage in Claude Code? Run /context to see what's loaded, /clear between unrelated tasks, disable unused MCP servers in /mcp, keep CLAUDE.md short, match the model to the task and send verbose work to subagents.

Is /compact or /clear better for saving tokens? /clear costs nothing and empties the context. /compact keeps a summary, but reads the whole conversation to write it, which on a long session is a large request.

How long should CLAUDE.md be? Anthropic suggests under 200 lines. Watch the size in bytes too: the file loads in full every session, and long lines add up.

Do MCP servers use tokens in Claude Code? Yes. With tool search, on by default, tool names and server instructions load at the start and full definitions load when used. Their results stay in the conversation.

What does "usage limit reached" mean in Claude Code? You've hit the five-hour session limit or the weekly limit, which you share with the Claude app. Switching model doesn't help with those. The message tells you when the window resets.

Free, no card needed

Start with a single note.

It is free and no card is needed. You sign in with a link sent by email, connect your AI and start building a memory that does not stay locked inside a chat.

Create your brain, freeGo to your brainA thousand notes, as many AIs as you want to connect.