Cost Optimization
Claude Code is powerful, but used carelessly, costs add up fast. Once you understand token efficiency and build the right habits, you can cut costs dramatically while keeping your productivity.
Understanding the Cost Structureβ
Subscription vs APIβ
| Method | Cost model | Where to check usage |
|---|---|---|
| Max/Pro subscription | Flat monthly fee, usage included | /stats |
| API key | Billed by token usage | /cost |
| Bedrock/Vertex | Billed by the cloud provider | Cloud console |
/cost shows API token usage and cost. If you're on a Max/Pro subscription, your usage is included in the subscription, so use /stats to review your usage patterns instead.
Average Cost Benchmarksβ
- Typical developer: about $13 per active day; 90% of users stay under $30 per active day
- Monthly: $150β250 per developer (varies widely with concurrent instances and level of automation)
These are enterprise-environment figures from the official docs (code.claude.com/docs/en/costs). If you're using a subscription, these amounts have nothing to do with your bill.
Building Token Intuitionβ
- 1 English word β 1.3 tokens
- 1 Korean character β 1.5β2 tokens
- 100 lines of code β 1,000β2,000 tokens
- A typical file β 500β5,000 tokens
The Main Drivers of Cost Growthβ
1. Unnecessarily Long Conversation Sessionsβ
The longer the conversation, the more all previous messages get included as input every single time.
# When the task changes, reset the context with /clear
> /clear
# Name the session before clearing so you can come back with /resume
> /rename auth-module
> /clear
2. Overly Broad Requestsβ
# Inefficient: exploring everything
> Analyze the project structure and find problems
# Efficient: scoped
> Analyze only the error handling patterns in the src/api/ directory
3. Repeating the Same Contextβ
Instead of re-explaining the same background every time, define it once in CLAUDE.md.
4. Retries Caused by Low-Quality Promptsβ
Vague request β wrong result β correction request β repeat. Be precise from the start.
Token-Saving Strategiesβ
Strategy 1: Context Managementβ
# Separate contexts between tasks with /clear
[Feature A done]
> /clear
# Compress the context with /compact (custom instructions supported)
> /compact Summarize with a focus on the code samples and API usage
You can also set default compaction instructions in CLAUDE.md:
# Compact instructions
When compacting, focus the summary on test output and code changes
Strategy 2: Optimize Model Selectionβ
# Use Haiku for simple tasks
claude --model claude-haiku-4-5-20251001 "Convert this JSON to a TypeScript interface"
# Switch models mid-session
> /model
- Sonnet: fits most coding tasks and is cheaper than Opus
- Opus: use for complex architecture decisions or multi-step reasoning
- Haiku: best for simple conversion, classification, and formatting tasks
Setting model: haiku in a subagent configuration cuts the cost of simple tasks.
Strategy 3: Tune Extended Thinkingβ
Extended Thinking is enabled by default (31,999-token budget), and thinking tokens are billed as output tokens:
- Adjust Opus's effort level via
/modelor/effort - Disable thinking in
/config - Cap the budget with an environment variable:
MAX_THINKING_TOKENS=8000
Note that Fable 5 is the exception β thinking can't be turned off (the /config toggle and MAX_THINKING_TOKENS=0 have no effect), and it always runs adaptively with no fixed budget. On Fable 5, control thinking depth via the effort level (for a one-off deeper pass, add ultrathink to your prompt). Its pricing is also 2x Opus 4.8 ($10/$50 vs $5/$25), so it's not recommended for cost-sensitive work.
Strategy 4: Reduce MCP Server Overheadβ
Each MCP server adds its tool definitions to your context:
- Prefer CLI tools:
gh,aws,gcloud,sentry-cli, and the like are more context-efficient - Disable unused servers: review in
/mcpand turn off what you don't need - Tool Search: when MCP tool descriptions exceed 10% of the context, they're lazy-loaded automatically. Set the threshold with
ENABLE_TOOL_SEARCH=auto:<N>
Strategy 5: Isolate Verbose Output in Subagentsβ
Delegate output-heavy work β running tests, fetching docs, processing logs β to subagents. The verbose output stays in the subagent's context, and only a summary returns to the main conversation.
Strategy 6: Use Hooks and Skillsβ
- Hooks: preprocess data before Claude sees it (e.g. extract only the errors from a 10,000-line log)
- Skills: provide domain knowledge without file exploration
// settings.json β a Hook that filters test output
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "~/.claude/hooks/filter-test-output.sh"
}
]
}
]
}
}
Strategy 7: Split CLAUDE.md into Skillsβ
- Keep
CLAUDE.mdwithin ~500 lines - Move specialized instructions (PR reviews, DB migrations, etc.) into Skills
- Skills load only when invoked, so your baseline context stays small
Team Cost Managementβ
Workspace Spend Limitsβ
When using the API, you can set an overall spend limit for the Claude Code Workspace in the Console. On first authentication, a "Claude Code" workspace is created automatically, enabling centralized cost tracking.
Recommended Rate Limits by Team Sizeβ
| Team size | TPM / user | RPM / user |
|---|---|---|
| 1β5 | 200kβ300k | 5β7 |
| 5β20 | 100kβ150k | 2.5β3.5 |
| 20β50 | 50kβ75k | 1.25β1.75 |
| 50β100 | 25kβ35k | 0.62β0.87 |
| 100β500 | 15kβ20k | 0.37β0.47 |
| 500+ | 10kβ15k | 0.25β0.35 |
As teams grow, the share of concurrent users drops, so per-user TPM decreases. Rate limits apply at the organization level.
Agent Teams Token Costβ
Agent Teams run multiple Claude Code instances simultaneously, each maintaining its own context window:
- In Plan mode, roughly 7x the token usage of a regular session
- Cost management tips:
- Use Sonnet for teammates (balance of capability and cost)
- Keep the team small
- Make spawn prompts specific
- Clean up the team when the work is done
Background Token Usageβ
Claude Code uses a small amount of tokens even while idle:
- Summarizing previous conversations (for
claude --resume) - Processing commands like
/cost
Typically under $0.04 per session.
Cost vs Productivityβ
Don't sacrifice productivity by obsessing over cost savings. Even if you spend $1 on Claude Code, it's more than worth it if it saves you 10 minutes on the task.
Tasks where you should spend freely on Claude:
- Complex refactorings that would take days
- Unfamiliar tech stacks
- When you're stuck identifying the cause of a bug
- Repetitive, tedious bulk work
Tasks where you should save the cost:
- Easy work you already know how to do
- Information lookups a simple search can answer
- Code under 10 lines that's faster to write yourself
Efficient Work Habitsβ
- Use Plan mode: press
Shift+Tabto plan and explore before implementing complex work - Change direction fast: interrupt a wrong direction with
Esc; roll back with/rewindorEsc+Esc - Provide verification criteria: supply test cases, screenshots, and expected output up front
- Test incrementally: write one file β test β move on
Checklistβ
- Describe your project context thoroughly in
CLAUDE.md(within ~500 lines) - Make
/cleara habit when switching tasks - Use the Haiku model for simple tasks
- Scope requests specifically instead of asking broadly
- Disable unused MCP servers
- Move specialized instructions into Skills
- Verify team Rate Limit settings
μ΄ μ±ν°λ₯Ό μλ£νμ ¨λμ?
νμ΅ μ§λλ₯Ό 체ν¬νμ¬ λμ λ‘λλ§΅ λ¬μ±λ₯ μ λμ¬λ³΄μΈμ.