Why Is Claude Code Slow? Official Causes and Developer Fixes

By Aakash Ahuja··22 min read

Why Is Claude Code Slow? Official Causes and Developer Fixes

If Claude Code has felt slow or degraded, you are not imagining it, and the cause is not a single bug. Claude Code does not simply "get worse" in one clean way. Developer complaints mix together several different problems: lower coding quality, slower-feeling sessions, forgetfulness, repeated tool choices, faster quota drain, actual service errors, and local context bloat.

The problem with most advice on this is that it has no baseline. "Claude Code is slow" is meaningless without knowing what normal looks like. So I measured it: 350 session files, 74,493 assistant turns, 35 days, read straight out of the transcripts Claude Code writes to ~/.claude/projects/. Those baselines are in the next section, and the rest of the article is a runbook for deciding which bucket your problem is in.

There is also a documented case of Anthropic breaking it themselves, which is worth understanding because it teaches you what the failure modes look like: Anthropic's April 23, 2026 engineering postmortem described three product-layer changes that degraded Claude Code, all resolved by April 20 in version 2.1.116. That is history now, but the diagnostic patterns it exposed still apply.


What a normal Claude Code session actually looks like

Before you conclude something is wrong, check your experience against measured behaviour. These are medians from 74,493 turns on client versions 2.1.187 through 2.1.222.

MeasureNormal range
Turn wall time, fresh session (under 25K context)1.1s median, 3.9s p90
Turn wall time, mid session (~150K context)3.5s median, 19.2s p90
Turn wall time, very large context (500K+)5.0s median, 25.5s p90
File read0.01s
Shell command0.12s median, 7.15s p90
Web search9.33s median
/compact at ~420K context123s median, 198s worst case
Three calibration points that resolve most "is this normal?" questions:

A 5x slowdown across a long session is expected. Turn time roughly quadruples from an empty session to 300K tokens of context, then flattens. If you are seeing that, nothing is broken.

Occasional 20-second turns are normal at every context size. The p90 sits at 19 to 22 seconds whether your context is 50K or 700K. A slow turn is not evidence of a problem.

Two minutes of nothing during /compact is normal. It is doing real work, and at large context it routinely takes over 90 seconds.

If your experience is materially worse than these numbers, work through the runbook below.


Short answer: what actually causes a slow or bad session

A slow or bad Claude Code session comes from one of these, roughly in order of how often it was the real answer in my own data:

  • A stale or bloated session context.
  • Large files or tool output filling the context window, which is Claude Code's working memory for the current session and grows with every tool call and turn.
  • Your own network. This is far more common than people think: of 375 API errors I logged across 35 days, 325 were local connection failures (ECONNRESET, ENOTFOUND). Claude Code retries these silently up to 10 times, so they present as unexplained stalls.
  • A reasoning effort setting that is wrong for the task in either direction.
  • Local environment problems, especially search or filesystem issues.
  • An outdated Claude Code version.
  • Account quota or API rate limits. Measurably rare: 12 rate-limit errors and 6 overload errors across 74,493 turns of heavy use.
  • Anthropic-side incidents, which do happen but were the least common cause in my data.
Note the ordering. Most developers check the status page first and their own network last. The data says to reverse that.



Jump to: Short answer · What Anthropic said · Classify the failure · 8 fixes · Diagnosis matrix · Common mistakes · FAQ · Takeaways


What did Anthropic officially say happened to Claude Code?

Anthropic's postmortem identified three separate changes. The important point for developers is that these were not all "latency bugs." Some affected answer quality, some affected memory-like continuity inside a session, and some affected token/cache behavior.

Reasoning effort changed from high to medium

Anthropic said it changed Claude Code's default reasoning effort from high to medium on March 4, 2026 for the then-current Sonnet and Opus models. The reason was latency: some users on high effort were seeing very long thinking times, enough that the UI could appear frozen. Anthropic later called this the wrong tradeoff and reverted it on April 7.

In Claude Code, reasoning effort controls the tradeoff between capability, latency, and token usage. The current ladder is low, medium, high, xhigh, and max. Anthropic's guidance is that low suits short, scoped, latency-sensitive tasks; medium reduces token usage while trading off some intelligence; high is the default and the minimum for intelligence-sensitive work; xhigh is the recommendation for most coding and agentic work; and max is for cases where correctness matters more than cost.

Worth knowing before you reach for the top of the ladder: on current Claude models, the lower settings are unusually strong. low and medium on the Claude 5 family often match what xhigh produced on the previous generation. In my own sessions 78% of turns ran at high, which was the wrong default for a lot of routine work and cost both time and tokens.

For developers, this means a "slow" experience can have two opposite causes:

  • The model is thinking more, producing better output but taking longer.
  • The model is thinking less, responding faster but producing weaker code.
If the complaint is "Claude is slower," check latency. If the complaint is "Claude is dumber," check effort level.


A caching optimization dropped older reasoning repeatedly

Anthropic said it shipped a March 26 optimization intended to reduce latency when users resumed sessions that had been idle for more than an hour. The intended behavior was to clear older thinking once after a session became stale. A bug caused older thinking to be cleared on every turn for the rest of the session.

That matters because Claude Code relies on conversation history, prior tool calls, prior edits, and previous reasoning to continue a multi-step coding task. Anthropic said the bug made Claude appear forgetful, repetitive, and prone to odd tool choices. It also caused cache misses, which Anthropic believes drove reports of usage limits draining faster than expected. (Postmortem)

For developers, this is the key lesson:

If a Claude Code session has become stale, repetitive, confused, or expensive, do not keep arguing with it inside the same thread. Compact, clear, rewind, or start a new focused session.

Claude Code's own docs make the same operational point from another angle: as context fills up, Claude Code clears older tool outputs first and summarizes conversation history if needed, but detailed instructions from early in the conversation can be lost. Persistent rules should go in CLAUDE.md, not only in the chat history.


A system prompt change reduced coding quality

Anthropic said it added a system prompt instruction on April 16 to reduce verbosity. One instruction limited text between tool calls and final responses. After broader ablation testing, Anthropic found a measurable drop in an evaluation and reverted the prompt as part of the April 20 release.

This is the most important product-design lesson from the incident: shorter answers are not automatically better answers for coding agents.

Coding work often needs enough reasoning, enough plan context, and enough explanation between tool calls to avoid shallow edits. The better instruction is:

"Be concise, but do not omit reasoning needed to make safe code changes."

That phrasing preserves the quality constraint.


Is Claude Code slow, overloaded, or just in a bad session?

Before changing settings, classify the failure.

SymptomLikely causeWhat to check first
Claude responds slowly but eventually gives strong answersHigh effort, large context, long tool run, or large output/effort, /context, task size
Claude responds quickly but makes shallow mistakesEffort too low or weak prompt/effort, model selection
Claude forgets earlier choices or repeats itselfStale/bloated session or context compaction issue/context, /compact, /clear, /rewind
Claude burns usage faster than expectedLarge context, cache misses, too many tool calls, agent teams/usage, /context, session structure
Claude shows 529 overload errorsAnthropic capacity issueClaude status, retry, switch model
Claude shows 429 errorsRate limit or credential/provider limit/status, provider console, concurrency
Search misses files or is slow on WSLFilesystem/search issueProject location, ripgrep, WSL filesystem
Claude Code crashes or resume failsVersion-specific bugclaude --version, changelog, update
Anthropic distinguishes infrastructure problems from account/request problems. Repeated 529 overload errors mean the API is temporarily at capacity across users and are not your usage limit; the recommendation is to check status and switch models if one model is under high load. A 429, by contrast, means a configured rate limit has been hit for your API key or provider project.


How to fix Claude Code slowness as a developer

1. Update Claude Code and check your version

Start with version sanity.

claude --version
claude update

This matters because some issues are version-specific. The Claude status page has reported crashes when resuming prior sessions with --resume or --continue in specific versions.

Do not debug prompts before you eliminate a bad client version.


2. Check the active model

Inside Claude Code, run:

/model

Confirm you are using the model you intended. A previous /model choice or environment variable may have selected a smaller model. Also check whether your plan supports the model you selected.


3. Raise or verify reasoning effort

Inside Claude Code, run:

/effort

For difficult debugging, multi-file refactors, architecture decisions, concurrency bugs, security-sensitive changes, or test-repair loops, do not run on low or medium effort unless you are deliberately optimizing for speed.

You can set effort using /effort, the /model picker, --effort, CLAUDE_CODE_EFFORT_LEVEL, settings, or skill/subagent frontmatter. The environment variable has the highest precedence. (Effort docs)

Practical defaults:

Task typeSuggested effort
Rename variable, small copy edit, simple commandlow or medium
Normal coding taskhigh (the default)
Multi-file debugging or agentic workxhigh
Hard architecture/debugging sessionxhigh or one-off ultrathink
Cost-sensitive bulk scriptingmedium, but verify output carefully
Do not blindly maximize effort. Anthropic says max can help demanding tasks but may show diminishing returns and is prone to overthinking.

Effort is also the most underused speed control, and the reason is not obvious. Lower effort does not only shorten thinking; it produces fewer and more consolidated tool calls. Since each tool call costs a full model turn (2 to 4 seconds in a typical session), a setting that removes five exploratory calls saves more wall-clock time than any amount of faster thinking. If a session feels sluggish on simple work, try lowering effort before you try switching model, since switching model mid-session also discards your prompt cache.


4. Use ultrathink only for hard one-off tasks

For a single hard turn, include:

ultrathink

Claude Code recognizes ultrathink as a one-off request for deeper reasoning without changing the session effort setting. Other phrases like "think hard" are passed through as normal prompt text and are not recognized as special keywords.

Use it for:

  • "Find the real root cause across these logs."
  • "Design the migration plan before editing code."
  • "Review this authentication flow for security bugs."
  • "Explain why the test passes locally but fails in CI."
Do not use it for every prompt.


5. Clear or compact stale context

Check context:

/context

Compact at a natural breakpoint:

/compact focus on the current bug, files changed, test output, and next steps

Clear when switching tasks:

/clear

Claude Code docs recommend /clear when switching to unrelated work because stale context wastes tokens on every later message. For a deeper look at how context accumulates mechanically over a session, see Why Claude Code gets slower the longer you use it. For long coding sessions, follow this rule:

One task, one session. New task, new context.

If you are debugging authentication middleware, do not keep the same Claude Code session alive when you move to frontend CSS, then database migrations, then CI config. That creates context pollution.


6. Reduce oversized files and tool output

If Claude Code reads giant files, dependency directories, generated build output, test logs, minified bundles, or massive JSON files, your context fills quickly.

Better prompt:

Read only src/auth/jwt_validator.py lines 80-180.
Do not scan the whole repository yet.
Identify why expired tokens are passing validation.

Worse prompt:

Search the whole repo and fix auth.

For auto-compaction thrashing, ask Claude to read oversized files in smaller chunks, use /compact with a focused instruction, move large-file work to a subagent, or run /clear if earlier conversation is no longer needed.


7. Diagnose local performance problems

Run:

/doctor

/doctor checks installation health, settings validity, MCP configuration, and context usage. If Claude Code will not start, run claude doctor from the shell. (Troubleshooting docs)

If search is slow or incomplete on WSL, filesystem read penalties across Windows/Linux boundaries may reduce search results. The recommended fixes are more specific searches, moving the project to the Linux filesystem under /home/, or running Claude Code natively on Windows.

If file discovery is broken, install ripgrep and configure Claude Code to use the system version:

brew install ripgrep          # macOS
sudo apt install ripgrep      # Ubuntu/Debian
winget install BurntSushi.ripgrep.MSVC  # Windows

Then set:

export USE_BUILTIN_RIPGREP=0

8. Check your network first, then Anthropic status

Conventional advice says to check the Claude Status page before rewriting prompts or reinstalling tools. That is correct but incomplete, because it skips the far more likely culprit.

Here is the actual error distribution from 74,493 turns over 35 days of heavy use:

Error typeCountShare
Local connection failure (ECONNRESET, ENOTFOUND)32586.7%
5xx server errors (500, 502, 503)215.6%
429 rate limited123.2%
529 overloaded61.6%
Nearly nine out of ten errors were my own network. Rate limiting and capacity problems together accounted for 18 events in five weeks of very heavy usage.

This matters diagnostically because Claude Code retries transient failures up to 10 times with exponential backoff before surfacing anything. So a dropped connection does not look like an error. It looks like Claude going quiet for 20 or 30 seconds and then continuing normally, which is exactly what people describe when they say Claude Code "feels slow."

ErrorMeaningDeveloper action
Silent stall, then resumesAlmost certainly a local connection drop being retriedCheck wifi, VPN, corporate proxy
529 overloadAnthropic capacity issueCheck status, retry later, switch model
429 rate limitYour API/provider/account limitCheck /status, reduce concurrency, request higher limits
Session/weekly limitSubscription quota exhaustedCheck /usage, wait for reset, buy or request extra usage
TimeoutLarge response, high load, network/proxy issueRetry, split prompt, raise timeout only if network or proxy is the issue
If you see an error message at all, automatic retries have already been exhausted, which makes it a genuinely different signal from a stall.


Claude Code slow-response diagnosis matrix

Use this as a practical runbook.

StepCommand/checkWhat it tells youAction
0Your own networkWhether the stalls are local, which they usually are87% of my logged API errors were local connection drops. Check this before the status page
1Claude Status pageWhether Anthropic has active incidentsIf degraded, switch model or wait
2claude --versionWhether your CLI may be outdatedRun claude update
3/modelWhether you are on the intended modelSwitch to expected model
4/effortWhether reasoning effort is too low/highRaise for hard work, lower for simple work
5/contextWhether session context is bloatedCompact, clear, or start fresh
6/usageWhether quota is near exhaustionReduce context, wait, or buy/request extra usage
7/doctorWhether install/config/MCP/context has issuesFix reported local issues
8/feedbackSends reproducible issue to AnthropicUse when issue persists after checks
---

Common mistakes developers make when Claude Code feels worse

Mistake 1: Treating quality degradation as a prompting problem only

Sometimes the problem is not your prompt. Anthropic's April 23 postmortem confirms that product-layer changes caused real quality issues for some users.

Still, do not jump from "Claude made a bad edit" to "the model is broken." First check model, effort, context, version, and status.


Mistake 2: Correcting bad output repeatedly in the same thread

If Claude makes a bad turn, replying with corrections can keep the bad attempt in context and anchor later answers.

Anthropic's error reference recommends rewinding when a response goes wrong. Press Esc twice or run /rewind, then rephrase the prompt with more specifics.

Better:

/rewind

Then:

Focus only on the failing test in tests/test_auth.py.
Do not edit production code yet.
First explain the failure path.

Mistake 3: Letting huge context accumulate across unrelated tasks

Long-running sessions feel convenient, but stale context costs tokens and can degrade relevance. Claude Code docs explicitly recommend /clear when switching to unrelated work.

A clean session is often faster than a heroic compacted session.

The proactive habit that prevents the problem: have the agent write its plan and decisions to a file (todo.md, DECISIONS.md) as it works, then /compact or /clear and tell it to re-read that file to continue. Because /compact summarizes and drops detail, write state to disk before you compact, otherwise the part of the context that mattered can disappear into the summary.


Mistake 4: Using Opus or high effort for every task

High effort and stronger models are not always the right default for every turn. For small edits, they can be slower and more expensive than needed.

Use effort as a control knob:

  • Simple task: lower effort.
  • Hard reasoning task: higher effort.
  • One hard turn: ultrathink.
  • Long-running task: monitor context and usage.
---

Mistake 5: Ignoring WSL and filesystem penalties

If Claude Code search is weak or slow on WSL, the issue may be filesystem placement, not model quality. Anthropic specifically recommends moving the project to the Linux filesystem under /home/ rather than working across /mnt/c/ when WSL filesystem penalties affect search.


Frequently Asked Questions About Claude Code Slowness

Why is Claude Code slow?

Claude Code can be slow because the model is using deeper reasoning, the session context is large, tool calls are expensive, the service is under load, or your local environment is slowing search and file access. Recent quality complaints were also tied to three Anthropic product-layer changes that were resolved by April 20, 2026.

Did Anthropic intentionally degrade Claude Code?

Anthropic said it does not intentionally degrade its models and said the API and inference layer were unaffected in the April 23 postmortem. The official explanation was three product-layer issues: effort default changes, a caching bug, and a system prompt change.

What should I check first when Claude Code feels worse?

Check /model, /effort, /context, claude --version, and the Claude Status page. If those are normal, run /doctor and consider whether your session has stale instructions, oversized files, or too much previous tool output.

Should I use high effort or xhigh effort in Claude Code?

Use high or xhigh for intelligence-sensitive coding tasks, debugging, refactoring, and agentic work. xhigh is the recommendation for most coding and agentic work on current Opus-tier models, while medium trades off some intelligence for lower token usage.

The direction people get wrong is upward. On the Claude 5 family, low and medium are considerably stronger than the same settings on earlier models, and they produce fewer exploratory tool calls, which is where most of the wall-clock time goes. In my own 74,493 turns, 78% ran at high, and a good share of that was routine work that did not need it.

How slow is too slow for Claude Code?

Use these measured medians as a baseline: 1.1 seconds per turn in a fresh session, 3.5 seconds at around 150K tokens of context, 5 seconds past 500K. Roughly one turn in ten will take 20 seconds at any context size, so a single slow turn means nothing. Sustained turns well beyond these numbers, or long silences that resolve on their own, point at your network rather than the model.

What is the difference between /compact and /clear?

/compact summarizes and preserves the useful parts of the current session, while /clear starts fresh. Use /compact when you are continuing the same task; use /clear when switching to unrelated work.

Why is Claude Code using my quota faster than expected?

Quota can drain faster when context is large, tool calls are repeated, cache misses happen, or agent teams/subagents are used heavily. Anthropic's postmortem said the March 26 caching bug likely drove some reports of usage limits draining faster than expected.

What does a 529 error mean in Claude Code?

A 529 overload error means the API is temporarily at capacity across users. It is not your usage limit and does not count against your quota; recommended actions are checking status, retrying later, or switching models.

What does a 429 error mean in Claude Code?

A 429 means you hit the configured rate limit for your API key, Amazon Bedrock project, or Google Vertex AI project. Recommended actions are checking /status, reviewing provider limits, reducing concurrency, or switching to a smaller model for high-volume scripted runs.


Key Takeaways

  • Know the baseline before you diagnose. Measured across 74,493 turns: 1.1s per turn in a fresh session, 3.5s at 150K context, 5s past 500K, with roughly one turn in ten taking 20 seconds at any size. A 5x slowdown over a long session is normal, not a fault.
  • Check your own network before the status page. 325 of my 375 API errors were local connection drops. Claude Code retries them silently, so they look like the model stalling.
  • Rate limits and overload are rare. 12 rate-limit and 6 overload errors across 35 days of heavy use.
  • /compact takes about two minutes at large context, and it is doing real work. That silence is not a hang.
  • "Slow" is ambiguous. Separate latency, lower answer quality, stale context, quota drain, rate limits, and service overload.
  • For hard coding work, verify /model and /effort before blaming prompts. Effort is also the most underused speed control, because lower effort means fewer tool calls and each tool call costs a full model turn.
  • Claude Code's April 2026 degradation was not one bug. Anthropic identified three product-layer issues: effort defaults, caching and context handling, and a system prompt change. All were resolved in 2.1.116.
  • For stale or confused sessions, use /context, /compact, /clear, or /rewind.
  • For local issues, run /doctor, check WSL filesystem placement, and install/use system ripgrep if search is broken.
  • For 529 errors, check status and switch models. For 429 errors, check credentials, provider limits, and concurrency.
  • Do not keep unrelated development tasks in one long Claude Code session.
---


Related reading

---

Use this checklist before opening a bug report or rewriting your workflow:

1. Check Claude Status.
  • Run claude --version.
  • Run claude update.
  • Check /model.
  • Check /effort.
  • Check /context.
  • Compact or clear stale sessions.
  • Run /doctor.
  • Use /feedback only after reproducing the issue with details.
AITutorialsMay 9, 2026
Share
Aakash Ahuja

Aakash Ahuja

Enterprise AI, Cybersecurity & Platform Engineering

Aakash writes about secure AI agents, microservices architecture, enterprise platforms, and production engineering. He has 20+ years of experience building and operating software systems across banking, cloud, cybersecurity, AI, and enterprise workflow automation. He is Director of Technology at itmtb Technologies and teaches AI, Big Data, and Reinforcement Learning at top institutes in India.