What to Do When Claude Code or Codex Is Down
Published October 5, 2026 · by the AQ team
When Claude Code or Codex stops responding, check the vendor's status page before changing anything on your machine: status.claude.com covers Claude Code and the Claude API, and status.openai.com covers Codex (the CLI, the web app, the API, and the VS Code extension). If there is no incident posted, the cause is more likely a usage limit, an expired login, or your own network than an outage. If there is an incident, nothing on your machine is broken or lost: the CLI is local software, your repository and uncommitted changes are on disk, and session history is stored locally, so you can wait it out or point a different agent CLI at the same checkout and keep working.
Check the official status pages first
Both vendors run real status pages with per-component detail, and both let you subscribe to updates so the next incident finds you instead of the other way around.
- Claude Code: status.claude.com (also reachable at status.anthropic.com). As of October 2026 it reports separately on claude.ai, the Claude Console, the Claude API (api.anthropic.com), Claude Code, and other surfaces, with incident states from degraded performance to major outage, and offers email, SMS, Slack, and RSS subscriptions.
- Codex: status.openai.com. As of October 2026 the Codex section lists four components (Web, API, CLI, VS Code extension) alongside the ChatGPT and API sections, with a subscribe option at the top.
Real incidents happen to both. On September 25, 2026, OpenAI's status page recorded a full Codex outage across all four components that lasted about 56 minutes; one mid-incident update noted that logging in with an API key would unblock access while ChatGPT-account auth was affected. On September 29, 2026, Anthropic's status page reported elevated errors across claude.ai, Claude Code, and the Claude API for about an hour. Short, sharp, and fully recovered is the common shape, which is why the right first move is reading the status page, not reinstalling anything. Crowd trackers such as Downdetector aggregate user reports, so treat them as a second opinion, never the diagnosis.
Outage or usage limit? They look similar until you read the message
The most common "is it down?" false alarm is a usage limit: the CLI stops making progress, but the service is fine. A usage limit is a cap built into your subscription plan, and it announces itself if you look.
Claude Code meters a rolling 5-hour session window plus weekly limits on Pro and Max plans, and when you hit one it says so in the session, including when the window resets. The /usage command in the CLI draws the progress bars for both windows. A message that names a limit and a reset time is not an outage, and the fixes are different: wait for the reset, spend the window differently, or move the work. The Claude Code usage limits guide covers how the windows work and what burns them fastest.
Codex draws on your ChatGPT plan's limits. As of October 2026, OpenAI documents Codex usage in a 5-hour window plus a weekly window, with how far each message goes depending on the model and task size; the /status command in the Codex CLI shows where you stand. A limit message with a reset time means your plan is spent, not that Codex is down.
API users get the distinction in the status code. On the Claude API, a 429 means you hit your rate limit, a 529 (overloaded_error) means Anthropic's systems are temporarily overloaded across users, and a 500 is an internal error; Anthropic's docs say to retry 529s and 500s with exponential backoff (increasing the wait between attempts). A burst of 529s or 500s across the board is outage-shaped; a 429 is about your account.
One more non-outage to rule out: authentication. An expired or revoked login produces errors that read like the service is broken. If the status page is green and you are not at a limit, sign out and back in before concluding anything.
| Symptom | Likely cause | What to do |
|---|---|---|
| Status page shows an incident | Vendor outage | Subscribe to the incident, switch CLIs or wait; change nothing locally |
| Message names a limit and a reset time | Usage limit | Wait for the reset or move the work; see the limits guide |
| 429 errors (API) | Your rate limit | Back off, batch requests, or raise your tier |
| 529 or 500 errors (API), status page green | Transient overload | Retry with exponential backoff; re-check the status page |
| Auth or sign-in errors, status page green | Expired login | Log out and log back in |
| Only your machine fails, browser apps work | Local network, VPN, or proxy | Check connectivity and proxy settings before blaming the vendor |
What an outage actually takes down (and what it does not)
Claude Code and Codex are local programs that call a hosted model API. When the vendor has an outage, the half on your machine is untouched:
- Your working tree. Every file the agent edited is on disk; the uncommitted diff is one git command away.
- Your git state. Branches, commits, and stashes are local. An API outage cannot lose committed work.
- Your session history. Both CLIs store transcripts locally and can reopen a past session once service returns (claude with the resume flag, codex with its resume command, as of October 2026). Resuming and searching Claude Code sessions covers the Claude side.
So the honest worst case of a one-hour outage is an hour of waiting, not lost work. The practical question is whether to wait at all.
Switching CLIs against the same worktree
The repository checkout is the real shared state, and nothing about it belongs to one vendor. If Claude Code is down and Codex is up (or the reverse), open the other CLI in the same directory and continue:
# In the same checkout the stalled agent was working in
git status # see what the previous agent left behind
git diff # the uncommitted work so far
git log --oneline -10
# Start the other CLI in that directory and ask it to catch up:
codex # or: claude
# "Read the uncommitted diff and recent commits, then continue: <the task>"
Two things make the handoff smooth. First, shared instructions: AGENTS.md is a plain markdown file in the repository root that tells coding agents how to build, test, and behave in that repo, and as of October 2026 both sides read it. Codex reads AGENTS.md natively, and Claude Code reads AGENTS.md when the repository has no CLAUDE.md (since v2.1.277, September 2026) or through a one-line import from CLAUDE.md, per Anthropic's docs; keeping AGENTS.md and CLAUDE.md in sync shows the patterns. Second, context that lives in files instead of chat: git history and the diff carry most of it, and carrying context between Claude Code and Codex covers rebuilding the rest. The CLIs do not share limits, so the stalled vendor's plan is irrelevant to the fallback.
Make the next outage boring
- Keep two CLIs installed and logged in. The switch above takes seconds only if the fallback CLI already has credentials.
- Put repo instructions in AGENTS.md. A file serves whichever agent shows up; instructions typed into one vendor's chat do not.
- Make the agent write plans into the repo. A markdown plan committed alongside the code lets any CLI (or human) continue mid-task. Executable specs takes this further.
- Commit early on agent branches. A work-in-progress commit turns "what was it doing?" into git log.
- Subscribe to both status pages. Knowing it is an outage within a minute saves twenty spent debugging your own setup.
- Isolate parallel tasks in worktrees. One stalled task should not block the rest; git worktrees keep them independent.
Where AQ fits
AQ is the multiplayer coding harness where engineering teams run AI coding agents like Claude Code and Codex together: shared live terminals, a code editor, and app previews, in your own cloud. An outage at Anthropic or OpenAI still pauses that vendor's CLI inside AQ, because agents run under your own CLI logins (your Claude and OpenAI accounts, with no markup on model usage). What AQ changes is everything around the stalled CLI: agents run as real CLIs in persistent tmux sessions on your team's VM, so the workspace, the worktree, and the session survive the incident, your laptop closing, or both, and resume from any device.
The switch this guide describes is one tab in AQ. Each workspace is an isolated git worktree, and Claude Code, Codex, Cursor Agent, Kimi, Grok, and plain shells all open against it, so when one vendor is down you start another CLI beside the stalled one in the same checkout, with the same files and git state in front of it. Teammates can open the workspace and see exactly where things stopped instead of asking, and the live dev-server preview keeps serving while the model API is away.
Pricing is two plans: Free is a personal sandbox for one person (AQ creates a private machine in an isolated network, nothing to install, no time limit), and Team is $50 per user per month early access (standard $200, billed monthly), covering VMs you connect from your own cloud or a dedicated always-on AQ-managed VM, with the rate locked for your first 12 months.
Plainly: AQ cannot make a model outage shorter. It makes the workaround (another CLI, same worktree, nothing lost, team aware) the default instead of a scramble.
Frequently asked questions
How do I check if Claude Code or Codex is down right now?
Open the vendor's official status page: status.claude.com reports Claude Code as its own component alongside claude.ai and the Claude API, and status.openai.com lists Codex's Web app, API, CLI, and VS Code extension separately, as of October 2026. Both offer subscriptions to incident updates. If the page is green, suspect a usage limit, an expired login, or your own network instead.
Claude Code says I hit a usage limit. Is that an outage?
No. A usage limit is your plan's cap, not a service failure: Claude Code meters a rolling 5-hour session window plus weekly limits, and the limit message includes when your window resets. Run /usage to see both progress bars. An outage shows up as errors and an incident on the status page; a limit shows up as a named limit with a reset time. The usage limits guide explains the windows.
Do I lose work if Claude Code or Codex goes down mid-task?
No. The CLI runs on your machine; only the model API is remote. Every edited file is on disk, your git branches and commits are local, and both CLIs keep session transcripts locally, so you can resume the conversation once service returns. The practical cost of an outage is waiting time, not lost work.
Can I switch to Codex while Claude Code is down (or the reverse)?
Yes, and the checkout makes it easy: open the other CLI in the same directory, have it read git diff and recent commits, and continue the task. As of October 2026 both read AGENTS.md (Codex natively; Claude Code when the repo has no CLAUDE.md, or via an import), so repo instructions carry over. Each CLI bills its own vendor's plan, so the stalled vendor's limits are irrelevant to the fallback.
Does upgrading my plan help during an outage?
No. Upgrading raises usage limits, which is the fix for limit lockouts, not for outages: during an incident every plan is equally affected. If the status page shows an incident, the useful moves are subscribing to the incident, switching to another CLI against the same checkout, or waiting; spending money changes nothing until the vendor resolves it.