Prompt Caching for Coding Agents, Explained
Published October 9, 2026 · by the AQ team
Prompt caching is a model-provider feature that stores the processed form of a request's unchanged beginning (the prefix) so the next request that starts the same way skips reprocessing it and pays a fraction of the normal input price. It is the single biggest reason a long coding agent session costs less than its token counts suggest: every turn, a CLI like Claude Code, Codex, or Gemini CLI re-sends the entire conversation (system prompt, tool definitions, every file read and command output so far), and with caching the provider re-reads all of that at 5 to 10 percent of the list input rate. As of October 2026, Anthropic, OpenAI, and Google all cache this way, with different prices, lifetimes, and defaults, and a handful of ordinary actions inside an agent loop silently throw the cache away.
Why coding agents depend on caching more than anything else
A model API is stateless: it remembers nothing between requests, so an agent harness sends the full history on every turn, and an agent task is rarely one request. A single "fix the failing test" task might be 30 model requests; at 150,000 tokens of accumulated context, that is about 4.5 million input tokens without caching. With caching, each request bills the unchanged prefix at the cached rate and only the newly appended content at full price.
Agent sessions are unusually cache-friendly because the conversation only ever grows at the end: the prefix is byte-for-byte identical to the previous request, which is exactly what the cache matches on. It is also why the failure mode is sharp. The match must be exact, so one changed byte early in the request reprocesses everything after it at full price. Two terms of art: a cache write is the request that stores a new prefix (some vendors charge extra for it), and a cache read (or hit) is a later request reusing it at the discounted rate.
What the vendors charge, as of October 2026
Figures below are from each vendor's own pricing and documentation pages in October 2026.
Anthropic caches automatically and also supports explicit cache_control breakpoints (up to 4 per request). Writing a cache entry costs 1.25 times the base input price for the default 5-minute lifetime, or 2 times for an optional 1-hour lifetime; reads cost 0.1 times base input on most models and 0.05 times on the newest (Claude Sonnet 5.5 and Opus 5.5). Concretely, Sonnet 5.5 input is $2.00 per million tokens, a 5-minute cache write is $2.50, and a cache read is $0.10. Every read resets the lifetime clock at no charge, so an active session stays warm indefinitely.
OpenAI caches automatically on prompts of 1,024 tokens or more. On GPT-5.6 and later models (including GPT-6.1 Sol), cache reads cost 0.1 times the uncached input rate (0.05 times on GPT-6.1 Sol), cache writes cost 1.25 times, and the cache lifetime is 30 minutes. On earlier models (GPT-5.5 and before) writes are free, in-memory retention lasts 5 to 10 minutes of inactivity up to an hour, and an extended 24-hour retention tier exists: it is the default for organizations without Zero Data Retention on supported models. Concretely, GPT-5.5 input is $5.00 per million tokens and cached input is $0.50; GPT-6.1 Sol is $2.00 and $0.10. A prompt_cache_key parameter helps route parallel sessions working on the same repository to the same cache.
Google enables implicit caching by default on Gemini 2.5 and newer models, with cached input at roughly a tenth of the base rate: Gemini 3.1 Pro Preview input is $2.00 per million tokens (prompts up to 200k) and cached input is $0.20. Google also offers explicit caching, where you pay a storage fee per million tokens per hour ($4.50 for Pro-class models, $1.00 for Flash-class) to pin context for as long as you want.
| Provider (October 2026) | Cache read price | Cache write price | Lifetime | Minimum prefix |
|---|---|---|---|---|
| Anthropic | 0.1x base input (0.05x on Sonnet 5.5 and Opus 5.5) | 1.25x (5 min) or 2x (1 hour) | 5 min or 1 hour, reset on every read | 512 to 4,096 tokens by model |
| OpenAI | 0.1x (0.05x on GPT-6.1 Sol) | 1.25x on GPT-5.6 and later; free on earlier models | 30 min on GPT-5.6 and later; 5 to 10 min, up to 24h retention, on earlier models | 1,024 tokens |
| Google (implicit) | About 0.1x base input | No write charge; explicit caching bills storage per hour | Short window; explicit caches live as long as paid | 2,048 to 4,096 tokens by model |
The write premium means caching is not automatically free money: a 5-minute write pays for itself after one read, while a 1-hour write needs its longer lifetime to actually get used. For agent work, where dozens of requests reuse each prefix, the math lands overwhelmingly on caching's side.
What breaks the cache in an agent loop
Everything below causes the next request to miss part or all of the cache: one slower, more expensive turn while the new prefix is written. The vendors document the same underlying rule: the request must match the cached prefix exactly, and it is matched in layers, with tool definitions and the system prompt ahead of the conversation, so a change high in the stack invalidates everything after it.
- Switching models mid-session. Each model has its own cache, so moving a 100k-token conversation to another model reprocesses all of it at full price even though not a byte of content changed. Claude Code now warns before a model switch while the cache is still warm.
- Changing the system prompt. It sits ahead of everything, so editing it invalidates the whole cache. Harnesses differ on instruction files: Claude Code reads CLAUDE.md once at session start, so a mid-session edit does not break the cache (and also does not take effect until a restart,
/clear, or/compact). - Adding or removing tools. Tool definitions are part of the prefix, so connecting or disconnecting an MCP server mid-session can invalidate everything: settle your MCP server setup before starting long work. OpenAI's docs list
tools,parallel_tool_calls,reasoning.effort, and output format among the request fields that change the prefix; Anthropic's list includes tool definitions,tool_choice, and reasoning settings. Unstable JSON key ordering in tool results breaks matching the same way, a real bug class in homegrown harnesses. - Compaction. Summarizing the conversation to fit the context window replaces the history with a shorter one that shares no prefix with the old, so the cache rebuilds by design. Compacting at a natural break between tasks beats letting it trigger mid-task.
- Letting the cache expire. Step away longer than the lifetime and the next turn reprocesses the whole conversation at full price: the "first message after lunch is slow and expensive" effect, and the burst of uncached input that greets a resumed overnight session.
Caching and rate limits are linked
Caching changes not just the bill but how far your limits stretch, and the vendors differ sharply here. On Anthropic's API, cache reads do not count toward input-tokens-per-minute rate limits on current models: with a 2 million ITPM limit and an 80 percent cache hit rate, you can process roughly 10 million total input tokens per minute. On OpenAI's API, cached input is cheaper but still counts toward tokens-per-minute limits. For subscription CLIs the same economics surface as how fast you burn the plan window: Anthropic's Claude Code documentation describes using the 1-hour cache lifetime for the main conversation on a subscription within plan usage, dropping to 5 minutes once you are drawing on paid usage credits. If you keep hitting usage limits, check your cache hit rate first: Claude Code shows a per-session ratio under /usage, and both vendors report per-request cache fields (cache_read_input_tokens on Anthropic, cached_tokens on OpenAI). A read share that stays high means the loop is healthy; a write number that stays high turn after turn means something in the prefix is churning, and finding it is usually worth more than any model discount. The arithmetic feeding a team budget is covered in coding agent token costs for teams.
Where AQ fits
AQ is the multiplayer coding harness where engineering teams run AI coding agents like Claude Code and Codex together: shared live terminals, a code editor, and app previews, in your own cloud. Agents on AQ run as the real CLIs (Claude Code, Codex, Cursor Agent, Kimi, Grok, or plain shells) in persistent tmux sessions on your team's VM, under each person's own Claude or OpenAI login (no inference markup, no shared vendor key): the caching economics on this page stay between your CLI and your provider.
The patterns that keep caches effective are the ones AQ is built around: sessions survive a closed laptop and resume from any device, so long work proceeds as one continuous conversation, each workspace gets its own isolated git worktree so parallel agents never trample each other's checkouts, and teammates open the same workspace to watch and steer the same live session rather than re-explaining context to a fresh one. Pricing is two plans: Free is a personal sandbox for one person (AQ creates a private machine in an isolated network, nothing to install, no time limit), and Team is $50 per user per month in early access (standard $200, billed monthly), covering VMs you connect from your own cloud or a dedicated always-on AQ-managed VM, with the rate locked for your first 12 months.
Frequently asked questions
Why is my coding agent session so expensive after a break?
The cache expired. Cached prefixes live for a limited window of inactivity (5 minutes to 1 hour on Anthropic, 30 minutes on OpenAI's newest models, up to 24 hours with extended retention on earlier ones, as of October 2026), so the first request after a long break reprocesses the entire conversation at the full input rate. Every request after that is cheap again. Compacting or starting fresh at natural breaks avoids carrying a huge history into that one uncached turn.
Does switching models mid-session cost extra?
Yes. Each model has its own cache, so switching models means the next request reads the whole conversation with zero cache hits, at the new model's full input price. On a 100k-token session that single turn can cost more than dozens of normal turns. Pick the model at the start of a task and switch at task boundaries, not in the middle.
Do I need to configure prompt caching in Claude Code, Codex, or Gemini CLI?
No. As of October 2026 all three vendors cache automatically: the CLIs mark or match prefixes for you, and the discount shows up in per-request usage fields. What you control is not whether caching happens but whether you break it: model switches, mid-session tool changes, and unstable prefixes are user-side choices. Claude Code additionally exposes settings for the cache lifetime and a per-session hit ratio under /usage.
Does prompt caching help with rate limits and usage limits?
On Anthropic's API, substantially: cache reads do not count toward input-tokens-per-minute limits on current models, so a high hit rate multiplies your effective throughput. On OpenAI's API, cached tokens are cheaper but still count toward tokens-per-minute limits. On subscription plans the benefit appears indirectly, as slower burn of your plan's allowance, which is why a poor cache hit rate is a common cause of hitting limits early.
Why does a one-line question in a long session still use so many tokens?
Because the API is stateless, the harness sends the entire conversation history with that one line. With a warm cache the history bills at the cached rate (5 to 10 percent of list price as of October 2026), which is cheap but not free, and if the cache has expired it bills at full price. Long-lived sessions are worth keeping warm, and worth compacting once most of their history is no longer needed.