Episode 7: Tokens and limits
What a token is, why every turn of an agent session sends the whole conversation again, how caching makes that affordable, what the context window holds, and how a plan's five-hour and weekly limits work, in Claude Code, Codex CLI and Antigravity CLI.

Downloads
AI for Research Efficiency, episode 7, with Douglas Hutchings. Narration: an AI-generated voice (ElevenLabs).
What a token is, why a long session costs more, and how to make a monthly plan last, in Claude Code, Codex CLI and Antigravity CLI. Plans, limits, prices and windows as of 4 October 2026, from Anthropic's, OpenAI's and Google's own pages (sources below), and from sessions recorded that day in Claude Code 2.1.289 and Codex 0.160.0. These change often: check each vendor's page before you rely on a number.
1. What a token is
A token is a piece of a word: a common word is one token, a rare name several. Each company splits text its own way, and each gives its own rule of thumb:
| OpenAI | Anthropic | ||
|---|---|---|---|
| A token, roughly | about 4 characters; 100 tokens, about 75 words | 1M tokens, about 555,000 words (its current tokenizer) | about 4 characters; 100 tokens, 60 to 80 words |
In the film, "Summarize the methane results in the Minamikawa paper." is 8 words and 13 tokens in tiktoken, OpenAI's open-source tokenizer (encoding o200k_base): Summ ar ize the methane results in the Min amik awa paper . Claude and Gemini split the same line differently.
What the agent reads is input; what it writes is output, and its thinking is billed as output in all three.
2. Every turn sends it all again
The model remembers nothing between turns, so the tool sends the whole conversation again each time: its own instructions and tools, your files, every answer so far. Anthropic: "The model doesn't remember anything between requests, so Claude Code re-sends the full context". OpenAI: "Every time you send a new message to an existing conversation, the conversation history is included as part of the prompt for the new turn". Antigravity's own example: asked to reply with one word, its agent read 30,384 input tokens; the second turn sent 278 new tokens and read 30,214 from the cache.
A cache makes that affordable: text sent before is read again at a tenth of the input price, or less. Prices per million tokens, at list prices, on 4 October 2026:
| Claude Opus 5.5 (Claude Code's default) | GPT-6.1 Sol (Codex's default) | Gemini 3.8 Flash (Antigravity's agents) | |
|---|---|---|---|
| Input | $4 | $2.00 | $0.75 |
| Read from the cache | $0.20 (0.05 times input) | $0.10 (0.05 times input) | $0.075, plus $0.50 per million tokens per hour of storage |
| Written to the cache | $5 (kept 5 minutes) or $8 (1 hour) | $2.50 | no write charge listed for implicit caching |
| Output, thinking included | $20 | $10.00 | $3.75 |
| Notes | up to 272K input; more above that | through 31 December 2026; doubled from 1 January 2027 |
Episode 3's run, priced (from its own log; Claude Code's estimate at list prices; it ran on a monthly plan, so nothing was charged and it drew on the plan's limits):
| Part | Tokens | Cost at list prices |
|---|---|---|
| New input | 916 | $0.004 |
| Read from the cache | 31,021,208 | $6.20 |
| Written to the cache | 2,085,946 | $10.87 |
| Output (78,510 of it thinking) | 391,000 | $7.82 |
| In all | 33,499,070 | $24.89 |
94% of the input came from the cache. Without it, the same run would have cost about $140, more than five times as much.
What starts the cache over: in Claude Code, switching models, changing the effort level (on most models), fast mode, and compacting; in Codex, compacting or summarizing; in any tool, a change near the start of the conversation.
3. The context window
How much the model can hold at once:
| Claude Code | Codex CLI | Antigravity CLI | |
|---|---|---|---|
| Default model and window | Opus 5.5: 1,000,000 tokens, on every plan | GPT-6.1 Sol: 272,000 tokens, of which Codex uses 95% (/status shows 258K) | Gemini 3.8 Flash: 1,048,576 tokens |
| When it fills | compacts by itself, at about 967K by default | compacts by itself | compacts by itself |
| Compact now | /compact | /compact | none (automatic only) |
| Start fresh | /clear (it costs nothing) | /clear, or /new | /clear |
| See what fills it | /context | /status | /context |
Before you type a word, the tool's own instructions and tools are already there: in the recorded Claude Code session, 31,300 tokens (3% of the window; that account has extra tools and skills installed, so a new one starts smaller). One question about one paper took it to 88,100 (9%).
4. A plan's limits
Most plans meter use in a five-hour window with a weekly cap on top, and every terminal, helper agent and cloud task on the account draws on the same allowance. None of the three publishes its allowance as a number of tokens.
| Claude (Claude Code) | ChatGPT (Codex) | Google (Antigravity) | |
|---|---|---|---|
| Plans | Pro $20 a month ($17 a month billed yearly); Max from $100 (5x or 20x Pro's usage) | Go $8; Plus $20; Pro $100, $200 or $500 | Free; Google AI Pro $19.99; Ultra $100 and $200 |
| Windows | a session limit that resets every five hours, and a weekly limit | a five-hour limit (Pro has none) and weekly limits | Pro: every five hours until the weekly limit; free: weekly |
| At the limit | waits, then carries on by itself ("Continuing automatically at 3:45pm · esc to cancel") | the running turn can finish; then wait, or buy credits | waits for the quota, unless AI credits are allowed (a setting) |
| Extra usage (paid) | usage credits, at API rates, when you turn them on | credits | AI credits, Pro and Ultra only |
A plan or a key: with an API key instead of a plan, every token is billed at API rates; in Claude Code a set ANTHROPIC_API_KEY is used instead of your subscription.
Fast modes spend faster: Claude's fast mode is paid usage credits only, outside the plan's limits; Codex's Fast mode uses the plan's limits at 2.5 times the standard rate.
5. See your usage
| Claude Code | Codex CLI | Antigravity CLI | |
|---|---|---|---|
| Plan limits | /usage | /status | /usage (or /quota) |
| Token history | /usage (Stats) | /usage | not documented |
| What fills the context | /context | /status | /context |
The trap: /usage shows plan limits in Claude Code and Antigravity, but your token history in Codex, whose limits are in /status. Codex's /usage menu also offers to redeem a rate-limit reset; choose View analytics to look only.
6. Make it last
- One task per session, and
/clearbetween tasks. - A smaller model, or lower effort, for simple jobs.
- Let the agent open the files it needs, rather than pasting everything in.
- For heavy reading, helpers: they hand back summaries.
- Before a big run, ask for an estimate.
- Before a deadline, check your usage.
7. Many agents multiply it
- Anthropic (June 2025): "agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats". Claude Code's documentation: agent teams use "approximately 7x more tokens than standard sessions when teammates run in plan mode".
- Google (May 2026): an operating system built by 93 subagents, "over 2.6B tokens", $916.92 at API prices.
- OpenAI publishes no multiplier: "subagent workflows consume more tokens than comparable single-agent runs".
Episode 8 is about running many agents at once, safely.
Not confirmed yet: how many tokens a plan's windows hold (no vendor publishes it), how cached tokens count against Claude's and Codex's windows (both say only that caching lowers usage), and Antigravity's usage panel as it appears when signed in (the film draws it from its documentation).
Try it
- Open your agent.
- Type
/usage(in Codex,/status). - Read the two windows: how much of the five hours and of the week is used, and when each resets.
Sources
Read on 4 October 2026 (copies kept with the project). The recorded screens come from Claude Code 2.1.289 and Codex 0.160.0 on the studio's server; Antigravity's from its documentation; episode 3's run from its own log.
- Anthropic (Claude Code, the Claude API, Claude's plans): Claude API docs: Glossary (opens in a new tab); Claude API docs: Models overview (opens in a new tab); Claude API docs: Pricing (opens in a new tab); Claude API docs: Prompt caching (opens in a new tab); Claude plans and pricing (opens in a new tab); How we built our multi-agent research system (2025-06-13) (opens in a new tab); Claude Help Center: Using Claude Code with your Pro or Max plan (opens in a new tab); Claude Help Center: What is the Max plan? (opens in a new tab); Claude Help Center: What is the Pro plan? (opens in a new tab); Claude Help Center: Usage credits for paid Claude plans (opens in a new tab); Claude Help Center: Usage limit best practices (opens in a new tab); Claude Code docs: Run agents in parallel (opens in a new tab); Claude Code docs: Commands (opens in a new tab); Claude Code docs: Manage costs effectively (opens in a new tab); Claude Code docs: Error reference (opens in a new tab); Claude Code docs: Fast mode (opens in a new tab); Claude Code docs: Interactive mode (opens in a new tab); Claude Code docs: Model configuration (opens in a new tab); Claude Code docs: How Claude Code uses prompt caching (opens in a new tab); Claude Code docs: Customize your status line (opens in a new tab); Claude Code docs: Dynamic workflows (opens in a new tab).
- OpenAI (Codex, the OpenAI API): Codex docs: Pricing (the plan cards, rendered in a browser) (opens in a new tab); Codex docs: Developer commands (CLI) (opens in a new tab); Codex docs: Models (opens in a new tab); Codex docs: Speed (opens in a new tab); Codex docs: Subagents (opens in a new tab); OpenAI API docs: Pricing (opens in a new tab); Unrolling the Codex agent loop (2026-01-23) (opens in a new tab); OpenAI Help Center: What are tokens and how to count them (opens in a new tab); OpenAI API docs: Prompt caching (opens in a new tab); tiktoken 0.12.0, OpenAI's open-source tokenizer, encoding o200k_base (run on the server, 2026-10-04) (opens in a new tab).
- Google (Antigravity, the Gemini API): Google Antigravity built an OS (2026-05-19) (opens in a new tab); Changes to Antigravity plans (2026-05-19) (opens in a new tab); Antigravity docs: Model quotas (/usage) (opens in a new tab); Antigravity docs: AI credits in the CLI (opens in a new tab); Antigravity docs: Headless mode (opens in a new tab); Antigravity docs: CLI reference (opens in a new tab); Antigravity docs: Home (opens in a new tab); Antigravity docs: Models (opens in a new tab); Antigravity docs: Plans and AI credits (opens in a new tab); Antigravity docs: Settings (opens in a new tab); Antigravity: Pricing (opens in a new tab); Gemini API docs: Text generation (opens in a new tab); Gemini API docs: Gemini 3.8 Flash (opens in a new tab); Gemini API docs: Pricing (opens in a new tab); Gemini API docs: Understand and count tokens (opens in a new tab); Google AI plans (opens in a new tab).
- The studio: episode 3's run log and its record (
03-agents-at-work/research/parts/run.md, sections 1.1 and 7); the recorded sessions (07-tokens-and-limits/research/parts/recordings.md); tiktoken (opens in a new tab) 0.12.0, run on the server.
Corrections
None so far. If you find something wrong, email doug,@douglashutchings.com.
Transcript
Every spoken line, by chapter
Tokens and limits
In episode 3, a lead agent and 17 helpers reviewed a folder of about 700,000 tokens.
Between them, they used 33.5 million. Where did the rest come from?
AI for Research Efficiency. Episode 7: Tokens and limits.
Last time, you published a page from your repository.
This time: what a token is, why a long session costs more, and how to make your plan last.
1. What a token is
One: what a token is.
A token is a piece of a word. A common word is one piece; a rare name splits into several.
This prompt is 8 words, and 13 tokens in OpenAI's open tokenizer. Each company splits text its own way.
What the agent reads is input. What it writes is output, and its thinking counts as output too.
2. Every turn sends it all again
Two: every turn sends it all again.
The model remembers nothing between turns. So each time, the tool sends the whole conversation again: its instructions, your files, every answer so far.
So each turn of a long session costs more than the last.
A cache makes that affordable: text sent before is read again for a tenth of the price, or less.
In episode 3, 94% of what the agents read came from the cache.
At list prices, the run comes to about $25; without the cache, more than 5 times that.
It ran on a monthly plan, so nothing was charged; it drew on the plan's limits.
Switching models in the middle, or compacting, starts the cache over.
3. The context window
Three: the context window.
That's how much the model can hold at once: a million tokens in Claude Code and Antigravity, about a quarter of that in Codex.
Before you type a word, the tool's own instructions and tools are already there. Here, 31,000 tokens.
When the window fills, the tool compacts: it swaps the conversation for a summary. /compact does it now, in Claude Code and Codex; Antigravity compacts by itself.
Or start fresh with /clear. In Claude Code, that costs nothing.
4. A plan's limits
Four: a plan's limits.
Most plans meter use in a 5-hour window, with a weekly cap on top. Every terminal and every helper draws on the same allowance.
At the limit, Claude Code waits and carries on by itself; Codex can finish the turn it's on; Antigravity waits, unless you let it spend credits.
Each company sells extra usage past the limit. That's paid; you decide whether to buy it.
And with an API key instead of a plan, every token is billed.
5. See your usage
Five: see your usage.
In Claude Code, type /usage: the session and the week, each with when it resets.
In Codex, /status shows the same two windows, and the context. Its /usage is your token history.
In Antigravity, /usage shows each group of models, with its 5-hour and weekly limits.
Careful: /usage means limits in two of the tools, and your history in the third.
6. Make it last
Six: make it last.
One task per session, and /clear between tasks.
A smaller model, or lower effort, for simple jobs.
Let the agent open the files it needs, rather than pasting everything in. For heavy reading, helpers hand back summaries.
Before a big run, ask for an estimate. Before a deadline, check your usage.
7. Many agents multiply it
Seven: many agents multiply it.
Anthropic measured about 4 times the tokens of a chat for one agent, and about 15 times for a team of agents.
Google's agents built an operating system: 93 helpers, over 2.6 billion tokens.
OpenAI publishes no figure. It says only that helpers use more.
Try it
Pause here and try it: open your usage view, /usage, or /status in Codex, and read the two windows.
Review and next
So: a token is a piece of a word, every turn sends it all again, and one allowance covers it all.
Next: many agents at once.
