Episode 7: Tokens and limits

What a token is, why every turn of an agent session sends the whole conversation again, how caching makes that affordable, what the context window holds, and how a plan's five-hour and weekly limits work, in Claude Code, Codex CLI and Antigravity CLI.

Length
5:24
As of
4 October 2026
Narration
An AI-generated voice (ElevenLabs)

Downloads

AI for Research Efficiency, episode 7, with Douglas Hutchings. Narration: an AI-generated voice (ElevenLabs).

What a token is, why a long session costs more, and how to make a monthly plan last, in Claude Code, Codex CLI and Antigravity CLI. Plans, limits, prices and windows as of 4 October 2026, from Anthropic's, OpenAI's and Google's own pages (sources below), and from sessions recorded that day in Claude Code 2.1.289 and Codex 0.160.0. These change often: check each vendor's page before you rely on a number.

1. What a token is

A token is a piece of a word: a common word is one token, a rare name several. Each company splits text its own way, and each gives its own rule of thumb:

OpenAIAnthropicGoogle
A token, roughlyabout 4 characters; 100 tokens, about 75 words1M tokens, about 555,000 words (its current tokenizer)about 4 characters; 100 tokens, 60 to 80 words

In the film, "Summarize the methane results in the Minamikawa paper." is 8 words and 13 tokens in tiktoken, OpenAI's open-source tokenizer (encoding o200k_base): Summ ar ize the methane results in the Min amik awa paper . Claude and Gemini split the same line differently.

What the agent reads is input; what it writes is output, and its thinking is billed as output in all three.

2. Every turn sends it all again

The model remembers nothing between turns, so the tool sends the whole conversation again each time: its own instructions and tools, your files, every answer so far. Anthropic: "The model doesn't remember anything between requests, so Claude Code re-sends the full context". OpenAI: "Every time you send a new message to an existing conversation, the conversation history is included as part of the prompt for the new turn". Antigravity's own example: asked to reply with one word, its agent read 30,384 input tokens; the second turn sent 278 new tokens and read 30,214 from the cache.

A cache makes that affordable: text sent before is read again at a tenth of the input price, or less. Prices per million tokens, at list prices, on 4 October 2026:

Claude Opus 5.5 (Claude Code's default)GPT-6.1 Sol (Codex's default)Gemini 3.8 Flash (Antigravity's agents)
Input$4$2.00$0.75
Read from the cache$0.20 (0.05 times input)$0.10 (0.05 times input)$0.075, plus $0.50 per million tokens per hour of storage
Written to the cache$5 (kept 5 minutes) or $8 (1 hour)$2.50no write charge listed for implicit caching
Output, thinking included$20$10.00$3.75
Notesup to 272K input; more above thatthrough 31 December 2026; doubled from 1 January 2027

Episode 3's run, priced (from its own log; Claude Code's estimate at list prices; it ran on a monthly plan, so nothing was charged and it drew on the plan's limits):

PartTokensCost at list prices
New input916$0.004
Read from the cache31,021,208$6.20
Written to the cache2,085,946$10.87
Output (78,510 of it thinking)391,000$7.82
In all33,499,070$24.89

94% of the input came from the cache. Without it, the same run would have cost about $140, more than five times as much.

What starts the cache over: in Claude Code, switching models, changing the effort level (on most models), fast mode, and compacting; in Codex, compacting or summarizing; in any tool, a change near the start of the conversation.

3. The context window

How much the model can hold at once:

Claude CodeCodex CLIAntigravity CLI
Default model and windowOpus 5.5: 1,000,000 tokens, on every planGPT-6.1 Sol: 272,000 tokens, of which Codex uses 95% (/status shows 258K)Gemini 3.8 Flash: 1,048,576 tokens
When it fillscompacts by itself, at about 967K by defaultcompacts by itselfcompacts by itself
Compact now/compact/compactnone (automatic only)
Start fresh/clear (it costs nothing)/clear, or /new/clear
See what fills it/context/status/context

Before you type a word, the tool's own instructions and tools are already there: in the recorded Claude Code session, 31,300 tokens (3% of the window; that account has extra tools and skills installed, so a new one starts smaller). One question about one paper took it to 88,100 (9%).

4. A plan's limits

Most plans meter use in a five-hour window with a weekly cap on top, and every terminal, helper agent and cloud task on the account draws on the same allowance. None of the three publishes its allowance as a number of tokens.

Claude (Claude Code)ChatGPT (Codex)Google (Antigravity)
PlansPro $20 a month ($17 a month billed yearly); Max from $100 (5x or 20x Pro's usage)Go $8; Plus $20; Pro $100, $200 or $500Free; Google AI Pro $19.99; Ultra $100 and $200
Windowsa session limit that resets every five hours, and a weekly limita five-hour limit (Pro has none) and weekly limitsPro: every five hours until the weekly limit; free: weekly
At the limitwaits, then carries on by itself ("Continuing automatically at 3:45pm · esc to cancel")the running turn can finish; then wait, or buy creditswaits for the quota, unless AI credits are allowed (a setting)
Extra usage (paid)usage credits, at API rates, when you turn them oncreditsAI credits, Pro and Ultra only

A plan or a key: with an API key instead of a plan, every token is billed at API rates; in Claude Code a set ANTHROPIC_API_KEY is used instead of your subscription.

Fast modes spend faster: Claude's fast mode is paid usage credits only, outside the plan's limits; Codex's Fast mode uses the plan's limits at 2.5 times the standard rate.

5. See your usage

Claude CodeCodex CLIAntigravity CLI
Plan limits/usage/status/usage (or /quota)
Token history/usage (Stats)/usagenot documented
What fills the context/context/status/context

The trap: /usage shows plan limits in Claude Code and Antigravity, but your token history in Codex, whose limits are in /status. Codex's /usage menu also offers to redeem a rate-limit reset; choose View analytics to look only.

6. Make it last

  1. One task per session, and /clear between tasks.
  2. A smaller model, or lower effort, for simple jobs.
  3. Let the agent open the files it needs, rather than pasting everything in.
  4. For heavy reading, helpers: they hand back summaries.
  5. Before a big run, ask for an estimate.
  6. Before a deadline, check your usage.

7. Many agents multiply it

  • Anthropic (June 2025): "agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats". Claude Code's documentation: agent teams use "approximately 7x more tokens than standard sessions when teammates run in plan mode".
  • Google (May 2026): an operating system built by 93 subagents, "over 2.6B tokens", $916.92 at API prices.
  • OpenAI publishes no multiplier: "subagent workflows consume more tokens than comparable single-agent runs".

Episode 8 is about running many agents at once, safely.

Not confirmed yet: how many tokens a plan's windows hold (no vendor publishes it), how cached tokens count against Claude's and Codex's windows (both say only that caching lowers usage), and Antigravity's usage panel as it appears when signed in (the film draws it from its documentation).

Try it

  1. Open your agent.
  2. Type /usage (in Codex, /status).
  3. Read the two windows: how much of the five hours and of the week is used, and when each resets.

Sources

Read on 4 October 2026 (copies kept with the project). The recorded screens come from Claude Code 2.1.289 and Codex 0.160.0 on the studio's server; Antigravity's from its documentation; episode 3's run from its own log.

Corrections

None so far. If you find something wrong, email doug@douglashutchings.com.

Transcript

Every spoken line, by chapter

Tokens and limits

In episode 3, a lead agent and 17 helpers reviewed a folder of about 700,000 tokens.

Between them, they used 33.5 million. Where did the rest come from?

AI for Research Efficiency. Episode 7: Tokens and limits.

Last time, you published a page from your repository.

This time: what a token is, why a long session costs more, and how to make your plan last.

1. What a token is

One: what a token is.

A token is a piece of a word. A common word is one piece; a rare name splits into several.

This prompt is 8 words, and 13 tokens in OpenAI's open tokenizer. Each company splits text its own way.

What the agent reads is input. What it writes is output, and its thinking counts as output too.

2. Every turn sends it all again

Two: every turn sends it all again.

The model remembers nothing between turns. So each time, the tool sends the whole conversation again: its instructions, your files, every answer so far.

So each turn of a long session costs more than the last.

A cache makes that affordable: text sent before is read again for a tenth of the price, or less.

In episode 3, 94% of what the agents read came from the cache.

At list prices, the run comes to about $25; without the cache, more than 5 times that.

It ran on a monthly plan, so nothing was charged; it drew on the plan's limits.

Switching models in the middle, or compacting, starts the cache over.

3. The context window

Three: the context window.

That's how much the model can hold at once: a million tokens in Claude Code and Antigravity, about a quarter of that in Codex.

Before you type a word, the tool's own instructions and tools are already there. Here, 31,000 tokens.

When the window fills, the tool compacts: it swaps the conversation for a summary. /compact does it now, in Claude Code and Codex; Antigravity compacts by itself.

Or start fresh with /clear. In Claude Code, that costs nothing.

4. A plan's limits

Four: a plan's limits.

Most plans meter use in a 5-hour window, with a weekly cap on top. Every terminal and every helper draws on the same allowance.

At the limit, Claude Code waits and carries on by itself; Codex can finish the turn it's on; Antigravity waits, unless you let it spend credits.

Each company sells extra usage past the limit. That's paid; you decide whether to buy it.

And with an API key instead of a plan, every token is billed.

5. See your usage

Five: see your usage.

In Claude Code, type /usage: the session and the week, each with when it resets.

In Codex, /status shows the same two windows, and the context. Its /usage is your token history.

In Antigravity, /usage shows each group of models, with its 5-hour and weekly limits.

Careful: /usage means limits in two of the tools, and your history in the third.

6. Make it last

Six: make it last.

One task per session, and /clear between tasks.

A smaller model, or lower effort, for simple jobs.

Let the agent open the files it needs, rather than pasting everything in. For heavy reading, helpers hand back summaries.

Before a big run, ask for an estimate. Before a deadline, check your usage.

7. Many agents multiply it

Seven: many agents multiply it.

Anthropic measured about 4 times the tokens of a chat for one agent, and about 15 times for a team of agents.

Google's agents built an operating system: 93 helpers, over 2.6 billion tokens.

OpenAI publishes no figure. It says only that helpers use more.

Try it

Pause here and try it: open your usage view, /usage, or /status in Codex, and read the two windows.

Review and next

So: a token is a piece of a word, every turn sends it all again, and one allowance covers it all.

Next: many agents at once.

AI for Research Efficiency, with Douglas Hutchings. Narration: an AI-generated voice (ElevenLabs).