AI for Research Efficiency, episode 7: Tokens and limits Narration: an AI-generated voice (ElevenLabs). [0:00] Tokens and limits In episode 3, a lead agent and 17 helpers reviewed a folder of about 700,000 tokens. Between them, they used 33.5 million. Where did the rest come from? AI for Research Efficiency. Episode 7: Tokens and limits. Last time, you published a page from your repository. This time: what a token is, why a long session costs more, and how to make your plan last. [0:36] 1. What a token is One: what a token is. A token is a piece of a word. A common word is one piece; a rare name splits into several. This prompt is 8 words, and 13 tokens in OpenAI's open tokenizer. Each company splits text its own way. What the agent reads is input. What it writes is output, and its thinking counts as output too. [1:07] 2. Every turn sends it all again Two: every turn sends it all again. The model remembers nothing between turns. So each time, the tool sends the whole conversation again: its instructions, your files, every answer so far. So each turn of a long session costs more than the last. A cache makes that affordable: text sent before is read again for a tenth of the price, or less. In episode 3, 94% of what the agents read came from the cache. At list prices, the run comes to about $25; without the cache, more than 5 times that. It ran on a monthly plan, so nothing was charged; it drew on the plan's limits. Switching models in the middle, or compacting, starts the cache over. [2:00] 3. The context window Three: the context window. That's how much the model can hold at once: a million tokens in Claude Code and Antigravity, about a quarter of that in Codex. Before you type a word, the tool's own instructions and tools are already there. Here, 31,000 tokens. When the window fills, the tool compacts: it swaps the conversation for a summary. /compact does it now, in Claude Code and Codex; Antigravity compacts by itself. Or start fresh with /clear. In Claude Code, that costs nothing. [2:40] 4. A plan's limits Four: a plan's limits. Most plans meter use in a 5-hour window, with a weekly cap on top. Every terminal and every helper draws on the same allowance. At the limit, Claude Code waits and carries on by itself; Codex can finish the turn it's on; Antigravity waits, unless you let it spend credits. Each company sells extra usage past the limit. That's paid; you decide whether to buy it. And with an API key instead of a plan, every token is billed. [3:16] 5. See your usage Five: see your usage. In Claude Code, type /usage: the session and the week, each with when it resets. In Codex, /status shows the same two windows, and the context. Its /usage is your token history. In Antigravity, /usage shows each group of models, with its 5-hour and weekly limits. Careful: /usage means limits in two of the tools, and your history in the third. [3:55] 6. Make it last Six: make it last. One task per session, and /clear between tasks. A smaller model, or lower effort, for simple jobs. Let the agent open the files it needs, rather than pasting everything in. For heavy reading, helpers hand back summaries. Before a big run, ask for an estimate. Before a deadline, check your usage. [4:24] 7. Many agents multiply it Seven: many agents multiply it. Anthropic measured about 4 times the tokens of a chat for one agent, and about 15 times for a team of agents. Google's agents built an operating system: 93 helpers, over 2.6 billion tokens. OpenAI publishes no figure. It says only that helpers use more. [4:52] Try it Pause here and try it: open your usage view, /usage, or /status in Codex, and read the two windows. [5:04] Review and next So: a token is a piece of a word, every turn sends it all again, and one allowance covers it all. Next: many agents at once.