WEBVTT

00:00:00.900 --> 00:00:02.082
In episode 3,

00:00:02.293 --> 00:00:08.044
a lead agent and 17 helpers reviewed
a folder of about 700,000 tokens.

00:00:09.604 --> 00:00:14.621
Between them, they used 33.5 million.
Where did the rest come from?

00:00:17.200 --> 00:00:22.436
AI for Research Efficiency.
Episode 7: Tokens and limits.

00:00:24.300 --> 00:00:27.319
Last time, you published a page
from your repository.

00:00:28.080 --> 00:00:32.372
This time: what a token is,
why a long session costs more,

00:00:32.643 --> 00:00:34.628
and how to make your plan last.

00:00:36.500 --> 00:00:38.646
One: what a token is.

00:00:39.561 --> 00:00:41.464
A token is a piece of a word.

00:00:41.854 --> 00:00:45.800
A common word is one piece;
a rare name splits into several.

00:00:47.507 --> 00:00:52.596
This prompt is 8 words, and 13
tokens in OpenAI's open tokenizer.

00:00:53.168 --> 00:00:55.652
Each company splits text its own way.

00:00:57.132 --> 00:00:59.035
What the agent reads is input.

00:00:59.306 --> 00:01:03.393
What it writes is output, and its
thinking counts as output too.

00:01:07.700 --> 00:01:10.671
Two: every turn sends it all again.

00:01:11.394 --> 00:01:13.795
The model remembers
nothing between turns.

00:01:14.106 --> 00:01:17.418
So each time, the tool sends
the whole conversation again:

00:01:17.788 --> 00:01:21.330
its instructions, your files,
every answer so far.

00:01:22.450 --> 00:01:26.078
So each turn of a long session
costs more than the last.

00:01:27.238 --> 00:01:28.899
A cache makes that affordable:

00:01:29.170 --> 00:01:33.353
text sent before is read again
for a tenth of the price, or less.

00:01:34.953 --> 00:01:39.680
In episode 3, 94% of what the
agents read came from the cache.

00:01:41.041 --> 00:01:44.251
At list prices,
the run comes to about $25;

00:01:44.502 --> 00:01:47.270
without the cache,
more than 5 times that.

00:01:48.394 --> 00:01:53.104
It ran on a monthly plan, so nothing was
charged; it drew on the plan's limits.

00:01:54.176 --> 00:01:58.244
Switching models in the middle,
or compacting, starts the cache over.

00:02:00.500 --> 00:02:02.636
Three: the context window.

00:02:03.416 --> 00:02:05.458
That's how much the
model can hold at once:

00:02:05.708 --> 00:02:08.501
a million tokens in Claude Code
and Antigravity,

00:02:08.771 --> 00:02:10.733
about a quarter of that in Codex.

00:02:12.153 --> 00:02:16.445
Before you type a word, the tool's own
instructions and tools are already there.

00:02:16.796 --> 00:02:19.202
Here, 31,000 tokens.

00:02:20.722 --> 00:02:26.212
When the window fills, the tool compacts:
it swaps the conversation for a summary.

00:02:26.262 --> 00:02:31.630
/compact does it now, in Claude Code and
Codex; Antigravity compacts by itself.

00:02:32.790 --> 00:02:37.650
Or start fresh with /clear.
In Claude Code, that costs nothing.

00:02:41.300 --> 00:02:43.449
Four: a plan's limits.

00:02:44.293 --> 00:02:48.519
Most plans meter use in a 5-hour window,
with a weekly cap on top.

00:02:48.910 --> 00:02:52.395
Every terminal and every helper
draws on the same allowance.

00:02:53.756 --> 00:02:56.960
At the limit, Claude Code waits
and carries on by itself;

00:02:57.290 --> 00:02:59.123
Codex can finish the turn it's on;

00:02:59.493 --> 00:03:02.597
Antigravity waits,
unless you let it spend credits.

00:03:03.784 --> 00:03:06.550
Each company sells extra
usage past the limit.

00:03:06.820 --> 00:03:09.405
That's paid;
you decide whether to buy it.

00:03:10.445 --> 00:03:14.573
And with an API key instead of
a plan, every token is billed.

00:03:17.300 --> 00:03:19.461
Five: see your usage.

00:03:20.264 --> 00:03:26.253
In Claude Code, type /usage: the session
and the week, each with when it resets.

00:03:28.030 --> 00:03:32.973
In Codex, /status shows the same
two windows, and the context.

00:03:33.484 --> 00:03:36.305
Its /usage is your token history.

00:03:38.026 --> 00:03:41.949
In Antigravity,
/usage shows each group of models,

00:03:42.000 --> 00:03:44.301
with its 5-hour and weekly limits.

00:03:45.905 --> 00:03:51.762
Careful: /usage means limits in two of
the tools, and your history in the third.

00:03:55.700 --> 00:03:57.836
Six: make it last.

00:03:58.738 --> 00:04:02.894
One task per session,
and /clear between tasks.

00:04:04.098 --> 00:04:07.570
A smaller model,
or lower effort, for simple jobs.

00:04:08.633 --> 00:04:12.426
Let the agent open the files it needs,
rather than pasting everything in.

00:04:12.778 --> 00:04:15.747
For heavy reading,
helpers hand back summaries.

00:04:16.802 --> 00:04:22.020
Before a big run, ask for an estimate.
Before a deadline, check your usage.

00:04:24.500 --> 00:04:27.403
Seven: many agents multiply it.

00:04:28.265 --> 00:04:32.291
Anthropic measured about 4 times
the tokens of a chat for one agent,

00:04:32.501 --> 00:04:35.405
and about 15 times
for a team of agents.

00:04:36.988 --> 00:04:39.232
Google's agents built
an operating system:

00:04:39.543 --> 00:04:43.510
93 helpers, over 2.6 billion tokens.

00:04:44.878 --> 00:04:49.482
OpenAI publishes no figure.
It says only that helpers use more.

00:04:53.300 --> 00:04:56.626
Pause here and try it:
open your usage view,

00:04:56.676 --> 00:05:01.162
/usage, or /status in Codex,
and read the two windows.

00:05:05.300 --> 00:05:07.824
So: a token is a piece of a word,

00:05:08.134 --> 00:05:12.500
every turn sends it all again,
and one allowance covers it all.

00:05:14.700 --> 00:05:17.357
Next: many agents at once.
