1
00:00:00,900 --> 00:00:02,082
In episode 3,

2
00:00:02,293 --> 00:00:08,044
a lead agent and 17 helpers reviewed
a folder of about 700,000 tokens.

3
00:00:09,604 --> 00:00:14,621
Between them, they used 33.5 million.
Where did the rest come from?

4
00:00:17,200 --> 00:00:22,436
AI for Research Efficiency.
Episode 7: Tokens and limits.

5
00:00:24,300 --> 00:00:27,319
Last time, you published a page
from your repository.

6
00:00:28,080 --> 00:00:32,372
This time: what a token is,
why a long session costs more,

7
00:00:32,643 --> 00:00:34,628
and how to make your plan last.

8
00:00:36,500 --> 00:00:38,646
One: what a token is.

9
00:00:39,561 --> 00:00:41,464
A token is a piece of a word.

10
00:00:41,854 --> 00:00:45,800
A common word is one piece;
a rare name splits into several.

11
00:00:47,507 --> 00:00:52,596
This prompt is 8 words, and 13
tokens in OpenAI's open tokenizer.

12
00:00:53,168 --> 00:00:55,652
Each company splits text its own way.

13
00:00:57,132 --> 00:00:59,035
What the agent reads is input.

14
00:00:59,306 --> 00:01:03,393
What it writes is output, and its
thinking counts as output too.

15
00:01:07,700 --> 00:01:10,671
Two: every turn sends it all again.

16
00:01:11,394 --> 00:01:13,795
The model remembers
nothing between turns.

17
00:01:14,106 --> 00:01:17,418
So each time, the tool sends
the whole conversation again:

18
00:01:17,788 --> 00:01:21,330
its instructions, your files,
every answer so far.

19
00:01:22,450 --> 00:01:26,078
So each turn of a long session
costs more than the last.

20
00:01:27,238 --> 00:01:28,899
A cache makes that affordable:

21
00:01:29,170 --> 00:01:33,353
text sent before is read again
for a tenth of the price, or less.

22
00:01:34,953 --> 00:01:39,680
In episode 3, 94% of what the
agents read came from the cache.

23
00:01:41,041 --> 00:01:44,251
At list prices,
the run comes to about $25;

24
00:01:44,502 --> 00:01:47,270
without the cache,
more than 5 times that.

25
00:01:48,394 --> 00:01:53,104
It ran on a monthly plan, so nothing was
charged; it drew on the plan's limits.

26
00:01:54,176 --> 00:01:58,244
Switching models in the middle,
or compacting, starts the cache over.

27
00:02:00,500 --> 00:02:02,636
Three: the context window.

28
00:02:03,416 --> 00:02:05,458
That's how much the
model can hold at once:

29
00:02:05,708 --> 00:02:08,501
a million tokens in Claude Code
and Antigravity,

30
00:02:08,771 --> 00:02:10,733
about a quarter of that in Codex.

31
00:02:12,153 --> 00:02:16,445
Before you type a word, the tool's own
instructions and tools are already there.

32
00:02:16,796 --> 00:02:19,202
Here, 31,000 tokens.

33
00:02:20,722 --> 00:02:26,212
When the window fills, the tool compacts:
it swaps the conversation for a summary.

34
00:02:26,262 --> 00:02:31,630
/compact does it now, in Claude Code and
Codex; Antigravity compacts by itself.

35
00:02:32,790 --> 00:02:37,650
Or start fresh with /clear.
In Claude Code, that costs nothing.

36
00:02:41,300 --> 00:02:43,449
Four: a plan's limits.

37
00:02:44,293 --> 00:02:48,519
Most plans meter use in a 5-hour window,
with a weekly cap on top.

38
00:02:48,910 --> 00:02:52,395
Every terminal and every helper
draws on the same allowance.

39
00:02:53,756 --> 00:02:56,960
At the limit, Claude Code waits
and carries on by itself;

40
00:02:57,290 --> 00:02:59,123
Codex can finish the turn it's on;

41
00:02:59,493 --> 00:03:02,597
Antigravity waits,
unless you let it spend credits.

42
00:03:03,784 --> 00:03:06,550
Each company sells extra
usage past the limit.

43
00:03:06,820 --> 00:03:09,405
That's paid;
you decide whether to buy it.

44
00:03:10,445 --> 00:03:14,573
And with an API key instead of
a plan, every token is billed.

45
00:03:17,300 --> 00:03:19,461
Five: see your usage.

46
00:03:20,264 --> 00:03:26,253
In Claude Code, type /usage: the session
and the week, each with when it resets.

47
00:03:28,030 --> 00:03:32,973
In Codex, /status shows the same
two windows, and the context.

48
00:03:33,484 --> 00:03:36,305
Its /usage is your token history.

49
00:03:38,026 --> 00:03:41,949
In Antigravity,
/usage shows each group of models,

50
00:03:42,000 --> 00:03:44,301
with its 5-hour and weekly limits.

51
00:03:45,905 --> 00:03:51,762
Careful: /usage means limits in two of
the tools, and your history in the third.

52
00:03:55,700 --> 00:03:57,836
Six: make it last.

53
00:03:58,738 --> 00:04:02,894
One task per session,
and /clear between tasks.

54
00:04:04,098 --> 00:04:07,570
A smaller model,
or lower effort, for simple jobs.

55
00:04:08,633 --> 00:04:12,426
Let the agent open the files it needs,
rather than pasting everything in.

56
00:04:12,778 --> 00:04:15,747
For heavy reading,
helpers hand back summaries.

57
00:04:16,802 --> 00:04:22,020
Before a big run, ask for an estimate.
Before a deadline, check your usage.

58
00:04:24,500 --> 00:04:27,403
Seven: many agents multiply it.

59
00:04:28,265 --> 00:04:32,291
Anthropic measured about 4 times
the tokens of a chat for one agent,

60
00:04:32,501 --> 00:04:35,405
and about 15 times
for a team of agents.

61
00:04:36,988 --> 00:04:39,232
Google's agents built
an operating system:

62
00:04:39,543 --> 00:04:43,510
93 helpers, over 2.6 billion tokens.

63
00:04:44,878 --> 00:04:49,482
OpenAI publishes no figure.
It says only that helpers use more.

64
00:04:53,300 --> 00:04:56,626
Pause here and try it:
open your usage view,

65
00:04:56,676 --> 00:05:01,162
/usage, or /status in Codex,
and read the two windows.

66
00:05:05,300 --> 00:05:07,824
So: a token is a piece of a word,

67
00:05:08,134 --> 00:05:12,500
every turn sends it all again,
and one allowance covers it all.

68
00:05:14,700 --> 00:05:17,357
Next: many agents at once.
