WEBVTT

00:00:00.900 --> 00:00:03.729
57 files on one research question:

00:00:04.100 --> 00:00:07.871
papers, spreadsheets,
Word tables and slide decks.

00:00:08.973 --> 00:00:11.479
More files than a browser
chat takes at once.

00:00:11.991 --> 00:00:17.304
Here, a team of AI agents reads them all,
and checks every number against its page.

00:00:19.600 --> 00:00:24.579
AI for Research Efficiency.
Episode 3: Agents at work.

00:00:26.700 --> 00:00:30.389
Last time, you installed an agent
and asked it a first question.

00:00:31.094 --> 00:00:34.837
In this episode, a literature
review by several agents at once.

00:00:35.287 --> 00:00:39.590
You write the rules, give each agent
one job, and check what they find.

00:00:41.300 --> 00:00:45.892
The question comes from Arkansas, which
plants nearly half of the country's rice.

00:00:46.594 --> 00:00:51.099
Rice fields are usually kept flooded,
and flooded soil gives off methane.

00:00:52.012 --> 00:00:55.932
Alternate wetting and drying lets
the water drain away between floods.

00:00:57.135 --> 00:01:00.903
How much methane and water does
that save? What happens to the yield?

00:01:01.234 --> 00:01:04.621
And do the Arkansas results agree
with the rest of the world?

00:01:07.700 --> 00:01:12.667
The folder holds copies of 57 files
that are free to share, in 8 formats.

00:01:13.038 --> 00:01:16.181
The PDFs alone run to 500 pages.

00:01:17.523 --> 00:01:20.590
A browser chat takes 10
to 40 files at a time,

00:01:20.721 --> 00:01:22.735
and it can't reach the
folder on your computer.

00:01:23.146 --> 00:01:26.594
And this job needs every
number traced back to its page.

00:01:29.300 --> 00:01:30.584
First, the rules.

00:01:31.036 --> 00:01:33.172
They go in a plain text
file in the folder,

00:01:33.263 --> 00:01:35.509
which the agent reads
before it starts.

00:01:36.490 --> 00:01:39.473
Claude Code reads a
file called CLAUDE.md.

00:01:40.245 --> 00:01:43.488
Codex and Antigravity read AGENTS.md.

00:01:44.761 --> 00:01:46.947
Write it as you'd brief
a research assistant:

00:01:47.198 --> 00:01:50.446
the question, what to extract,
and how to cite it.

00:01:51.367 --> 00:01:55.271
Every number with its page,
sheet and cell, or slide.

00:01:55.602 --> 00:01:59.456
Quotes, word for word.
Anything uncertain, marked.

00:01:59.786 --> 00:02:01.948
And never change the source files.

00:02:05.300 --> 00:02:06.463
Then the jobs.

00:02:06.874 --> 00:02:09.390
Each kind of agent gets a
short description of its own:

00:02:09.741 --> 00:02:12.127
a reader, a checker and a writer.

00:02:13.116 --> 00:02:14.698
Readers pull out the evidence.

00:02:15.089 --> 00:02:18.324
Checkers trust nothing:
they open every source again.

00:02:18.715 --> 00:02:21.739
The writer drafts the answer
from checked rows only.

00:02:24.500 --> 00:02:27.429
Then one request, in plain words.

00:02:31.232 --> 00:02:33.735
The run is long,
so it works in auto mode,

00:02:33.825 --> 00:02:38.370
the faster speed from episode two.
That's why the folder holds only copies.

00:02:41.300 --> 00:02:44.068
The lead agent splits
the files into 8 batches,

00:02:44.159 --> 00:02:47.066
by topic,
and starts 8 readers at once.

00:02:48.387 --> 00:02:50.351
Each reader works
through its own batch.

00:02:50.822 --> 00:02:55.040
From a PDF, it takes the text
with the table columns in place,

00:02:55.090 --> 00:02:57.846
or looks at the page as an image.

00:02:57.896 --> 00:03:00.941
For Word, Excel and PowerPoint,
it writes a short Python script.

00:03:02.102 --> 00:03:06.073
Every finding becomes a row:
one number, for one outcome,

00:03:06.204 --> 00:03:10.073
with its file, its exact place,
and the words it came from.

00:03:11.414 --> 00:03:17.377
In under 6 minutes, the readers are
done: 667 rows, from 40 of the files.

00:03:17.527 --> 00:03:20.098
The other 17 had
nothing on the question,

00:03:20.249 --> 00:03:23.090
or only pictures,
and the readers said which.

00:03:26.900 --> 00:03:29.461
Then 8 checkers, again at once.

00:03:30.032 --> 00:03:35.275
Each one opens every cited page, cell
or slide, and compares it with the row.

00:03:36.665 --> 00:03:41.908
573 rows are confirmed,
94 corrected, and none rejected.

00:03:42.438 --> 00:03:47.040
Most corrections are small: the right
page, or how a study is described.

00:03:48.420 --> 00:03:49.742
One changes the meaning.

00:03:50.313 --> 00:03:55.249
A reader filed an EPA sentence, "as
much as 50 percent", as a comparison:

00:03:55.780 --> 00:04:00.626
flooding fields in winter, or not.
And it said the study was in California.

00:04:02.090 --> 00:04:03.691
The checker read the page again.

00:04:04.201 --> 00:04:08.214
The sentence gives the share of a year's
methane released during winter flooding,

00:04:08.364 --> 00:04:12.907
and the page never says where. The
right number, but the wrong claim.

00:04:17.300 --> 00:04:20.407
Last, the writer drafts the
answer from checked rows.

00:04:20.838 --> 00:04:24.376
The lead agent then checks that
every statement carries its citation,

00:04:24.547 --> 00:04:26.250
and fixes the ones that don't.

00:04:27.417 --> 00:04:32.844
Across 10 meta-analyses, drying
cut methane by 31 to 62 percent:

00:04:33.155 --> 00:04:34.837
about half, on average.

00:04:35.777 --> 00:04:39.527
It saved a quarter to two fifths of
the water, depending on what's counted.

00:04:39.998 --> 00:04:43.346
Yield held on average,
unless the drying was severe.

00:04:44.269 --> 00:04:48.742
Nitrous oxide went up, so the climate
gain is smaller than the methane cut.

00:04:49.883 --> 00:04:55.007
And in Arkansas, field trials saved 34
to 50 percent of the irrigation water,

00:04:55.057 --> 00:04:59.411
with no loss of yield: in line with the
world's studies of irrigation alone.

00:04:59.761 --> 00:05:01.492
But the folder has no Arkansas

00:05:01.542 --> 00:05:04.364
measurement of nitrous oxide
under this practice.

00:05:07.700 --> 00:05:08.683
Now it's your turn.

00:05:09.156 --> 00:05:13.671
Pick a sentence, follow its citation
to the row, and the row to the page.

00:05:15.234 --> 00:05:19.120
Here's the water figure: row E0159.

00:05:19.170 --> 00:05:20.883
Page 2, Table 1:

00:05:21.254 --> 00:05:25.179
−25.7%, in the Water use column.

00:05:26.619 --> 00:05:29.636
Do that for every figure
your conclusions rest on.

00:05:31.700 --> 00:05:33.945
The agents also listed
what they couldn't do:

00:05:34.457 --> 00:05:38.526
numbers shown only in charts, and
key papers missing from the folder.

00:05:39.606 --> 00:05:43.957
This run took 17 and a half minutes.
That's one example,

00:05:44.007 --> 00:05:47.415
on one folder.

00:05:50.900 --> 00:05:54.382
Pause now, and try the practice
folder from the companion page.

00:05:54.873 --> 00:05:58.545
Start your agent in the mode that
asks first, and paste the request.

00:05:58.956 --> 00:06:02.138
You should get a table
where every row names its page.

00:06:05.200 --> 00:06:08.925
So: write the rules,
give each agent one job,

00:06:09.276 --> 00:06:12.439
check against the source,
and follow the trail yourself.

00:06:14.900 --> 00:06:15.923
The rules file,

00:06:16.014 --> 00:06:20.250
the job descriptions and the practice
folder are all on the companion page.
