Episode 3: Agents at work: a literature review

A literature review of 57 files by several AI agents at once, replayed from a real run: you write the rules, give each agent one job, and check every number they find against its page.

Length
6:26
As of
2 October 2026
Narration
An AI-generated voice (ElevenLabs)

Downloads

AI for Research Efficiency, episode 3, with Douglas Hutchings. Narration: an AI-generated voice (ElevenLabs).

One real run, recorded on 2 October 2026: a review of 57 files (papers, spreadsheets, Word tables, slide decks) by eight reader agents and eight checker agents working at the same time, and a writer that drafts the answer from checked rows only. Every row in the evidence table names its file and its page, sheet and cell, or slide.

Try it: the practice folder

Download the practice folder (2.8 MB): 10 of the run's files in five formats (PDF, Word, Excel, PowerPoint, CSV) and a saved web page, all openly licensed; the rules file, the three agent files and the request from the recorded run; a README with each file's license; and the recorded run's own outputs in example-run/ (evidence.csv, checks.md, summary.md) to compare with yours, with the full review's list of its 57 sources (MANIFEST.csv: titles, links and licenses).

  1. Unzip it, open a terminal in the folder, and set up Python as the README shows (the agents write short Python scripts to read Word, Excel and PowerPoint files).
  2. Start your agent in the folder (Claude Code: claude), and paste the request from PROMPT.txt.
  3. Start in the mode that asks before each step; when you trust what it does, switch to the faster mode.
  4. Check the result: pick a sentence in summary.md, follow its row in evidence.csv, and open the page.

You should get a table where every row names its page.

The pattern

  1. Copies only. Put copies of the files in sources/; the agents are told never to change them.
  2. The rules, in a plain text file. Claude Code reads CLAUDE.md; Codex and Antigravity read AGENTS.md (Antigravity also reads GEMINI.md). Write it as you would brief a research assistant: the question, what to extract, how to cite it.
  3. One job per agent. A reader, a checker and a writer, each described in its own short file (Claude Code keeps them in .claude/agents/; other tools have their own ways to define helpers).
  4. Check against the source. The checkers open every cited page, sheet or slide again; you follow the trail for the sentences that matter.

The rules file of the recorded run

Markdown
# Literature review: alternate wetting and drying (AWD) in rice ## The question Compared with continuous flooding, how much does alternate wetting and drying cut methane and irrigation water, whathappens to yield, and do Arkansas results agree with the global evidence? ## The folder - `sources/`: the source files (PDF, Word, Excel, PowerPoint, CSV, one saved web page, one zip of CSV files). They are  copies. Never change, move, rename or delete anything in `sources/`. `sources/MANIFEST.csv` gives each file's title,  authors, year, license and link.- `work/`: scratch space for converted text, page images, scripts and each agent's partial results. Keep temporary  files here, not in `/tmp`.- The outputs, at the top of the folder: `evidence.csv`, `checks.md`, `summary.md`. ## Tools - Python: `.venv/bin/python` (pandas, openpyxl, xlrd, python-docx, python-pptx, pdfplumber and markitdown are  installed). Save every script you write in `work/scripts/` so someone else can run it again.- PDF: `pdfinfo file.pdf` (page count); `pdftotext -layout -f 3 -l 3 file.pdf -` (page 3, keeping table columns);  `pdftoppm -png -r 110 -f 3 -l 3 file.pdf work/pages/name` (an image of page 3, which you can open and look at).- Word, Excel, PowerPoint: python-docx, openpyxl or xlrd, python-pptx, or `.venv/bin/python -m markitdown file`.- Page numbers are the PDF's own page numbers (1 = the first page of the file). If the printed journal page differs,  give both: `p. 2 (journal p. 526)`. ## What to extract: `evidence.csv` One row per finding: one number for one outcome in one comparison. The columns, in this order: | Column | What goes in it ||---|---|| `row_id` | left blank by readers; numbered when the rows are combined || `source_file` | the file name in `sources/` || `locator` | exactly where: `p. 4, Table 2`, `p. 7, para 2`, `sheet Data, cell H12`, `slide 9` || `study_design` | meta-analysis, review, field experiment, on-farm trial, model, dataset, report, calculation, other || `study_location` | country, state or site; `global` for syntheses || `comparison` | what is compared, for example `AWD vs continuous flooding` || `outcome` | `methane`, `water`, `yield`, `nitrous oxide` or `other` (say what in `notes`, for example global warming potential) || `effect_value` | the number exactly as printed, with its sign: `-48.2`, `66`, `+1.3` || `unit` | as printed: `%`, `kg CH4 ha-1`, `kg CH4-C ha-1`, `mm`, `t ha-1` || `quote_or_cell` | the sentence quoted word for word, or the table cell with its row and column headings || `confidence` | `high`, `medium` or `low` || `notes` | anything needed to read the number correctly (n, baseline, how it was averaged, which water) or anything uncertain || `reader` | which reader extracted the row || `check_status` | left blank by readers; the checker writes `confirmed`, `corrected` or `rejected` || `check_note` | the checker's reason and what it opened | ## Rules 1. Every row cites its exact page, sheet and cell, or slide. No row without a locator.2. Quotes are word for word. Numbers are copied exactly as printed: sign, decimals and unit. Do not round or convert   units. A number you calculate yourself (for example a mean over a spreadsheet's rows) gets its own row with   `study_design` = `calculation` and the script's path in `notes`.3. Read tables from the page image or from `pdftotext -layout`, never from plain text alone (plain text can mix up the   columns), and name the row and the column of the cell in `quote_or_cell`.4. Mark anything uncertain: `confidence` = `low`, and say why in `notes`. If a file cannot be read (for example it is   only pictures), say so in your report. Do not guess.5. Extract only what bears on the question: methane, nitrous oxide, global warming potential, water and yield under AWD   (or another water-saving regime) compared with continuous flooding, and the Arkansas and US context (state emissions,   rice area). A file with nothing relevant gets no rows; say so in your report.6. Write only in `work/` and the three output files. ## The output files - `evidence.csv`: every row, with the checker's `check_status` and `check_note`.- `checks.md`: the checker's log: for every row, confirmed, corrected (old value, new value, why) or rejected (why), and  what was opened (page image, layout text, cell, slide).- `summary.md`: the answer to the question, drafted only from confirmed or corrected rows, every statement followed by  its citation (file, locator). Ranges where the sources differ. What could not be read, and what is missing.

One of the agent files: the reader

Markdown
---name: readerdescription: Reads one batch of source files and extracts evidence rows for the review in CLAUDE.md. Give each reader its own batch and name; several readers can run at the same time.tools: Read, Bash, Write, Glob, Grep--- You are a careful research assistant extracting evidence for the literature review described in CLAUDE.md (read itfirst). You are given a batch of files in `sources/` and a name for your output, for example `reader-3`. For each file in your batch: 1. Check what it is (`sources/MANIFEST.csv` has the title) and whether it bears on the question.2. Read it with the right tool, and save any script you write in `work/scripts/`:   - PDF: `pdftotext -layout` for the text, one page at a time when you need the page number. For every table or figure     you take a number from, render that page with `pdftoppm -png -r 110 -f N -l N` into `work/pages/` and look at the     image with the Read tool, so that you take the number from the right row and the right column.   - Word (.docx): python-docx (paragraphs and tables), or markitdown.   - Excel (.xlsx, .xls) and CSV: pandas with openpyxl or xlrd. Print the sheet names, the headings and the rows you     use, and record each cell's address (sheet, column letter, row number).   - PowerPoint (.pptx): python-pptx (the text, tables and charts on each slide, with slide numbers).   - Zip: list it, and read the files inside from Python (extract into `work/` if you must, never into `sources/`).   - Web page (.html): markitdown, or read the text.3. Write evidence rows exactly as CLAUDE.md says. Write your rows to `work/rows/<your name>.csv` with the columns from CLAUDE.md, in that order (leave `row_id`,`check_status` and `check_note` empty, and put your name in `reader`). Write a short report to`work/rows/<your name>.md`: for each file, what it is, how you read it (the tool or script), how many rows you took,and anything you could not read or were unsure of. Your final message: the same report in brief (not the rows themselves).

The request, word for word

I'm reviewing the evidence on alternate wetting and drying in rice. The question, the rules and the format of the evidence table are in CLAUDE.md, and the source files are in sources/ (PDFs, Word, Excel, PowerPoint and CSV files; MANIFEST.csv lists them). Please do the review with the agents in .claude/agents: 1. Split the files into batches and run several reader agents in parallel, one batch each, so that every file gets read.2. Combine their rows into evidence.csv and number them.3. Have the checker verify every row against its source page, sheet or slide (you can run several checkers in parallel on different files). Put the checked rows in evidence.csv and the checker's log in checks.md.4. Then have the writer draft summary.md from the checked rows only. Don't change anything in sources/. When you're done, tell me how many rows were extracted, how many the checker confirmed, corrected and rejected, and anything that couldn't be read.

How the run was made

  • Claude Code 2.1.288 in auto mode, on the studio's Linux server, 2 October 2026, 23:30 to 23:48 UTC: 17 minutes 29 seconds, 507 tool calls. Eight readers ran at once (0:25 to 5:39), then eight checkers (6:15 to 10:39), then the writer.
  • Result: 667 rows from 40 of the 57 files; the checkers confirmed 573, corrected 94 and rejected none. The other 17 files gave no rows, each with a reason (nothing on the question, or pictures only).
  • Claude Code reported an estimated $24.89 at API list prices; the run used a monthly plan, where it counts against the plan's usage limits instead.
  • Claude Code refused the writer's own file write ("Subagents should return findings as text"), so the lead agent saved the writer's draft as summary.md, then checked every cited row and number with a script of its own, and fixed an uncited figure before finishing.
  • The review folder sat inside the studio's own repository, so the studio's instructions file was also in the agent's view; the log shows no effect on the review. A run in a folder of its own is cleaner.
  • The film replays the run from its log, as a time-lapse with the run's clock on screen; nothing in it is staged.

The answer, in brief

From checked rows only, each with its citation in summary.md: across 10 meta-analyses, alternate wetting and drying cut methane by 31 to 62% (mean 50.9%); water use fell by 25.7% (irrigation plus rainfall) to 39.7% (irrigation only); yield changed by -5.4% to +11% (mean +1.3%), with losses of about 23% under severe drying; nitrous oxide rose by 37 to 445%. In Arkansas field trials, irrigation water fell by 34.4 to 50.1% with no loss of yield; the folder has no Arkansas measurement of nitrous oxide under the practice.

Sources and licenses

Corrections

None so far. If you find something wrong, email doug@douglashutchings.com.

Transcript

Every spoken line, by chapter

Agents at work

57 files on one research question: papers, spreadsheets, Word tables and slide decks.

More files than a browser chat takes at once. Here, a team of AI agents reads them all, and checks every number against its page.

AI for Research Efficiency. Episode 3: Agents at work.

Last time, you installed an agent and asked it a first question.

In this episode, a literature review by several agents at once. You write the rules, give each agent one job, and check what they find.

The question

The question comes from Arkansas, which plants nearly half of the country's rice.

Rice fields are usually kept flooded, and flooded soil gives off methane.

Alternate wetting and drying lets the water drain away between floods.

How much methane and water does that save? What happens to the yield? And do the Arkansas results agree with the rest of the world?

The folder

The folder holds copies of 57 files that are free to share, in 8 formats. The PDFs alone run to 500 pages.

A browser chat takes 10 to 40 files at a time, and it can't reach the folder on your computer. And this job needs every number traced back to its page.

The rules

First, the rules. They go in a plain text file in the folder, which the agent reads before it starts.

Claude Code reads a file called CLAUDE.md. Codex and Antigravity read AGENTS.md.

Write it as you'd brief a research assistant: the question, what to extract, and how to cite it.

Every number with its page, sheet and cell, or slide. Quotes, word for word. Anything uncertain, marked. And never change the source files.

The jobs

Then the jobs. Each kind of agent gets a short description of its own: a reader, a checker and a writer.

Readers pull out the evidence. Checkers trust nothing: they open every source again. The writer drafts the answer from checked rows only.

The request

Then one request, in plain words.

The run is long, so it works in auto mode, the faster speed from episode two. That's why the folder holds only copies.

Eight readers

The lead agent splits the files into 8 batches, by topic, and starts 8 readers at once.

Each reader works through its own batch. From a PDF, it takes the text with the table columns in place, or looks at the page as an image. For Word, Excel and PowerPoint, it writes a short Python script.

Every finding becomes a row: one number, for one outcome, with its file, its exact place, and the words it came from.

In under 6 minutes, the readers are done: 667 rows, from 40 of the files. The other 17 had nothing on the question, or only pictures, and the readers said which.

Eight checkers

Then 8 checkers, again at once. Each one opens every cited page, cell or slide, and compares it with the row.

573 rows are confirmed, 94 corrected, and none rejected. Most corrections are small: the right page, or how a study is described.

One changes the meaning. A reader filed an EPA sentence, "as much as 50 percent", as a comparison: flooding fields in winter, or not. And it said the study was in California.

The checker read the page again. The sentence gives the share of a year's methane released during winter flooding, and the page never says where. The right number, but the wrong claim.

The answer

Last, the writer drafts the answer from checked rows. The lead agent then checks that every statement carries its citation, and fixes the ones that don't.

Across 10 meta-analyses, drying cut methane by 31 to 62 percent: about half, on average.

It saved a quarter to two fifths of the water, depending on what's counted. Yield held on average, unless the drying was severe.

Nitrous oxide went up, so the climate gain is smaller than the methane cut.

And in Arkansas, field trials saved 34 to 50 percent of the irrigation water, with no loss of yield: in line with the world's studies of irrigation alone. But the folder has no Arkansas measurement of nitrous oxide under this practice.

Check it yourself

Now it's your turn. Pick a sentence, follow its citation to the row, and the row to the page.

Here's the water figure: row E0159. Page 2, Table 1: −25.7%, in the Water use column.

Do that for every figure your conclusions rest on.

What it couldn't do

The agents also listed what they couldn't do: numbers shown only in charts, and key papers missing from the folder.

This run took 17 and a half minutes. That's one example, on one folder.

Try it

Pause now, and try the practice folder from the companion page. Start your agent in the mode that asks first, and paste the request. You should get a table where every row names its page.

Review

So: write the rules, give each agent one job, check against the source, and follow the trail yourself.

The rules file, the job descriptions and the practice folder are all on the companion page.

AI for Research Efficiency, with Douglas Hutchings. Narration: an AI-generated voice (ElevenLabs).