Episode 11: Map a research field with OpenAlex

A real, recorded run: an AI agent maps the research on alternate wetting and drying in rice from OpenAlex (works per year, the leading institutions and journals, and who works with whom in Arkansas), a second agent re-checks every count, and the result becomes one page.

Length
6:00
As of
4 October 2026
Narration
An AI-generated voice (ElevenLabs)

Downloads

AI for Research Efficiency, episode 11, with Douglas Hutchings. Narration: an AI-generated voice (ElevenLabs).

One real run, recorded on 4 October 2026: an AI agent (Claude Code, in auto mode) maps the research on alternate wetting and drying in rice from OpenAlex, with no key: works per year, the institutions and journals that publish the most, and who writes with whom in Arkansas. A second agent re-checks every count with its own calls before the page is made. Nothing was staged. Below: the instruction file, the checker's agent file and the request, word for word, so you can run the same thing on your own field.

The instruction file (CLAUDE.md)

Put it at the top of an empty folder. Change the question, the field and the Arkansas part to your own.

Markdown
# Field map: alternate wetting and drying (AWD) in rice ## The questionWhat does the research on alternate wetting and drying (AWD) in rice look like, and where does Arkansas fit in it?Map the field from OpenAlex: works per year, the institutions and journals that publish the most on it, and who workswith whom on it in Arkansas. ## The data: OpenAlex- The OpenAlex API, with no key: https://api.openalex.org. Its data is CC0.- Read https://help.openalex.org/api/llm-quick-reference/ before writing any query.- Without a key the allowance is small: $0.10 a day (1,000 credits; a list call costs 1 credit, a search 10). Keep this  whole job under 300 credits. Check what each call costs (the `meta` block and the X-RateLimit headers) and estimate  before paging through results. Prefer `group_by` for counts.- Resolve every institution, journal and author to its OpenAlex ID before filtering by it; never filter by a name.- Log every API URL you call, with its cost and time, in `work/calls.csv`; keep the raw responses you use in `work/raw/`.- Write scripts in `work/scripts/`, not one-off commands, for anything that ends up in a figure or a number: they must  re-run. ## The fieldDefine it with one search, and put the exact query at the top of `field.md`: works on alternate wetting and drying inrice (also called AWD or intermittent irrigation). Look at 20 random results by title to see that the query finds thefield and not something else, and say what you found. ## What to make1. `data/works_per_year.csv`: works per year, 2000 to 2025.2. `data/institutions.csv`: the 15 institutions with the most works in the field.3. `data/venues.csv`: the 10 journals (sources) with the most works.4. `data/arkansas_works.csv` and `data/coauthors.csv`: the works with at least one author at an Arkansas institution   (the University of Arkansas and its Division of Agriculture, Arkansas State University, and the USDA research units   in Arkansas: resolve each to its ID), and who wrote with whom (pairs of authors, with the number of works together).5. `figures/`: one PNG for each view (per year, institutions, journals, the co-author network), plainly labelled.6. `index.html`: one page with the four figures, a short paragraph on each, and under each the query, the date of the   data and the OpenAlex links, so a reader can check it.7. `field.md`: the query, the counts, what was checked, and anything that looked wrong. ## CheckingBefore the page is made, the checker agent re-runs every count with its own OpenAlex calls and compares them with thefiles. Each comparison, and any difference with its explanation, goes in `checks.md`. ## Rules- Use `.venv/bin/python` (it has requests, pandas, matplotlib and networkx). Install nothing.- Write only inside this folder.- Put the date of the data on every figure and on the page.

On Windows, the folder's Python is .venv\Scripts\python.exe, not .venv/bin/python: change that line, or let the agent set up the environment (then take out "Install nothing", and approve each step it asks about).

The checker (.claude/agents/checker.md)

A second agent with its own instructions, in Claude Code's agent folder. In Codex or Antigravity, ask in your request for a second agent (a subagent) that checks the counts with its own calls, and paste these steps into it.

Markdown
---name: checkerdescription: Re-checks the field map's numbers against OpenAlex with its own calls, before anything is published.tools: Read, Bash, Write, Glob, Grep---You check the numbers in data/ against OpenAlex, independently of the scripts that made them. 1. Write your own small script in work/check/ (do not reuse work/scripts/) that makes its own OpenAlex calls: group_by   for the counts, and single works by ID (free) for spot checks.2. Compare every count in data/works_per_year.csv, data/institutions.csv and data/venues.csv with OpenAlex's.3. Check that the works in data/arkansas_works.csv really have an author at one of the Arkansas institutions named in   field.md: fetch 10 of them by ID and read their authorships.4. Check that data/coauthors.csv can be rebuilt from data/arkansas_works.csv. Keep your calls under 60 credits and log them in work/calls.csv. Do not change the data files. Report every comparisonas a table (what was checked, the URL, the value in the file, the value from OpenAlex, the verdict), and explain anydifference.

The request

I'd like a map of the research field on alternate wetting and drying in rice, made from OpenAlex. What I want and the rules are in CLAUDE.md. Please: 1. Read OpenAlex's quick reference, then define the field with one search and check a sample of what it finds.2. Estimate what the calls will cost before you page through anything, and stay inside the budget in CLAUDE.md.3. Build the four views (works per year, the leading institutions, the leading journals, and the Arkansas co-author network), with a script for each.4. Have the checker agent verify the numbers against OpenAlex before you make the page.5. Make index.html with the four figures. When you're done, tell me what it cost, what the checker found, and anything that looked wrong.

Set it up

  1. A folder with CLAUDE.md and .claude/agents/checker.md (above).

  2. Python in that folder, with the four libraries the instruction file names:

    Windows

    python -m venv .venv.venv\Scripts\pip install requests pandas matplotlib networkx

    Mac

    On a Mac:

    python3 -m venv .venv.venv/bin/pip install requests pandas matplotlib networkx
  3. Start your agent in the folder and paste the request. The run in the film took 7 min 32 s.

What the run made

Download the recorded run's files (0.7 MB): the instruction file, the checker, the request (PROMPT.txt) and everything below. Left out: the agent's log and the saved OpenAlex answers (they hold authors' email addresses from their affiliations). To re-run it, run work/scripts/01 to 08 in order (they fetch again from OpenAlex, no key needed), then the checker's work/check/fetch.py and compare.py.

FileWhat it holds
field.mdthe query, the counts, what was checked, and what looked wrong
data/works_per_year.csvworks per year, 2000 to 2025
data/institutions.csv, data/venues.csvthe institutions and journals with the most works (ties kept)
data/arkansas_works.csv, data/coauthors.csvthe works with an Arkansas author, and who wrote with whom
figures/one PNG per view
checks.mdthe checker's comparisons, each with its OpenAlex address
index.htmlthe page: four figures, each with its query, the date of the data and the API addresses
work/calls.csv, work/scripts/, work/check/every call with its cost; the scripts (run them in order: the saved answers are left out of the download)

In this run (OpenAlex, retrieved 4 October 2026): 2,945 works match the query, 2,494 of them from 2000 to 2025; 16 in 2000, 65 in 2010, 325 in 2025. The International Rice Research Institute has the most (138); Agricultural Water Management is the leading journal (109). 44 works have an author at an Arkansas institution (1.5% of the field), with 126 authors and 501 pairs of co-authors. The checker matched every count; it found the year groups add up to 2,920, 25 short of the total (most likely works with no year).

How to check it

  • Open the addresses. Every count comes from a URL in work/calls.csv and under each figure on the page; paste one into a browser and compare.
  • Read the query and the sample. One search defines the field; the run read 20 random results and said how many sit at the edge. Change the query, and the map changes.
  • Read "What looked wrong" in field.md, then the checker's report, before you share anything.
  • Check any name you quote. OpenAlex matches authors and affiliations automatically: in this run some researchers had two or three author IDs, and one Arkansas author had none.

What a map like this can and cannot say

  • The counts are OpenAlex's matches for one search, not a list of everything published, and not a ranking of quality.
  • Recent years are still filling in, and repository and dataset records have grown, so the last years are not directly comparable with earlier ones.
  • An institution's count depends on how OpenAlex matched its authors' affiliations (the University of Arkansas's Division of Agriculture, for example, appears under "University of Arkansas System").

What it costs (as of 4 October 2026)

OpenAlex's data is free to reuse (CC0). The API is free to try with no key: $0.10 of usage a day, which OpenAlex counts as 1,000 credits; a page of results costs 1 credit, a search 10, a single record nothing. A free key raises the daily budget ten times. Over the budget, or above 100 requests a second, the API answers 429 Too Many Requests. This run made 47 calls for 112 credits. Keep a key in an .env file, never in the chat (episode 9).

Not confirmed yet

  • The same run in Codex CLI and Antigravity CLI (only Claude Code was recorded); their way of defining a second agent.
  • OpenAlex's allowances and prices change: read its pricing page on the day.

Sources

Corrections

None so far. If you find something wrong, email doug@douglashutchings.com.

Transcript

Every spoken line, by chapter

A field map

Who studies alternate wetting and drying in rice, where, and with whom? One request, and an agent maps the whole field from OpenAlex.

Then a second agent checks every number, before anyone sees the page.

AI for Research Efficiency. Episode 11: map a research field with OpenAlex.

Last time, an agent fetched research data through APIs, starting with OpenAlex, which needs no key.

This time, one real run, from a question to a checked map and a page: write the rules, let it estimate and fetch, and check before you share.

The rules

The rules go in an instruction file, as in episode 3: the question, the source, and a budget.

Without a key, OpenAlex gives you 1,000 credits a day. A page of results costs 1; a search costs 10. So the file says: stay under 300, and estimate before paging through anything.

It asks for scripts that re-run, and a log of every address it calls, with what it cost.

And it lists what to make: works per year, the leading institutions and journals, and who writes with whom in Arkansas.

The run

Then the request, and the run begins. The clock is the run's own: 7½ minutes, shown faster.

First it reads OpenAlex's guide for agents. Then it writes one small module that every script shares: it logs each call, and keeps each answer.

One search defines the field

Then one search defines the field: alternate wetting and drying, or intermittent irrigation, or AWD, together with rice or paddy.

It finds 2,945 works. Before going on, it reads 20 of them at random.

All 20 are about rice and its water. About 13 are squarely on the practice, and about 7 sit at the edge. So the search finds the field, with a margin around it.

Cost, as it goes

Every answer states its cost, so it counts as it goes. The counts by year, by institution and by journal are one grouped search each: 10 credits apiece.

A wrong field name was refused for free. And to confirm each institution, it fetched single records, which cost nothing.

Three views

Works per year: 16 in 2000, 65 in 2010, and 325 last year.

The institutions that publish most: the International Rice Research Institute leads, with 138. UC Davis is the only one in the United States on the list.

And the journals: Agricultural Water Management, then Field Crops Research.

Arkansas: who writes with whom

For Arkansas, it never filters by a name. It looks up each institution's OpenAlex number first: 11 records, from Fayetteville to Stuttgart and Jonesboro.

44 works have an author there: 1.5% of the field. 126 authors, and 501 pairs.

Each dot is an author, each line a paper written together. The agent sees two groups: the university with the USDA unit in Jonesboro, on methane and water, and the rice research center in Stuttgart.

What looked wrong

Before any check, it writes down what looked wrong.

Some researchers have two or three OpenAlex profiles, so they appear more than once. One paper lists its four authors twice, and the author with the Arkansas address has no profile at all.

And the last few years jump. The agent notes that more datasets and repository records may be part of it, so recent years aren't comparable with earlier ones.

The checker

Then the checker: a second agent, with its own script and its own calls, that doesn't reuse the first one's code.

Every count matched: 26 years, 16 institutions, 11 journals, the 44 works, and all 501 pairs, rebuilt from scratch.

It found one thing: the yearly counts add up to 25 fewer than the total. Most likely, those works have no year. The agent added that to its notes.

The page

Only then the page: four figures, each with its query, the date of the data, and the addresses to check it.

It's on your computer for now. To share it, put the folder on GitHub and publish it with Vercel, as in episodes 5 and 6.

Read it for what it is

Read the map for what it is. One search defines the field. The counts are OpenAlex's matches, not a ranking of quality. And check any name you quote.

The whole run: 7½ minutes, 47 calls, and 112 credits of the day's 1,000. A free key gives 10 times as many.

Try it

Pause here and try it: copy the instruction file from the companion page, put in your own question and field, and read the checker's report before you make the page.

Review and next

So: define the field with one search, estimate before you page, keep the scripts, and check every number before you share.

Next: make a narrated explainer of your paper.

AI for Research Efficiency, with Douglas Hutchings. Narration: an AI-generated voice (ElevenLabs).