AI for Research Efficiency, episode 12: Make a narrated explainer of your paper Narration: an AI-generated voice (ElevenLabs). [0:00] From a paper to a video This is a 2025 review by Kazunori Minamikawa in Paddy and Water Environment. Flooded rice fields release methane, a greenhouse gas made by soil microbes. An agent made this explainer from one paper, in about five minutes: the script, the slides, the captions, and the video. Every sentence points to the page it came from. And the paid voice? The agent never saw the key. AI for Research Efficiency. Episode 12: Make a narrated explainer of your paper. Last time, you mapped a research field with OpenAlex. This time, one paper becomes a short narrated video, for a lab's page, a class, or a grant's broader impacts. One real run, replayed. [0:55] 1. The folder and its rules The folder: the paper, an instruction file, and three empty folders, for the work, the voice, and the finished video. The paper is a 2025 review of water management in rice, licensed CC BY: anyone may adapt it, with credit. It's the review behind episodes 3 and 4. The instruction file sets the rules. About 150 words, one sentence per line, each ending with its page. Only what the paper says, and every number as printed. And one rule for the voice: the agent has no key, and must not look for one. It writes a request, and waits for the audio. [1:38] 2. A script, page by page The prompt asks for the script first, checked against the paper; then the voice request, the slides, and the video. In its first 10 seconds, the agent splits the paper into one text file per page, so each sentence can carry its page. 35 seconds in, a first draft: 215 words, too long for a minute. It cuts. Two more drafts. Then it checks each figure against the page it cites, and writes down every check. At 82 seconds: 159 words in 13 sentences, each with its page. [2:16] 3. Every sentence, pinned to its page Methane fell by 31 to 62%: page 1, in the abstract. Both gases together, 33.6 to 56.2% lower: page 3. Safe AWD raised yield by 8.9%, over a 5-year experiment: page 5. Page 5 says one thing more: the experiment was at a single site. The first draft said so; cutting for length took it out. The numbers are right, but the context is thinner. That's the check only you can make: does each sentence still mean what the paper means? [2:57] 4. The voice, and the key Now the voice. The agent writes a request: the text, a stock voice, and the model. Then it waits. You make the paid request yourself, outside the agent's session, with the key from your own .env file. The companion page has a short script for it. First an estimate: 966 characters, about 107 credits. About 20 seconds later, the audio is back, with the start and end of every character. It cost 106 credits. The agent never saw the key. It saw two files appear. [3:38] 5. Slides, captions, video Meanwhile, the agent drew 12 slides in Python: one idea each, the key number large. It looked at them all at once, and fixed what it saw: a clipped number, a drawing over a title, and text running off two slides. The character timings become captions: a sentence at a time, at most two lines. Then ffmpeg puts it together: each slide while its sentences are spoken, the voice, and the captions burned in. It checks the result: the length matches the voice, every slide appears, and the captions can be read. It made their boxes darker until they could. 5 minutes and 7 seconds, start to finish. [4:26] 6. What you check Its report is candid: the video runs 89 seconds, not 60. This voice reads slowly, and the paper's address alone takes 12 seconds. And it names what it could not check: it cannot listen. So listen yourself; every number should sound the way the paper prints it. Before you share it: each sentence against its page, the voice against the numbers, the credit to the paper, and a line that says the voice is AI. [4:57] 7. The voice service's terms About the voice service. ElevenLabs' free plan can't be used commercially, and anything you publish from it must have elevenlabs.io in its title. Paid plans include a commercial license, and what you make during a subscription stays licensed after it ends. Never clone someone's voice without their consent, and never hide that a voice is AI. This series is narrated with ElevenLabs, on a paid plan we use and value. A free voice that runs on your own computer is on the companion page. [5:31] The explainer, in full Here is what the run made, in full. This is a 2025 review by Kazunori Minamikawa in Paddy and Water Environment. Flooded rice fields release methane, a greenhouse gas made by soil microbes. Draining fields now and then, as in alternate wetting and drying, or A W D, cuts it. The review gathers 11 meta-analyses, studies that pool many experiments. Compared with continuous flooding, methane fell by 31 to 62%. But nitrous oxide, another greenhouse gas, rose by 37 to 445%. Counting both gases, their combined climate impact fell by 33.6 to 56.2%. Rice yield ranged from a 5.4% loss to an 11% gain. Yields generally held if water fell no more than 15 cm below the soil surface. With this milder approach, called safe A W D, significant yield losses are unlikely. In Vietnam's Mekong Delta, safe A W D raised mean yield by 8.9% over a 5-year experiment. Tuned to the crop and place, water management can save water, cut emissions, and keep or raise yields. Read the paper at doi dot org slash ten point one zero zero seven slash s one zero three three three dash zero two five dash zero one zero four five dash four. [7:07] Try it Try it with your own paper: copy the instruction file from the companion page, ask for the script first, and check every sentence against its page before you make the voice. [7:21] Review and next So: pin every sentence to its page, keep the key with you, listen before you share, and say the voice is AI. Next: how this series is made.