WEBVTT

00:00:12.700 --> 00:00:16.591
An agent made this explainer from
one paper, in about five minutes:

00:00:16.902 --> 00:00:20.351
the script, the slides,
the captions, and the video.

00:00:20.781 --> 00:00:23.247
Every sentence points
to the page it came from.

00:00:23.638 --> 00:00:26.825
And the paid voice? The
agent never saw the key.

00:00:29.200 --> 00:00:32.250
AI for Research Efficiency.

00:00:32.300 --> 00:00:35.610
Episode 12: Make a narrated
explainer of your paper.

00:00:38.700 --> 00:00:41.946
Last time, you mapped a
research field with OpenAlex.

00:00:42.485 --> 00:00:47.035
This time, one paper becomes a short
narrated video, for a lab's page,

00:00:47.085 --> 00:00:52.254
a class, or a grant's broader
impacts. One real run, replayed.

00:00:55.700 --> 00:00:57.382
The folder: the paper,

00:00:57.512 --> 00:01:00.326
an instruction file,
and three empty folders,

00:01:00.476 --> 00:01:03.420
for the work, the voice,
and the finished video.

00:01:04.130 --> 00:01:07.398
The paper is a 2025 review

00:01:07.448 --> 00:01:11.347
of water management in rice,
licensed CC BY:

00:01:11.397 --> 00:01:14.984
anyone may adapt it, with credit.
It's the review behind episodes 3 and 4.

00:01:15.834 --> 00:01:17.837
The instruction file sets the rules.

00:01:18.248 --> 00:01:22.806
About 150 words, one sentence per
line, each ending with its page.

00:01:23.237 --> 00:01:26.682
Only what the paper says,
and every number as printed.

00:01:27.512 --> 00:01:32.501
And one rule for the voice: the agent
has no key, and must not look for one.

00:01:32.933 --> 00:01:35.417
It writes a request,
and waits for the audio.

00:01:38.900 --> 00:01:42.385
The prompt asks for the script first,
checked against the paper;

00:01:42.776 --> 00:01:45.921
then the voice request,
the slides, and the video.

00:01:46.611 --> 00:01:51.394
In its first 10 seconds, the agent splits
the paper into one text file per page,

00:01:51.605 --> 00:01:53.906
so each sentence can carry its page.

00:01:54.825 --> 00:01:57.287
35 seconds in, a first draft:

00:01:57.597 --> 00:02:01.660
215 words, too long
for a minute. It cuts.

00:02:02.570 --> 00:02:03.572
Two more drafts.

00:02:03.983 --> 00:02:06.619
Then it checks each figure
against the page it cites,

00:02:06.710 --> 00:02:08.533
and writes down every check.

00:02:09.423 --> 00:02:12.470
At 82 seconds: 159 words

00:02:12.520 --> 00:02:15.466
in 13 sentences, each with its page.

00:02:17.300 --> 00:02:22.620
Methane fell by 31 to 62%:
page 1, in the abstract.

00:02:23.737 --> 00:02:25.760
Both gases together,
33.6 to 56.2% lower:

00:02:25.810 --> 00:02:29.836
page 3.

00:02:30.907 --> 00:02:34.813
Safe AWD raised yield by 8.9%,
over a 5-year experiment:

00:02:34.863 --> 00:02:37.546
page 5.

00:02:38.956 --> 00:02:43.220
Page 5 says one thing more:
the experiment was at a single site.

00:02:43.591 --> 00:02:46.994
The first draft said so;
cutting for length took it out.

00:02:48.270 --> 00:02:50.792
The numbers are right,
but the context is thinner.

00:02:51.223 --> 00:02:53.095
That's the check only you can make:

00:02:53.445 --> 00:02:56.347
does each sentence still
mean what the paper means?

00:02:58.100 --> 00:02:59.001
Now the voice.

00:02:59.471 --> 00:03:03.825
The agent writes a request: the text,
a stock voice, and the model.

00:03:04.215 --> 00:03:05.196
Then it waits.

00:03:05.961 --> 00:03:09.406
You make the paid request yourself,
outside the agent's session,

00:03:09.536 --> 00:03:14.863
with the key from your own .env file. The
companion page has a short script for it.

00:03:15.699 --> 00:03:21.038
First an estimate:
966 characters, about 107 credits.

00:03:22.151 --> 00:03:24.859
About 20 seconds later,
the audio is back,

00:03:25.070 --> 00:03:30.124
with the start and end of every
character. It cost 106 credits.

00:03:30.843 --> 00:03:34.954
The agent never saw the key.
It saw two files appear.

00:03:38.900 --> 00:03:41.985
Meanwhile,
the agent drew 12 slides in Python:

00:03:42.356 --> 00:03:45.221
one idea each, the key number large.

00:03:45.942 --> 00:03:48.589
It looked at them all at once,
and fixed what it saw:

00:03:49.041 --> 00:03:54.015
a clipped number, a drawing over a title,
and text running off two slides.

00:03:54.926 --> 00:04:00.477
The character timings become captions:
a sentence at a time, at most two lines.

00:04:01.348 --> 00:04:06.462
Then ffmpeg puts it together: each
slide while its sentences are spoken,

00:04:06.713 --> 00:04:09.159
the voice,
and the captions burned in.

00:04:10.050 --> 00:04:14.555
It checks the result: the length
matches the voice, every slide appears,

00:04:14.785 --> 00:04:19.070
and the captions can be read. It made
their boxes darker until they could.

00:04:19.756 --> 00:04:22.966
5 minutes and 7 seconds,
start to finish.

00:04:26.900 --> 00:04:31.545
Its report is candid:
the video runs 89 seconds, not 60.

00:04:31.956 --> 00:04:36.620
This voice reads slowly, and the
paper's address alone takes 12 seconds.

00:04:37.566 --> 00:04:40.855
And it names what it could not check:
it cannot listen.

00:04:41.266 --> 00:04:45.557
So listen yourself; every number should
sound the way the paper prints it.

00:04:46.567 --> 00:04:49.790
Before you share it:
each sentence against its page,

00:04:49.960 --> 00:04:52.993
the voice against the numbers,
the credit to the paper,

00:04:53.283 --> 00:04:56.105
and a line that says the voice is AI.

00:04:58.100 --> 00:04:59.441
About the voice service.

00:04:59.871 --> 00:05:02.583
ElevenLabs' free plan
can't be used commercially,

00:05:02.753 --> 00:05:07.516
and anything you publish from it
must have elevenlabs.io in its title.

00:05:08.181 --> 00:05:10.364
Paid plans include a
commercial license,

00:05:10.454 --> 00:05:14.259
and what you make during a subscription
stays licensed after it ends.

00:05:14.970 --> 00:05:17.434
Never clone someone's voice
without their consent,

00:05:17.644 --> 00:05:20.067
and never hide that a voice is AI.

00:05:21.073 --> 00:05:25.966
This series is narrated with ElevenLabs,
on a paid plan we use and value.

00:05:26.277 --> 00:05:30.046
A free voice that runs on your own
computer is on the companion page.

00:05:31.700 --> 00:05:34.111
Here is what the run made, in full.

00:07:07.700 --> 00:07:11.867
Try it with your own paper: copy the
instruction file from the companion page,

00:07:12.118 --> 00:07:13.450
ask for the script first,

00:07:13.701 --> 00:07:17.547
and check every sentence against
its page before you make the voice.

00:07:22.100 --> 00:07:26.103
So: pin every sentence to its page,
keep the key with you,

00:07:26.433 --> 00:07:30.136
listen before you share,
and say the voice is AI.

00:07:31.500 --> 00:07:34.078
Next: how this series is made.
