A podcast clip maker that reads the second hour too

Two hours of conversation is not one long transcript to skim. It is four windows read in parallel, three minutes of overlap between them, and one blind pass that ranks the finalists against each other.

Windows, not one long read

What a single-pass podcast clip generator misses

Read in one call, a two-hour transcript anchors on what it saw first. Measured across two full episodes, every pick landed in the first 57% of the video and nothing after it. The windows exist because of that measurement, and they are the reason a ninety-minute answer can still win.

Measured, then fixed
Every pick came from the first 57%

About thirty minutes each

A two-hour episode becomes four windows, and six is the cap. Each is read independently and in parallel, and no single window may put forward more than five candidates.

Three minutes of bleed

Windows overlap on every inner edge, so a moment sitting across a boundary is whole inside one of them instead of being halved by both.

Best of each, before second of any

Local scores are only comparable inside the window that produced them, so the moments go forward in rotation — one from every window before a second from any.

One judge, reading blind

A final pass sees the finished candidates side by side, with no timecode and no clue which part of the episode each came from, and ranks them on one shared scale.

Four stages

Between a two-hour transcript and a podcast clip

Finding a moment and cutting it are two jobs. They used to be one call, and the cut was decided by the stage least equipped to decide it — so they were separated, and a fourth stage was added to check the result.

Discover

The sentence where the point lands

Per window, in parallel. It returns an anchor and one line saying what the moment is — and no boundaries at all, so a good moment is never lost because no tidy range happened to fit around it.

Cut

Ninety seconds either side

One call per moment, reading the neighbourhood closely. A clip may be assembled from up to three pieces in order: the setup from a minute earlier, the moment itself, and the detour between them dropped.

Cold read

The clip’s own lines, nothing else

No transcript around it, no title, no timecodes. If a stranger cannot follow it, the report goes back and the cut is made again. One that still does not hold stays in the list capped at 60, with a line naming what is missing.

Rank

Hook, flow, value, travel

Four scores set by the only stage that has seen every finalist together, plus the reason it gave. That is what you read before you decide anything.

The model is never asked for a number of seconds. It quotes the words it wants to open and close on, and the times are read off those words — so a boundary can land anywhere, but never inside a word. A cut over 100 seconds is refused and made again rather than trimmed down to fit.

The intro is read, not skipped by the clock

An episode opens on material that must never be clipped: a cold-open teaser cut from later in the conversation, the show open, a sponsor read at the top, the pleasantries while the guest settles in. A fixed “ignore the first sixty seconds” is wrong in both directions — some shows are talking at 0:20, some are still doing housekeeping at 4:00 — and a teaser is the worst case of all, because it is made of the strongest lines in the episode, out of order.

So a model reads the opening — up to ten minutes of it, or ninety sentences — and reports the first sentence of real content. Everything before that is hidden from selection, and any candidate that would start inside it is dropped even after the cut. If the answer would swallow more than 80% of the episode it is ignored. It is on by default; you can pin it to a number of seconds or turn it off per video.

What the four scores mean
Start to finish

How to clip a podcast here

Four steps, and you are the third one. Nothing is rendered, and nothing is published, until you say so.

1 — The recording

If the episode also went up on YouTube, paste the watch address and skip the six-gigabyte round trip — it is fetched once on our side. If it lives anywhere else, on Riverside, Zencastr, Spotify or an RSS feed, it is an export and an upload. Reading two hours transfers about 60 MB: the audio, not the video.

2 — Six candidates

Six by default, anywhere from one to twelve if you ask. Each arrives with its score, its timecodes and the sentences that made it a candidate.

3 — You keep or you pass

Move the in and out points if you disagree. Dragging a handle clears an assembled multi-piece cut, because those gaps were reasoned about on a window that no longer exists.

4 — Only then, the cut

Just the seconds you approved are downloaded. Forty seconds of video fetched for a forty-second clip, and the MP4 lands in your library.

Links are YouTube only — youtube.com and youtu.be, watch pages, Shorts, live and embed — so for a podcast host, a shared drive or a meeting recording the way in is the file. A paste asks you to confirm the episode is yours to use before the import starts, and stores the sentence you were shown, the version of the terms, the time and the address. Imports are capped by the day as well: three on the trial, twenty on Starter, fifty on Creator, a hundred on Studio. Uploads are not capped.

The render

From podcast to shorts, two looks ship by default

A house style rather than a settings panel: you choose the shape once and everything you approved comes out that way.

Boxed 1:1, the text behind the speaker

The subject is matted out and composited over the words, so the type sits behind them rather than on top. Keywords as they land, the whole passage, or both at once.

Full 9:16, word-by-word captions

Edge-to-edge vertical with karaoke highlighting. Two people across a table is the hard case, and it is handled by speaker tracking rather than by a fixed centre crop.

Framing follows whoever is actually speaking, and a banner guard measures the source for a lower third baked into the footage — it punches in only when it finds one, so a chyron on your show does not narrow every clip that never had one. The one combination that does not exist is 9:16 with the text behind: the compositing geometry is written against the 1:1 box, and the options panel corrects the pairing instead of failing the job.

How much a plan reads

Reading a video costs credits, in proportion to how long it runs; a clip does not, until a generous per-plan batch is used up. There is no cap on how long a single episode may be, on any plan. Starter reads 240 credits a month — about four hours — with 80 rendered clips included, then three credits each past that. Creator reads 660 credits, about eleven hours, with 220 rendered clips included. Studio reads 1,620 credits, about twenty-seven hours, with 540 rendered clips included, then three credits each past that. The free trial is thirty credits, once — about thirty minutes of video, ten rendered clips included, no card.

Re-running the selection on an episode you have already analysed costs nothing: the transcript is cached against the source, so asking for twelve moments instead of six, or turning the intro detection off and looking again, spends no minutes. Only a fresh transcription is charged a second time.

Questions

Written for someone with a two-hour episode already recorded.

How long an episode can it take?

There is no cap on how long a single episode can be, on any plan — a four-hour episode just spends four hours of credits. What differs by plan is the credits themselves: 30 once on the free trial (about 30 minutes), 240 a month on Starter (about 4 hours), 660 on Creator (about 11 hours), 1,620 on Studio (about 27 hours). A batch of clips comes with each plan too — 10, 80, 220 and 540 — and only past that does an export cost 3 credits.

How many clips does it find in one episode?

Six by default, and you can ask for anywhere between one and twelve. It returns fewer when the material does not hold that many: a weak moment included to hit a number takes the slot a good one could have had.

Does it post the clips for me?

No. There is no integration with YouTube, TikTok, Instagram or any podcast host, and nothing is published on your behalf. The finished MP4 waits in your library for you to download and upload wherever you want it.

Is this a podcast clip editor?

Not in the timeline sense. What you edit is the decision: which candidates survive, where each one starts and ends, and the style everything is rendered in. There is no track, no keyframe, and no way to re-cut inside a finished clip by hand.

How does it frame two people sitting across a table?

Speaker tracking runs on the approved window and the crop follows whoever is actually talking, so the frame moves with the conversation instead of sitting on a wide two-shot. A banner guard measures the source for a lower third baked into the footage and punches in only when it finds one.

Does running the analysis again cost minutes?

No, as long as the transcript is still cached against that source. Changing the number of moments, or the intro handling, and looking again is free; only a fresh transcription is charged again.

Which part is the AI actually deciding?

Which passages are worth publishing, where each one starts and ends, whether a stranger could follow it, and how the finalists rank against each other. It never decides that a clip is finished — that is the one judgement it hands back to you, and nothing renders before you make it.

Start with the episode whose best moment is at ninety minutes