A Riverside alternative for the clips, not the recording

Riverside runs the interview and hands you the files. This one starts at the recording: upload the export, or paste the YouTube address if the episode is already published, and it ranks every passage worth cutting before it cuts any of them.

Where the overlap stops

It does not record anything

No studio, no guest link, no separate tracks, no local capture. The first thing it asks for is a recording that already exists, as a file or as a YouTube address — so these two are substitutes for each other on the last step only, and on nothing before it.

The recording is ahead of you
Then keep what records it

A remote interview needs something capturing both ends of the call. Nothing here does that, and no feature list further down this page changes it.

The episodes already exist
Then this is the half that is left

Recorded anywhere, edited anywhere, published or not. It reads the file and hands back the passages worth cutting, ranked, before it cuts one.

Both are true
Then they stack, in that order

Export the episode from wherever the session was recorded, upload the file, and the moment list is the next step. Neither one replaces the other.

What differs once the recording exists

Both of them cut short clips out of a long conversation, and that is the whole of the overlap. Everything above the last row belongs to one of them alone.

What differsInTheClipsRiverside
What it is firstA reader. It starts from a recording that already exists and never touches a camera or a microphone.A recording studio in a browser: a link for the guest, and a track captured on each end.
How a video gets inTwo ways, and neither is the lesser one: a file you send — mp4, mov, mkv or webm, up to 8 GB — or a YouTube address you paste, which is read without a copy of the video being kept — audio for the transcription, then only the seconds each render needs. Links are YouTube only and counted by the day; uploads are not counted, and a file does not depend on an address staying alive. No cap on how long a single video may be, on any paid plan.It is already there, because the session that made it ran on the same platform.
Before a clip existsA ranked candidate list: hook, flow, value and travel, the timecodes, and the sentences that made each one a candidate. Nothing renders until you keep it.Clips are produced from the recording and reviewed in the editor the recording landed in.
What shipsFive shapes, all fixed rather than laid out by hand: boxed 1:1 with words behind the speaker, full 9:16 with word-by-word captions, and three more built for footage with a camera over it. No layout editor, no free placement of a caption block.Clips in several aspect ratios, with captions and layout adjusted in its own editor.
What you pay forCredits, one balance: a minute of source analysed spends one, and an export past the plan’s included batch spends three.Its own plans, on its own site — a number copied into this table would be stale by the time you read it.

If the recording matters to you more than the selection does, the first row is an argument for keeping what you have and pointing this at the exports. Their column is from Riverside’s own public description of the product, read on ; ours is from the code that runs this one.

Once it has the video

A finished episode, read end to end

A two-hour interview is not read in one pass. It is split into roughly thirty-minute windows scored in parallel with three minutes of overlap, and then one judge ranks the pooled finalists — because a single pass anchors on what it reads first and never picks from the back half.

Only the audio is read to transcribe it

ffmpeg pulls the audio stream out of the stored file at 16 kHz mono, so a two-hour export is transcribed without the video being fetched for it.

The intro is detected and left out

A model reads the opening and finds where the content starts, so the ranking is not spent on the cold open, the sting and the welcome.

A boundary never lands inside a word

The model names the words it opens and closes on, verbatim, and the times are read off those words. Anywhere is allowed except mid-word.

The frame follows whoever is talking

Speaker tracking is TalkNet-ASD, so a two-up interview exported as one file stays on the person speaking rather than on whoever is in shot.

Delivery

What comes out, and what stays yours

Boxed 1:1, text behind. A square of the speaker on a full-height canvas, with the words composited behind them: the subject is matte-cut and sits over the text, so the line reads and the face is never covered.

Full-frame 9:16, karaoke. Edge to edge vertical, with word-by-word highlighting timed off the transcript. Pick one, save it as a preset, and every moment you approve comes out the same way.

Banner guard. It measures the footage for a lower third baked into the export and punches in only if it finds one — which is worth knowing if your episodes ship with a name strap already burned in. Left on always, it would cost you framing on every clip that did not need it.

Nothing is posted for you. The MP4 lands in your library and downloads in one click. No account is connected to anything, so where the clip goes and when is a decision that never leaves your hands.

Read how the look is built

Questions

The ones worth answering before you move a two-hour file across the internet.

Can it record my interview the way Riverside does?

No. There is no studio, no guest link and no local track on each end. It starts from a recording you already have, which is why it is only an alternative for the clipping half of that job.

How does an episode get in?

Two ways. Paste the YouTube address if the episode is already published there — a watch page, a Short, a live URL or an embed — and no copy of it is kept here: the transcription reads an audio-only stream, and each render afterwards fetches only the seconds of its own window, so a two-hour recording never has to be downloaded in order to be uploaded again. The trade is that a link goes on depending on the episode staying at that address. Or send the file yourself: mp4, mov, mkv or webm, up to 8 GB. You confirm the video is yours to use before an import starts, and imports are counted by the day — three on the trial, twenty on Starter, fifty on Creator, a hundred on Studio — while uploads are not counted. Links are YouTube only, so a Riverside export comes in as a file.

How long can the episode be?

There is no cap on any paid plan — Starter, Creator and Studio all read a video of any length; it simply spends credits by the minute. The free trial is thirty minutes of video, granted once. A two-hour interview spends half of the four hours Starter includes each month.

It is a two-person interview. Which face does it follow?

Whoever is speaking. Speaker tracking is TalkNet-ASD, so the crop moves with the voice rather than sitting on whoever happens to be on screen. One upload is one video, though: there is nothing that recombines separate participant tracks, so send the composed export rather than the raw files.

Will it only pick moments from the first twenty minutes?

No, and that is the reason for the windows. Thirty minutes at a time are scored in parallel with three minutes of overlap so nothing falls between two of them, and a final pass ranks the finalists from every window against each other rather than against their neighbours.

How long do you keep the file I upload?

As long as the plan keeps the clips: seven days on the trial, ninety on Starter, a year on Creator, and on Studio for as long as the subscription stays active. That way you can still cut a new moment out of an episode whose clips are still there. Deleting a source removes the file, its transcript and its clips together.

The episodes are recorded. The clips are still inside them.