An AI video clipping tool that argues its case before it cuts anything
Most of this category hands you finished files with a score on each and asks you to trust the number. This one reads the whole recording, ranks every passage worth cutting, and shows you the sentences and the reason behind each candidate — before a single clip is rendered. And there is no generation anywhere in it: every word in a clip is a word somebody actually said.
A list you read, not a folder you did not ask for
The output of the AI part of this is a ranked list of candidates. The output of the product is whichever of them you decide to keep.
A reason, not just a score
Every candidate carries four scores — hook, flow, value and travel — plus a one-line reason and the sentences it was drawn from. A number with nothing behind it is not an argument, it is a guess dressed up as one.
Checked before you see it
A second pass reads only the clip’s own lines — no title, no surrounding transcript, no timecodes — and if a stranger could not follow it, the cut is made again. One that still fails keeps its place with its score capped and a line naming what is missing, rather than being quietly removed.
Zero files until you decide
Keep it, pass on it, or drag the in and out points yourself. Nothing is rendered before that decision, so there is no folder of clips to sort through — there is a list, and then there is what you kept.
It reads what was recorded. It does not write anything new
There is no script, no stock footage, no synthetic voice and no avatar. A candidate is a passage that already exists in the video you gave it. The transcript is read, ranked and quoted from — never written. If a sentence is not already in your recording, it cannot end up in a clip.
That is a limit against one kind of search and a promise against another. If what you want is a video assembled out of a prompt with nobody ever filmed, this is the wrong tool and no page here will pretend otherwise. If what you want is your own footage, judged on its own words and timed to the frame, that is the entire product.
- Transcription runs in over thirty languages, and captions are timed word by word in the language spoken
- There is no translation and no dubbing — the words on screen are the words in the audio
- The transcript downloads as SRT, WebVTT or plain text, and for a clip it is cut to what actually plays — filler removed, timings closed up — so it lines up with the MP4 beside it
Two ways in: a YouTube link, or the file itself
A watch page, a Short, a live URL or an embed — youtube.com or youtu.be — is fetched once on our side. Anything else is a file: mp4, mov, mkv or webm, up to 8 GB.
The link path asks one question
You confirm the video is yours to use before the import starts — refused rather than assumed — and what is stored is the sentence you were shown, the version of the terms, the time and the address.
The upload path asks nothing
A file you already have is already yours to work with, so there is no declaration to sign. It goes straight from your browser into storage.
One is capped by the day, the other is not
Link imports are limited per day by plan — three on the trial, twenty, fifty, a hundred — because pasting a link costs you nothing and costs us a fetch. An upload rate-limits itself on your own connection.
How a recording becomes a ranked list
A two-hour transcript read in one pass anchors on what it saw first. This is built against that failure specifically.
The opening is read, not guessed at
A model reads up to the first ten minutes and reports where the actual content starts, rather than skipping a fixed sixty seconds and hoping. Everything before that point is excluded from selection.
Windows, not one long read
The rest is split into windows of about thirty minutes, read independently and in parallel, with three minutes of overlap on every inner edge so nothing sitting on a seam gets cut in half by both sides.
Best of each before second of any
Candidates go forward in rotation across the windows, then a single pass ranks the pooled finalists against each other on one shared scale.
The boundary is a word, never part of one
The model quotes the exact words it wants to open and close on, and the times are read off those words — so a cut can land anywhere except inside one.
Five shapes, two caption treatments, and one crop that follows the voice
A preset carries the two shapes that need nothing extra; the other three appear only where a camera was actually found in the picture.
Boxed 1:1 and Full 9:16
The two standing shapes, and the only pair a saved preset can carry. Text-behind captions ship on Boxed 1:1 only — that pairing on 9:16 is genuinely unbuilt.
Cam bubble and Face over screen
Both need a camera composited over the picture, and both are offered only where the shape watcher found one in that clip.
The whole picture
The source dropped into a vertical frame and otherwise left alone, its own camera where it already was.
The frame follows whoever is talking
Speaker tracking (TalkNet-ASD) runs on both default renders, so a two-person conversation keeps the crop on the person actually speaking.
Karaoke captions are editable — font, size, colour, outline, shadow, capitalisation — and can be saved as a named style you reuse. What is fixed is position: no free placement, no logo layer.
An MP4, or the cut without the look
A clip leaves as an MP4 by default, MP4 is the finished file most people want. It can also leave as an XML timeline or a CMX3600 EDL for Premiere or Resolve instead — the cut and the assembly, not the composited captions or the per-frame reframe, since no timeline format can hold a decision made per frame. That is a timeline export, not a timeline editor: there is no track view and no keyframes inside this product.
A publishing calendar holds a date and a status, not a connection. You can plan a clip onto a day and mark it published once it is, with a link if you want one — but nothing is posted anywhere for you, and no account credentials for Instagram, TikTok or anywhere else are ever asked for or stored.
A credit buys a minute of reading, and clips are not free past a point
One credit is one minute of source analysed. Every plan also includes a generous batch of clip exports — about twenty per hour of allowance — and only an export past that batch spends credits too, at three each. Reading the video is still what the price is mostly buying: a forty-second clip out of a two-hour recording costs the same three credits whether the recording was ten minutes or four hours long, because rendering only ever touches the seconds you approved.
- Free trial: 30 credits once, no card, does not renew — 10 rendered clips included, exports marked, clips kept 7 days
- Starter, $19 a month: 240 credits (4 hours of video), 80 rendered clips included then 3 credits each, any length of video, clips kept 90 days
- Creator, $49 a month: 660 credits (11 hours), 220 rendered clips included, clips kept a year
- Studio, $99 a month: 1,620 credits (27 hours), 540 rendered clips included, clips kept while the subscription is active
- No paid plan watermarks anything — only the free trial marks its exports
- A year, where offered, is billed for eight months of the plan
Questions people ask before picking a tool in this category
What makes this different from other AI video clipping tools?
It hands back a ranked list of candidates — each with four scores, the sentences behind it and a one-line reason — before anything is rendered, rather than a batch of finished files you have to judge after the fact. You keep, pass, or adjust each one, and nothing is cut until you say so.
Does this generate a video, or only cut from footage I already have?
Only cuts. There is no script-to-video, no stock footage, no synthetic voice and no avatar anywhere in it. Every clip is built from seconds that were actually recorded, in the recording you gave it.
How does the AI decide what to clip?
It transcribes the whole recording, finds and excludes the intro, then reads the rest in roughly thirty-minute windows in parallel with three minutes of overlap between them. Each candidate is scored on hook, flow, value and travel, checked by a cold read that sees only its own lines, and re-cut if a stranger could not follow it.
Can I paste a YouTube link, or do I have to upload a file?
Either. A youtube.com or youtu.be address — a watch page, a Short, a live URL or an embed — is fetched once on our side after you confirm the video is yours to use. An upload asks nothing and takes mp4, mov, mkv or webm up to 8 GB.
What do I actually get back — an MP4, or something else?
An MP4 by default. If you want to finish the edit somewhere else, it can also leave as an XML timeline or a CMX3600 EDL for Premiere or Resolve — the cut, not the composited look, since a timeline cannot hold a decision made per frame.
Does it post the clips to social media for me?
No. A publishing calendar lets you plan a date and mark a clip published, but it holds no account credentials for any platform and posts nothing on your behalf. The MP4 sits in your library and you upload it yourself.
Is there a free way to try it?
Thirty minutes of video, once, with no card. Ten clips are included, every shape and caption style is available, and the trial does not renew. Trial exports carry a small mark; no paid plan does.
The same engine, described from a more specific angle
This page argues the category. These argue one way of using it.

AI reel generator
It generates nothing. Every word in a clip is a word somebody actually said — which is the point, or the deal-breaker.

Clipping software
What the whole pipeline does, end to end: transcript, ranked candidates, your decision, then the cut.

YouTube clip maker
Paste the watch URL and it is fetched once on our side, or upload the file. Only the passages you approve are ever cut.

Highlight video maker
Finding the highlights is the whole product. It does not stitch them into a reel — each one is its own clip.