Video to text, and the text is a file you keep
The whole recording is read end to end and comes back as words with timings on them — download it as a subtitle file, a WebVTT track or plain prose. The same read also hands you the passages worth cutting, ranked, so one pass answers both questions.
Three files, and none of them costs a credit
The words come out of the product, not just onto the screen. A recording that has been analysed has a transcript behind it, and the Download button on the transcript pane hands it over in three shapes: SRT for anything with an upload form, WebVTT for a player on your own site, and plain text for reading, pasting or searching. Nothing is charged for any of them, on any plan, because there is nothing to charge for — the file is arithmetic over a transcript already paid for once.
The timings are measured, not divided up afterwards. Every word carries its own start and end, straight from the read, which is what the captions in a clip are drawn from. A subtitle file built that way lands on the syllable rather than on an even share of a paragraph, and that difference is the whole reason a transcript is worth having as a file rather than as a wall of prose.
Cues are re-cut for reading, because a speaking turn is not a subtitle. The read hands back speaking turns, and one of those is routinely thirty seconds and four hundred characters — a fine paragraph and a useless caption. The export cuts them into lines of about forty characters, two at most, seven seconds at most, broken where the sentences break and timed from the words themselves.
A clip gets its own file, cut to what actually plays. Export the subtitles beside a rendered clip and they are mapped onto the clip’s own clock — filler you removed is gone and everything after it moves up. A file that merely subtracts the start time is right until the first cut and drifts for the rest of the video, which nobody notices until it is posted.
A YouTube address or a file, and one thing neither of them is
Both ways in are first-class. What is not offered is a URL from anywhere else, and saying so here is cheaper than finding out after you have pasted one.
Paste a YouTube address
A watch page, a Short, a live URL or an embed, from youtube.com or youtu.be. No copy of the video is kept: the read takes an audio-only stream, and a render later fetches only the seconds of its own window.
Or send the file
mp4, mov, mkv or webm, up to 8 GB, from the browser straight to storage. A file does not depend on an address staying alive, which a link does — that is the honest trade between the two.
Another platform’s URL is not a way in
A TikTok, Twitch, Vimeo, Zoom or Drive address is refused. For those the way in is the file, exported from wherever it lives and uploaded here.
The rights question comes first
An import asks you to confirm the video is yours to use before it starts, and stores the sentence you were shown, the version of the terms, the time and the address.
Links are counted by the day, uploads are not
Three a day on the trial, twenty on Starter, fifty on Creator, a hundred on Studio. Pasting costs you nothing and costs us a fetch; an upload limits itself on your own connection.
Any length of recording
No paid plan caps how long a single video may be. What bounds the work is the credit balance, and a four-hour file costs exactly what four hours of shorter files cost.
Four things that are in the file, and one that is not
The parts worth knowing before you decide whether this is the transcription tool for you.
Read once. Everything else on this page is something done with that one read.
Over thirty languages, detected rather than declared
The language is worked out from the audio unless you name one, and the words come back in the language spoken.
Speaker turns, where there is more than one voice
Turns are labelled and carried through into the plain-text export as paragraphs with a name and a timestamp — which is what makes an interview readable rather than a single block.
A misheard word can be corrected, and it costs nothing
Replace what is between two marks and the replacement occupies exactly the seconds the old words did. Nothing is re-transcribed, no model is called, no credits are spent, and the next export carries the fix.
What it does not carry: another language
There is no translation and no dubbing anywhere in this product. The words in the file are the words in the audio, and a page here that suggested otherwise would be selling something that does not exist.
The same read also tells you which forty seconds are worth posting
This is a clipping product that happens to transcribe well. A transcription service reads the recording and stops. Here the read is the first half of a longer job: the same words are scored for a hook, for flow, for value and for travel, and what comes back beside the transcript is a ranked list of passages with timecodes and one line of reasoning each.
You are not billed twice for the same reading. One credit is one minute of source analysed, and that minute buys the transcript and the candidate list together. If you were going to pay a transcription service and then decide by hand which bits to cut, the second half of that job is already done and already paid for.
Nothing is cut until you say so, and nothing is posted at all. The candidates are a proposal. Rendering is a separate decision, and the finished MP4 waits in your library — there is no integration with YouTube, TikTok, Instagram, Twitch or Facebook for posting, and none planned. Reading a video from a platform and publishing one back to it are different connections, and only the first exists here.
How the passages are chosenThe reading is metered. The file is not.
Exports of the words cost nothing on any plan, including the trial.
Getting the text means analysing the video, and that is the charge. Analysing a minute of source costs one credit. A thirty-minute talk is thirty credits, an hour is sixty, and the transcript, the corrections and every download that follows are included in that one read rather than metered separately.
The free trial is thirty credits, once, with no card. Thirty minutes of video — one half-hour recording, read in full, with the transcript downloadable in all three formats. Clips exported on the trial carry a small mark; the transcript files do not carry anything, because there is nothing to mark on a text file.
Above it the plans are 4, 11 and 27 hours of reading a month. Starter is 240 credits at $19, Creator 660 at $49, Studio 1,620 at $99, and a year costs eight months. Exporting a clip past the batch each plan includes costs three credits; downloading the words never does.
Questions people ask first
What formats does the transcript download in?
Three. SRT, which is what every upload form and every editor takes; WebVTT, for a player on your own site; and plain text, laid out by speaker with a timestamp on each turn, for reading or pasting. All three come from the same transcript, and none of them costs a credit on any plan.
Do the subtitles line up with a clip I exported?
Yes, and that is deliberate rather than incidental. A clip is a window with holes in it — filler you cut, or a multi-piece assembly — so the file for a clip is mapped onto what actually plays: the removed seconds are gone and everything after them moves up. Exporting the whole recording instead gives you the words on the original clock.
Can I fix a word it got wrong before I export?
Yes. Select the stretch, type what was actually said, and the replacement occupies exactly the seconds the old words did — so no clip boundary and no timecode anywhere else moves. Nothing is re-transcribed and nothing is charged. A clip already rendered keeps the old words burnt into it; the fix lands on the next render and on every export after the correction.
Will it translate the video into another language?
No. There is no translation and no dubbing in this product at any tier. The language is detected from the audio and the transcript comes back in the language spoken, which is also why the captions on a clip are always the speaker’s own words.
Can I upload an mp3 or a voice memo?
The product is built around video and that is what it is tested on. The upload box will take an audio file, but the path has not been proven end to end and nothing here is going to promise it works — if the recording is audio only, the honest answer today is that this is not the tool for it yet.
How accurate is it, and in which languages?
It reads over thirty languages and works out which one it is hearing rather than asking you to declare it. Accuracy is whatever the audio allows: a close microphone in a quiet room transcribes close to cleanly, a phone in a busy street does not, and no transcription tool of any price changes that. What this one adds is that anything it gets wrong can be corrected in place and re-exported for nothing.
Is there a limit on how long the video can be?
No paid plan caps the length of a single video. The limit is the credit balance: a minute read costs one credit, so a two-hour interview spends 120 of them. On the free trial the thirty credits are the bound, which is one half-hour recording.
The rest of what happens to those words
The transcript is the first half of the job. These are the pages about the second half.

AI video clipping tool
The category term, answered with the difference: candidates and reasons first, cuts only after you approve.

Podcast clip maker
Two hours read in thirty-minute windows with three minutes of overlap, so the back half is ranked too.

YouTube clip maker
Paste the watch URL and it is fetched once on our side, or upload the file. Only the passages you approve are ever cut.

Clipping software
What the whole pipeline does, end to end: transcript, ranked candidates, your decision, then the cut.
