A subtitle editor for the words, not for the styling

Fix what was misheard, and the correction lands in the SRT, the VTT and the captions burnt into the next render. What it will not do is on this page too, in the first section.

The boundary, before anything else

This edits what the subtitles say. It does not edit how they look.

What you can change: the words. Select a stretch, type what was actually said, and it is replaced. A name the transcription heard wrong, a technical term it guessed at, a brand it spelled phonetically — those are the corrections this exists for, and they are the ones that matter, because a misheard name is burnt into every clip that contains it.

What you cannot change: the type, the position, the colours, the timing by hand. There is no timeline to drag a cue along, no font picker and no per-cue styling. Two caption treatments ship — words drawn one at a time, or text set behind the speaker — and that is the whole of the choice. If you came looking for a place to restyle a subtitle track, this is the wrong tool and it is better to say so now than after an account exists.

The timings look after themselves, and that is deliberate. A replacement occupies exactly the seconds the old words did, subdivided across the new words in proportion to their length. Nothing outside the edit moves by a millisecond — no clip boundary, no other cue, no moment elsewhere in the recording. That is what makes correcting a word safe enough to do without checking the rest.

How a correction behaves

Four things that happen when you fix a word

Worth knowing before you fix forty of them.

It costs nothing

No model is called and nothing is re-transcribed — the bytes on disk change and the next read of them is different. No credits are spent on a correction, ever.

Every export afterwards carries it

The SRT, the WebVTT, the plain text and the captions drawn onto the next render all come from the same words, so fixing it once fixes it everywhere it has not already happened.

A clip already rendered keeps the old words

That file exists with the mistake burnt into it. The fix lands on the next render, which is the honest behaviour and the only affordable one — re-rendering everything on a typo would be a bill nobody agreed to.

The paragraphs rebuild themselves

The blocks a person reads are recomputed from the corrected words rather than edited alongside them. Two representations of the same sentence maintained separately is two representations that eventually disagree.

The file at the end

Cues cut for reading, and a clip file that matches the clip

The export does not simply write the transcription engine’s utterances out. Those are speaking turns — routinely thirty seconds and four hundred characters, which is a fine paragraph and an unusable subtitle. The words are packed instead into cues of about forty characters over two lines, seven seconds at most, broken where sentences close, where the voice changes and where somebody genuinely stops talking.

And a clip gets a file cut to what actually plays. A rendered clip is a window with holes in it — filler removed, or several pieces assembled — so its subtitle file is mapped onto the clip’s own clock: the cut seconds are gone and everything after them moves up. A file that merely subtracts the start time is correct until the first cut and drifts for the rest of the video.

Three formats, and none of them is metered. SRT for anything with an upload form, WebVTT for a player on your own site, plain text for reading. Downloading them costs no credits on any plan, including the free trial, because the file is arithmetic over a transcript that was already paid for once.

What this site is

The subtitles are a side effect of what this product is really for

It reads a long recording and finds the passages worth cutting. Upload a file or paste a YouTube address, the whole thing is transcribed, and every passage worth posting comes back as a ranked proposal — four scores, its timecodes and the sentences behind it. Nothing is rendered until you approve it, and nothing is ever published on your behalf.

The transcript exists because the selector needs it, and you get it too. Which is why correcting a word is free and why the export is not metered: both are second uses of a read that has already happened. If all you want is a subtitle file for a video you already have, that works — but you are using the smaller half of the product.

Questions people ask first

Can I change the font, the colour or where the subtitles sit?

No. Two caption treatments ship — words drawn one at a time, or text set behind the speaker — and there is no font picker, no colour control and no positioning. If restyling a subtitle track is what you need, this is the wrong tool, and that is worth knowing before rather than after.

Can I retime a cue by hand?

Not directly, and the reason is that the timings come from the audio rather than from a timeline. A correction occupies exactly the span the old words did, so what you can change is what is said inside a stretch of seconds, not which seconds it occupies.

Does correcting a word cost anything?

No credits at all. Nothing is re-transcribed and no model is called — the stored words change, and everything read from them afterwards is different. It is the cheapest operation in the product.

Will it fix a clip I already exported?

No. That MP4 exists with the old words burnt into it, and the fix lands on the next render instead. Re-rendering everything whenever a word changed would be a cost nobody agreed to, so the honest behaviour is the one that leaves the old file alone.

Can I upload an SRT I already have and edit that?

No. The words here come from reading a video, so the recording is the way in — a file you upload or a YouTube address you paste. There is no path that takes an existing subtitle file as input.

What formats can I download?

SRT, WebVTT and plain text, from the whole recording or from any clip you have rendered. The clip version is cut to what actually plays, so it lines up with the MP4 beside it.

Fix it once, and every export after it is right.