A Descript alternative for the clips, not for the edit
Descript is where you cut an episode by cutting its transcript. This reads the episode you already cut and says which forty seconds of it are worth posting — with the score, the timecodes and the sentences behind each candidate.
One is a wider tool than the other, and that is the answer
Both start from a transcript and the resemblance ends there. Descript is an editor: everything you can do to a video is on the table, and finding a clip is one job among many. This does a single job — it reads a long recording and hands back the passages that survive being watched on their own.
The deliverable is the episode
You are removing ten minutes of tangent, laying in B-roll, fixing the audio and exporting one finished video. That is an edit, and it wants an editor. Descript.
The episode exists, you want the clips
It is published, or it is about to be, and somewhere in two hours are the four passages worth posting. That is a reading problem before it is an editing one. This.
Most weeks it is both
In that order. Nothing here reads a project file or a timeline — it takes the finished video, so it sits after the edit rather than in place of it.
What differs, written out
No ticks and no dashes: a mark next to somebody else’s product is a claim about software we do not run. Each cell says what each tool does, and you decide which sentence describes your week.
| What differs | InTheClips | Descript |
|---|---|---|
| The unit of work | One long recording in, a ranked list of candidate clips out. There is no project to assemble and nothing to lay out. | A project you assemble. The transcript is the document, and deleting a line deletes the video under it. |
| Who finds the moment | A discovery pass reads the transcript in thirty-minute windows, in parallel, with three minutes of overlap, then one blind judge ranks the pooled finalists on a single scale. | You do, by reading. Seeing everything that was said and choosing from it is precisely what a text editor for video is for. |
| What arrives with a candidate | Four scores — hook, flow, value, travel — the timecodes, and the sentences that made it a candidate, so you can disagree with the reasoning rather than only with the result. | The passage you highlighted, where you highlighted it. |
| Where a cut lands | The model names the words to open and close on and the times are read off those words, so a boundary can fall anywhere except inside a word. | Wherever you put it, to the word, by hand. |
| What you may change afterwards | The in and out points, and a tick list of the dead air and hesitation sounds it proposes to remove. Nothing else — this is not an editor. | Anything. That is the point of it. |
| What comes out | An MP4 in the library: boxed 1:1 with the words composited behind the speaker, or full-frame 9:16 with word-by-word captions. | Whatever you export from the sequence you built, in the shape you set up. |
Their column is from Descript’s own public description of the product, read on ; ours is from the code that runs this one.
What reading a two-hour episode consists of
How it choosesThe check you cannot run on your own edit
A cold reader gets the clip and nothing else. Once a moment has been cut, a second pass receives only the clip’s own lines — no transcript around them, no title, no timecodes — and answers as somebody who has just landed on it. If it cannot follow, the report goes back to the cut and the boundaries are chosen again, up to twice. A clip that still does not stand alone stays in the list with its score capped at 60 and a line saying what is missing.
It is the one reading you are disqualified from. You know who the guest is and what was said in the ten minutes before, so on a timeline every clip you cut makes sense to you. That is why a clip that opened on “well, like you said there” gets caught here and not in the export.
One edit is made for you, and it is small. A moment can be assembled from up to three pieces in chronological order — the setup from a minute earlier, the moment itself, and the detour between them dropped. The holes are added by the API at render time and never taken from a client, so the clip that gets cut is the clip that was checked.
Read a long episode this wayWhere this one stops
An alternative to Descript that quietly implied it could also record, edit and export your episode would be found out on day one. So here is the edge, written down.
There is no timeline
No tracks, no B-roll, no multitrack, no recording, no transitions. Moving the in and out points of a moment is the whole of the manual control, on purpose.
It generates nothing
No synthetic voice, no video, no rewriting of what you said. Every word in a clip is a word you actually spoke, at the time you spoke it.
One video at a time
A back catalogue goes through it one file after another. There is no batch that swallows a hundred episodes overnight.
One pairing is not built
Text behind the speaker is written against the square box, so 9:16 with text behind does not exist. The options panel corrects the pairing rather than letting a job fail twenty minutes in.
Questions
Asked by people who already have Descript open in another tab.
Is this a replacement for Descript?
Not if the deliverable is the episode. There is no timeline here, no multitrack and no recording, so anything that ends with you exporting a finished long-form video still wants an editor. What it replaces is the part of the week spent scrubbing a finished episode looking for the forty seconds worth posting.
Can I fix a moment it got slightly wrong?
You can move the in and out points. Dragging a handle also makes the clip one continuous stretch again: the assembled pieces are dropped, because they were reasoned about a window that no longer exists, and keeping them would remove seconds nobody can explain afterwards.
Does it remove filler words like Descript does?
Dead air and non-lexical hesitation sounds — uh, um, er, hm — are computed from the word timings before any video is touched, itemised with what would go, and removed only for the ones you leave ticked. Words that carry meaning, like so or right, are never on that list: cutting one can break the sentence around it.
What do I give it?
A finished video, and there are two ways to hand it over. Upload the file — mp4, mov, mkv or webm — or paste a YouTube address, which is read on our side without a copy of the video being kept — the transcription takes an audio-only stream, and each render afterwards fetches only the seconds of its own window, so a long recording is never downloaded in order to be uploaded again. Links are YouTube only and counted by the day — three on the trial, a hundred on Studio — while uploads are not counted at all. No paid plan caps how long a single video may be; the free trial is bounded by its thirty credits, which is thirty minutes of video.
How is it billed, if the clips are unlimited?
By minutes of source video analysed. Transcription and the selection pass both cost in proportion to how long a video is; rendering does not, because it only ever touches the seconds you approved. So the hours are the meter and the clips that come out of them are not counted.
What does the finished clip look like?
Boxed 1:1 with the text composited behind the speaker, who is matte-cut and sits over the words, or full-frame 9:16 with word-by-word captions timed off the transcript. In both, the crop follows whoever is actually talking, and a banner guard punches in only when it measures a lower third baked into the footage.
The other tools people arrive from
Same structure, different starting point — and, for anyone weighing the yearly price rather than the tool, what a discount here is actually worth.

