Read how it decides

Written from the engineering: what the score means, how a window is cut, and where both of them stop.

Sentences, not seconds. The selector picks a range of whole sentences from the transcript. That is how the window is defined, so a clip cannot open mid-word or end on a dangling clause — and there is no snapping step, because there is nothing to snap.

The intro is removed first. Detected and excluded before anything is ranked, so the ranking is not spent on the part where you say your own name.

Windows in parallel. Roughly thirty minutes at a time, with three minutes of overlap so nothing worth keeping falls between two windows. A single pass over a two-hour transcript anchors on what it reads first and never picks anything from the back half.

One judge at the end, told nothing. The finalists from every window go into one pool, and a final pass ranks them against each other rather than against their neighbours — without being told which window a clip came from or what it scored on the way in. A judge that can see the earlier number mostly agrees with it, which is a way of ranking nothing twice.

The cold read. Every finalist is then read once more as if by somebody who never saw the video. One that a stranger could not follow on its own is capped at 60 — COLD_FAIL_CAP in packages/engine/src/find.js — and the list says what they would be missing, rather than quietly ranking it lower for a reason nobody can see. It is the step that separates this from a ranker.

See the moment list
Field guide

What the four pillars ask

Hook. Whether the first three seconds earn the next three. It is the only pillar that can be judged from the opening line alone, which is why it is first.

Flow. Whether the passage holds together end to end with no edit. A moment that needs a jump cut to make sense is not one moment.

Value. Whether someone is better off for having watched it. This is the pillar that separates a good line from a useful one.

Travel. Whether the shape matches clips that already went viral. The rubric was distilled from real shorts that beat their own channel’s median views, so it is a prior about what travelled — not a prediction that this one will.

Read them separately. A 90 that is all hook and no value is a different clip from a 90 spread evenly, and only one of them survives being watched twice.

See a score open up

Boxed 1:1, text behind. A square of the speaker on a full-height canvas. The words are composited behind them: the subject is matte-cut and sits over the text, so the line reads and the face is never covered.

Full-frame 9:16, karaoke. Edge to edge vertical, with word-by-word highlighting timed off the transcript.

The frame follows the voice. Speaker tracking is TalkNet-ASD, so on a two-person interview the crop stays on whoever is actually talking rather than whoever happens to be in shot.

Banner guard. It measures the footage for a baked-in lower third or chyron and punches in only if it finds one. Left on always, it would cost you framing on every clip that did not need it.

Both styles are on every plan

Documentation

Shipped with your workspace, written for whoever has to sign this off.

Architecture

The app, the queue and the engine, and which of them runs where.

Selection

Windowing, overlap, the finalist pool, and the final judge.

Transcription

The audio-only stream, word timings, and how sentences are formed.

Rendering

Speaker tracking, the matte cut, captions, and the banner guard.

Keys and data

Which provider sees what, how long it is kept, and what is never sent at all.

Presets

The render options, and how a house style is pinned to a name.

Changelog

July 2026
Selection runs over roughly thirty-minute windows in parallel with three minutes of overlap; a final judge ranks the pooled finalists.
July 2026
The intro is detected and excluded by default, before anything is ranked.
July 2026
Banner guard measures the footage and only punches in when it finds a baked-in lower third.
July 2026
Speaker tracking moved to TalkNet-ASD: the frame follows whoever is talking.
July 2026
Only the seconds of an approved moment are downloaded; transcription reads an audio-only stream.

Bring a video you think it will read wrong