Read how it decides
Written from the engineering: what the score means, how a window is cut, and where both of them stop.
How a moment is chosen
Sentences, not seconds. The selector picks a range of whole sentences from the transcript. That is how the window is defined, so a clip cannot open mid-word or end on a dangling clause — and there is no snapping step, because there is nothing to snap.
The intro is removed first. Detected and excluded before anything is ranked, so the ranking is not spent on the part where you say your own name.
Windows in parallel. Roughly thirty minutes at a time, with three minutes of overlap so nothing worth keeping falls between two windows. A single pass over a two-hour transcript anchors on what it reads first and never picks anything from the back half.
One judge at the end, told nothing. The finalists from every window go into one pool, and a final pass ranks them against each other rather than against their neighbours — without being told which window a clip came from or what it scored on the way in. A judge that can see the earlier number mostly agrees with it, which is a way of ranking nothing twice.
The cold read. Every finalist is then read once more as if by somebody who never saw the video. One that a stranger could not follow on its own is capped at 60 — COLD_FAIL_CAP in packages/engine/src/find.js — and the list says what they would be missing, rather than quietly ranking it lower for a reason nobody can see. It is the step that separates this from a ranker.
What the four pillars ask
Hook. Whether the first three seconds earn the next three. It is the only pillar that can be judged from the opening line alone, which is why it is first.
Flow. Whether the passage holds together end to end with no edit. A moment that needs a jump cut to make sense is not one moment.
Value. Whether someone is better off for having watched it. This is the pillar that separates a good line from a useful one.
Travel. Whether the shape matches clips that already went viral. The rubric was distilled from real shorts that beat their own channel’s median views, so it is a prior about what travelled — not a prediction that this one will.
Read them separately. A 90 that is all hook and no value is a different clip from a 90 spread evenly, and only one of them survives being watched twice.
See a score open upThe two standing shapes
Boxed 1:1, text behind. A square of the speaker on a full-height canvas. The words are composited behind them: the subject is matte-cut and sits over the text, so the line reads and the face is never covered.
Full-frame 9:16, karaoke. Edge to edge vertical, with word-by-word highlighting timed off the transcript.
The frame follows the voice. Speaker tracking is TalkNet-ASD, so on a two-person interview the crop stays on whoever is actually talking rather than whoever happens to be in shot.
Both styles are on every planDocumentation
Shipped with your workspace, written for whoever has to sign this off.
The app, the queue and the engine, and which of them runs where.
Windowing, overlap, the finalist pool, and the final judge.
The audio-only stream, word timings, and how sentences are formed.
Speaker tracking, the matte cut, captions, and the banner guard.
Which provider sees what, how long it is kept, and what is never sent at all.
The render options, and how a house style is pinned to a name.


