The aspect ratio is decided by watching the clip, not by a crop box
Five shapes come out of this, and which one a given clip gets is chosen by a model looking at that footage — a composited camera, how many faces share a frame, whether the moment is carried by words or by picture. Nobody drags a box, and there is no box to drag until after it has already guessed.
A short gets cut and framed — nothing is relabelled whole
This is not a general video aspect ratio converter. Paste a 16:9 recording and you will not get the same footage back, unchanged, in 1:1. What comes out is a short cut from inside it, framed for whichever of five shapes fits what the camera actually shows during that moment. If what you want is one whole file relabelled into a new ratio with nothing cut, that is a different job than the one this does — this tool finds the passage worth publishing first, and the shape decision belongs to that passage, not to the source file.
And every shape delivers the same file dimensions. All five formats render to the same 1080×1920 vertical file — Boxed 1:1 included, which sits a square of the speaker inside that full-height canvas with room for text above and below it, rather than handing back a literally square video. What changes between the five is what fills the frame, never the width and height of the file it arrives in.
The whole pipeline, end to endA vision pass over the clip, never the transcript
A podcast and a gameplay stream can share the exact same kind of transcript — one person talking — and still need opposite frames. What tells them apart is visible in a single glance and invisible in the words, so the picture is what gets asked.
Facts, plus one judgement
The model reports what it can check by looking — is a separate source composited over the picture, what is inside it, how many real faces ever share a single frame — and answers exactly one opinion: is this clip carried by what is said, or by what is shown.
A camera over a game or a screen
When a separate source is composited in and it is a face, the shape becomes Cam bubble or Face over screen. Offered only where that camera was actually found — never guessed at because a preset asked for one.
Nobody to frame
A screen recording, a slide, gameplay with no webcam over it — there is no subject to track, so the source is dropped into the vertical frame and otherwise left alone.
One head, or a room full
A single speaker explaining something loses nothing shrunk to a square, so it is boxed and the freed space goes to text. Two people, or a moment carried by the picture rather than the words, go full-frame instead — a square cannot hold a two-shot.
The five shapes, and the two a preset may keep
Every clip can be pushed into any of the five by hand. Only two of them are a standing preference across every video you cut — the other three are facts about one clip.
Boxed 1:1
A square of the speaker on the full-height canvas, with the space above and below given to text.
Full 9:16
Edge to edge, no square, no bands. The frame follows whoever is talking across the whole vertical file.
Cam bubble
The screen fills the frame and the camera rides over it as a circle, the way it looked live.
Face over screen
The camera sits across the top, the screen filling the rest — for when the talking matters as much as what is on screen.
The whole picture
Dropped into the vertical frame and otherwise left alone, its own camera wherever it already was.
A saved preset can only carry Boxed 1:1 or Full 9:16. Cam bubble and Face over screen are answers about one clip’s footage, not a preference — save one of them and every future video with no camera in the picture would quietly fall back to something nobody chose.
Two caption treatments, and one pairing that is not built
Text behind the speaker is Boxed 1:1 only. The subject is matte-cut and composited over the words, which only means something on a frame built to hold one subject behind them. On a full-frame shape the compositing geometry for it simply is not written yet — not refused, unbuilt — so choosing a full-frame shape sets the captions to karaoke instead, and the options panel corrects the pairing rather than letting a render fail after it has already started.
Karaoke runs on all five shapes. Word-by-word highlighting, timed off the transcript, is the one caption style every shape can carry — font, size, colour, outline and shadow are all yours to change, and a look you build can be saved and reused by name.
Where each treatment is built and where it is notThe shape decides what the canvas is made of. Something else decides where to point it
Choosing between the five shapes is one question. Where the crop looks, second by second, inside whichever shape you land on, is a separate mechanism reading the same footage for a different answer.
Speaker tracking, not a fixed crop
TalkNet-ASD reads who is actually speaking and the frame follows them — on both Full 9:16 and Boxed 1:1, the two shapes that ship by default.
A two-shot, held rather than cut down
When two people are far enough apart in frame to need it, the crop stacks them for exactly that stretch instead of picking one and losing the other.
Two questions, two passes
The vision pass that picks the shape and the tracker that points the crop inside it read the same clip for different things — what the frame is made of, and where inside it to look.
How to change a clip’s aspect ratio here, step by step
You are needed for one decision, and it comes after the model has already made its own.
Give it the video
Paste a YouTube watch, Short, live or embed address, or upload the file — mp4, mov, mkv or webm, up to 8 GB. The link path asks you to confirm the video is yours to use; an upload asks for nothing.
It reads before it looks
The whole recording is transcribed and ranked in windows first. Only the passages that make the shortlist get a vision pass over their own footage — reading pixels for a moment nobody keeps would be wasted.
The shape is proposed, not fixed
What the vision pass found pre-fills the format panel. Pick any of the other four instead, or a different caption treatment, and the panel corrects any pairing the renderer cannot build rather than queuing a job that fails.
It cuts what you approved
Only the seconds inside the approved window are read out of the stored video, at the shape you kept. The MP4 waits in your library — nothing is posted anywhere on your behalf.
Every shape is in the free trial, not held back for later
Thirty minutes of video, once, with no card — all five shapes and both caption treatments included, and as many clips out of it as you keep. Trial clips carry a small mark in the corner and are kept for seven days; no paid plan marks the clips it exports.
Questions people ask before pasting a video
Can this convert my whole video from 16:9 to 1:1, unedited?
No — that is a different job. This finds a short worth publishing inside a longer recording, and frames that short in one of five shapes. There is no path that takes a whole file and hands the same footage back relabelled into a new ratio.
How does it decide which of the five shapes to use?
A vision pass looks at the candidate clip itself — whether a separate camera is composited over a game or a screen, how many real faces ever share one frame, whether the moment is carried by what is said or by what is shown — and a fixed table turns those facts into a shape. It never reads the transcript to make this call, because a podcast and a gameplay stream can produce the same kind of transcript and need opposite frames.
Is Boxed 1:1 actually a square video file?
No. Every shape, Boxed included, renders to the same 1080×1920 vertical file. Boxed sits a square crop of the speaker inside that full-height canvas with room for text above and below it — the five shapes differ in what fills the frame, not in the file’s dimensions.
Can I pick a different shape than the one it chose?
Yes. The proposed shape pre-fills the options panel and you can choose any of the other four before rendering. Cam bubble and Face over screen only appear where a camera was actually found in the footage — a saved preset can only hold Boxed 1:1 or Full 9:16, since the other three are facts about one clip rather than a standing preference.
Can I put text behind the speaker on a vertical 9:16 clip?
Not yet. That composite is built against the Boxed 1:1 geometry, and lifting it to a full-frame crop is not done. Choosing a full-frame shape sets the captions to karaoke instead, and the panel corrects the pairing rather than letting a render fail once it has started.
Does picking a different shape cost anything extra?
No shape costs more than another — a render is a render, whatever it comes out framed as. What is metered is minutes of source analysed and the export itself: free while your plan’s included rendered clips last, three credits after. Re-rendering the same moment in a different shape is billed exactly like exporting any other clip, not as a special charge for changing your mind.
Where the shape matters most
The same shape decision, from the angle a different search arrives at.

YouTube Shorts maker
What a Short needs that a clip does not: a survivable opening, a vertical frame, captions sized for a thumb.

Clipping software
What the whole pipeline does, end to end: transcript, ranked candidates, your decision, then the cut.

TikTok clip maker
Vertical 9:16 with word-by-word captions. It hands you the MP4 — the upload to TikTok stays yours.