Two hours in.
Six worth posting.The transcript finds them. Watching decides.
A talking format is the one case where the transcript is genuinely strong. Which is why finding candidates is the easy half. The hard half is knowing which of them survives on its own, out of context, in a feed - and no wall of text can tell you that. Cutlist reviews every candidate with the frames and the audio in hand before it ships one.
The transcript is the shortlist
Words tell you where the ideas are. They do not tell you whether the idea was delivered. A tool that only reads the transcript hands you the shortlist and calls it the answer. Cutlist treats it as the shortlist, then goes and watches.
- motion
- audio
- scene cut
- nominated
Almost no motion, and that is fine — here the work is in the words and in the two places the room actually reacts. The frames still matter for the other half: they are what catches the lower third, the slide, and the fact that the second-best moment on the transcript is a sponsor read.
Illustrative shape · your own curves are measured, not drawn
What the words get right
Topic, structure, the moment somebody makes a claim worth arguing with. On a scripted interview this is most of the way there, and Talk mode leans on it deliberately.
What the words cannot see
A four-second pause that kills the pacing. A guest reading off a laptop. The host talking over the punchline. A joke that got nothing back. All of it invisible in text, all of it fatal in a clip.
The review step
Every shortlisted moment is sampled as frames and judged against the audio texture at the same time - not a face detector, a model looking at the actual pictures alongside what the room sounded like.
It shows its working
Each clip scores 0 to 99 and reports what it saw, what it heard and why it landed there. You can check the number instead of trusting it.
The opening line, stated up front
Every result reports the line it opens on. If that sentence does not stand up cold, the clip will not either - and you can see that before you watch a single one.
The score is a judgement, not a forecast. It rates whether a moment stands alone, not whether an audience will share it - nobody can predict that from a file.
Talk mode, and a brief
Set the footage type to Talk and selection leads with the transcript, which is the right order for a conversation. The review still watches. Everything else in the panel is optional; drop-and-go still works.
Find these moments
Plain words. "The parts where they disagree", "practical advice only", "anything about the funding round". Selection narrows to what you describe.
Clip length
Under 30 for a single line, 30-60 for one idea fully landed, 60-90 when the set-up earns its keep. Or Auto, and each moment takes what it needs.
Only this part
Skip the cold open and the ad read. Give it 4:00 to 1:58:00 of a two-hour episode and only that stretch is processed. The trimmed ends still have to be downloaded before anything can seek into the file, so they carry a small handling charge rather than nothing at all - the estimate itemises it.
The estimate, before the button
A live credit figure computed by the same code that bills you. The number you see is the number you pay.
Upload the finished edit rather than the raw record. Working from the raw file, it will surface moments you already decided to cut.
Word-timed, six ways
A feed autoplays muted, and a talking-head clip with no captions is a person moving their mouth. Timings come from the transcript at word level, so the highlight lands on the word being spoken instead of drifting a beat behind it. Burned into the video, with an SRT alongside for platforms that want their own.
Bold
Three words a line, the spoken one in your accent colour. The default.
Punch
One very large word at a time. For a line that lands on one word.
Clean
Sentence case, more words per line. For a considered answer, not a zinger.
Block
A solid card behind the text, so it holds over a bright studio wall.
Brand
Every word in your accent colour, for cuts that should read as yours.
Subtle
Small and top of frame, so it never fights a lower third in the source.
Size and height
0.6x to 1.6x, and vertical position anywhere in the frame. Previewed before anything renders, so a caption never lands on a chin.
Strike a word, the footage goes
Open the Text tab and the clip is a transcript. Click a word to cut it and the video goes with it - the silence around it included - and the remainder is rejoined seamlessly. Losing an um, a false start or a whole rambling sentence takes about as long as reading it.
Trim by hand too
Drag the handles, nudge in tenths of a second, or use the arrow keys. The preview scrubs with you.
Move the opening line
If a clip opens mid-thought, pull the in-point to the sentence that works cold. The clip is usually fine. The first four seconds were not.
Restyle without re-running
Caption style, colour, size, height, aspect ratio and reframing mode all change on a finished clip. Detection never runs again.
A re-cut is 18 credits
The moment is already known, so none of the model half of the start-up charge applies: no detection, no sweep, no review, no ranking pass. What is left is the minute being processed and the episode being fetched again, because a fresh worker has no copy of it. Re-cutting a minute-long clip out of a two-hour episode costs 18 credits against the 348 the episode itself cost. Experimenting is cheap on purpose.
Face-centred, and honest about it
Face mode samples frames across the clip, finds the largest face in each one and crops there. When there is no face to lock onto - a hand gesture, a whiteboard, a wide of the room - it falls back to tracking whatever moves.
Face
The default for talking heads. It centres the speaker rather than the middle of the frame, so a guest sitting camera-left is not half cropped out of a vertical.
Fit
The whole picture on a blurred bed. The right answer for a two-shot where both people matter and cropping either one is worse than the letterbox.
Manual
Place the crop yourself. For a locked-off camera where you already know which third of the frame you want.
9:16, 1:1, 16:9
Each rendered at full resolution rather than upscaled from a smaller cut. Change the shape on a finished clip and it re-cuts from the original file - no second upload.
The crop is one position for the whole clip, not an operator following a cut. If a clip crosses a camera change - host wide, then guest close - Face picks the middle ground and suits one angle better than the other. Use Fit there, or trim to a single angle. And there is no multi-speaker split screen: it reframes the picture you give it, it does not stack two people into one vertical.
A post caption, in a voice you pick
Every clip still needs something written under it. Cutlist writes that too, describing what is genuinely in the clip rather than inventing a hook for it. Pick the voice once in the brand kit and every new project inherits it.
Relaxed
Conversational and understated, like telling a friend about it.
Punchy
Short and declarative. Often one line, front-loading the surprise.
Fun
Playful and quick, with an emoji or two where they earn their place.
Corporate
Measured and credible. No slang, no emoji, no exclamation marks.
Analytical
Leads with the number or the specific claim, then why it matters.
Storyteller
Sets the scene in a sentence, then lands the moment.
Each one comes with alt text describing the visuals for people who cannot see them, and a few specific hashtags - never #viral, never #fyp. Copy it from the clip card or find it as a text file next to the video in the zip.
It will not invent a statistic, a name or a claim that is not in the footage, which makes it less breathless than a copywriter would be.
What a two-hour episode costs
One credit is one minute of source, rounded up, plus a start-up charge for the work that happens once however long the episode is. You are billed for the footage the pipeline reads, not for the files that come out: nothing you do to a finished clip costs anything, and asking for fifteen clips instead of five moves only the review half of the start-up, never the episode.
- 120 source minutes
- see + listen review (x1.8 a minute)
- start-up, once a job (+132)
Full detection, plus the see and listen review on every candidate. Captions and post captions included, and nothing we render carries a watermark.
- 120 source minutes
- anchored, detection skipped (x0.7 a minute)
- captions need a transcript (x1.4)
- start-up, once a job (+1)
Paste your show notes and detection is skipped entirely. Nobody should pay a model to work out something they already wrote down. Captions still need a transcript, and that part is billed when you ask for it.
A weekly two-hour show at the settings above is four full passes a month, so it wants Studio: $249 a month for 2,400 credits, which is 6 full episodes with 312 credits left over for re-cuts. Creator at $79 holds 2 of them, which is the right plan for a fortnightly show or for anyone using chapter markers - pasting your own puts the same episode at 119 credits instead of 348, and four of those fit inside Creator comfortably.
Three shapes of the same job
Podcasters
One show, one voice, one look. Set the brand kit once and every episode inherits the logo, the colour, the caption style and the caption voice.
Agencies cutting client shows
A different look per client. A brand kit belongs to the account, so the honest shape is one account per client - or set the style per project at upload and override on any clip.
Interview YouTubers
One long-form upload, then the vertical cuts. The same moment can go out 9:16 and 16:9 - switch the shape on a finished clip and it re-cuts from the source you already uploaded.
What it will not do
Multi-speaker split screen
One frame, one crop. Two speakers side by side in a vertical clip is not built. Fit keeps the whole two-shot instead.
Intro and outro cards
The brand kit covers logo, colour, caption style and voice. Topping and tailing a clip is not built.
Publishing for you
Download and post yourself. Nothing goes to a social account on your behalf, and there is no scheduler.
Generated b-roll
Deliberate, not missing. Your clip is your footage - no stock inserts, no synthetic filler.
Clips live 30 days and the source upload is removed within 2 days of a finished render. This is working output, not an archive - download anything you want to keep.