Features.All of them, plainly.
Where something has a limit, the limit is written down next to it. Anything not on this page is not built yet.
Hover anything to see what it does. All of it is included on every plan, including the free trial.
It opens the video
Most clipping tools score a transcript and never look at a single frame. Cutlist judges every candidate on three inputs at once: the pictures inside it, what the room actually sounded like, and the words if there are any.
Every score is the sum of what was seen, heard and read. Click a moment to move the playhead.
- hook
- 0.34
- Does the first frame stop a thumb?
- standalone
- 0.22
- Does it make sense with nothing before it?
- payoff
- 0.20
- Does it land, or end on a setup?
- value
- 0.14
- Is there something in it worth the watch?
- flow
- 0.10
- Does it hold together start to end?
The number on a clip is 70% these five axes, weighted as above, and 30% the model’s holistic read — which catches what no axis names. It is a judgment about the footage, answerable from the footage. It is not a view count and does not pretend to be one.
Works with no dialogue at all
When nobody is talking, sampled frames become the detector. A wide shot of a pitch is not a dead end, it is the input.
Rejects what is not really a clip
Commercials, slates, dissolves and ten seconds of black are thrown out at review, before anything renders.
Reads on-screen text
A scoreboard, a lower third, a slide title. Context a transcript never contains, used both to judge and to title.
Scores 0 to 99, and shows its working
Every clip reports what it saw, what it heard and why it landed where it did. You can check the number rather than trust it.
Prompt it if you want something specific
Ask for "every lead change" or "the parts where they disagree" and selection narrows to that.
The score is a judgement, not a forecast. It rates whether a moment stands alone, not whether an audience will share it - nobody can predict that from a file.
Point it before it runs
Give it a file, a link to one, or the page a video sits on, and a guidance panel appears: what the footage is, what to look for, how long a clip should run, what part of the video to even consider. Every control maps to a real engine knob, every one has a working default, and the shortest path is still drop-and-go.
What is this footage
Talk leads with the transcript; Action leads with the pictures. Auto reads the footage and decides - and tells you which way it went.
Find these moments
Plain words: "every touchdown by number 7", "the parts where they disagree". The model reads, watches and listens for what you describe.
Clip length
Under 30 for hooks and single plays, 30-60 for one idea fully landed, 60-90 when a build-up earns its keep. Or Auto, and each moment takes what it needs.
Spoken language
Auto-detected by default. Override it when a noisy open or a bilingual stream fools the detector.
Only this part
Give it 10:00 to 30:00 of a three-hour broadcast and only that stretch is processed, and very nearly all of what you save is saved. The file itself still has to come down the wire before ffmpeg can seek into it, so the rest of it carries a small handling charge - itemised in the estimate, and a fraction of what reading it would have cost. Skip the pre-show, the halftime, the outro.
The price, before the button
A live credit estimate updates as you set options, computed by the same code that bills you. The number you see is the number you pay.
The crop follows the action
Face-detection reframing has nothing to hold onto in a wide shot, so it centres the frame and hopes. Cutlist tracks whatever is actually moving and pans with it, smoothed so it reads like an operator rather than a jump cut.
The crop tracks whatever is moving, so the subject stays in frame across the shot.
Follow
Motion tracking. Built for sport, stage and anything where the subject crosses the frame.
Face
Centres on the speaker for talking heads, and falls back to motion when there is no face to find.
Fit
The whole picture on a blurred bed. Nothing is ever cropped out, which is the safe choice for busy footage.
Manual
Place the crop yourself with a slider, for a fixed camera in a known corner of the frame.
Vertical, square or wide
9:16, 1:1 and 16:9, each rendered at full resolution rather than upscaled from a smaller cut.
Word-timed, six ways
Timings come from the transcript at word level, so the highlight lands on the word being spoken instead of drifting behind it. Burned into the video, with an SRT alongside for platforms that want their own.
HE TAKES IT
Word-level timings from the transcript, so the highlight lands on the word being said. Burned into the video, with an SRT alongside it.
Bold
Three words a line, the spoken one in your accent colour. The default.
Punch
One very large word at a time. Maximum attention, minimum reading.
Clean
Sentence case, more words per line, for a conversation rather than a highlight.
Block
A solid colour card behind the text, so it holds on bright or busy footage.
Brand
Every word in your accent colour, for cuts that should read as yours.
Subtle
Small and top of frame, so it never fights a lower third already in the source.
Adjustable
Size from 0.6x to 1.6x and vertical position anywhere you want it, previewed before rendering.
Change anything, re-cut cheap
An automatic cut is a first draft. The editor gives you the controls that change the outcome, and re-cutting is billed at the anchored rate because the moment is already known.
Trim
Drag the handles, nudge in tenths of a second, or use the arrow keys. The preview scrubs with you.
Edit by text
Where there is speech, click a word in the transcript to cut it. The footage goes with it, silence included, and the remainder is rejoined seamlessly.
Reshape and reframe
Switch aspect ratio or reframing mode on a finished clip without going back to the source.
Restyle captions
Style, colour, size and height, with a live preview of where the words will sit.
Cut the ums and the dead air
One click strikes the hesitation sounds and closes the long pauses, so a rambling take reads tight. It works by cutting your own footage - the struck words and trimmed silences run through the exact span mechanism the text editor already renders through - so nothing is ever added, smoothed over or generated.
Hesitations only
Um, uh, er, hmm and their spellings, plus an immediate stutter repeat. It leaves "like", "so" and "you know" alone, because those carry meaning as often as not.
Closes the dead air
A silence longer than the gate you set is trimmed back to it, so the pauses between kept words stop dragging without clipping the delivery.
Your footage, not a rewrite
Struck words and trimmed pauses become seams the clip is rejoined across. Picture and audio are cut together - never synthesised, never bridged with filler.
Every cut is visible
It marks what it would remove in the transcript and shows the seconds saved before you commit. Strike more by hand, or put a word back, one click each.
Tighten needs words to work from, so it acts where there is speech. It is deliberately conservative - it would rather leave an "um" than delete a real word - so a very loose take may still want a pass by hand.
Clean up the sound you recorded
Levels drift and rooms hum. Enhance audio evens the loudness and pulls down the background noise on the audio already in your footage, so a clip sounds finished. It cleans what you captured - it never adds a voice, a word or a sound that was not there.
Even loudness
A quiet passage and a loud one are brought to one consistent level, so a viewer is not reaching for the volume between clips.
Less background noise
Steady hiss, hum and room tone are reduced, so speech sits forward instead of sounding underwater.
Your sound, cleaned not faked
It only processes the track you recorded. No synthetic voice-over, no generated ambience, no words put in anyone's mouth.
One switch
On when a clip should sound polished, off when the raw room is the point. Off leaves the audio exactly as it was shot.
Enhancement, not restoration. It improves what was captured; it cannot recover a voice that was never on the mic, and audio that clipped at the source stays clipped.
Set it once, everything inherits
Your logo, your colour, your caption treatment and your caption voice, saved to the account and applied as the defaults on every new project. Any single clip can still override them.
Logo on every clip
Burned in at render, any corner, with size as a percentage of the frame so it looks right on vertical and wide alike. Opacity is yours to set.
Accent colour
Drives the highlighted caption word and the Block and Brand styles.
Default caption style
The look you use most, pre-selected on every upload.
Default post voice
So your captions read consistently without picking a voice each time.
Live preview
Placement, size and opacity are previewed on a mock clip before anything renders.
No intro or outro cards yet, and no per-project kits - a brand kit belongs to the account.
Bring your own moments
If you already know where the good parts are - chapters, a run of show, notes you took while watching - paste them. Detection is skipped entirely, so the start-up charge drops to the delivery and the renders, and the source bills at the anchored rate.
Almost any format
0:45 Title, 12:34 - Title, ranges like 1:02:11-1:02:40, bracketed times, or bare seconds. One per line.
Counted before you spend
The upload box tells you how many moments it recognised as you type, so a mistyped line is visible immediately.
Exactly what you asked for
Supplied moments are never re-ranked, deduplicated or trimmed to a clip limit. What you list is what you get.
A fraction of the credits
No scanning and no detection model call. Captions still need a transcript, and that part is billed when you ask for it.
Take everything, no watermark
Clips are yours the moment they exist. Nothing we render carries a watermark on any plan, including the free trial.
One zip
Every clip and every caption file for a project, named from the clip titles.
Per clip
Download the video or the SRT on its own.
100 GB per account
Shared across all your projects, with usage shown in the studio.
Clips live 30 days
Working output, not an archive. Each clip counts down to its expiry so nothing vanishes unannounced.
Source uploads are removed within two days of a finished render. Download anything you want to keep.
What is not built
Intro and outro cards
The brand kit covers logo, colour, caption style and voice. Topping and tailing a clip is not built.
Team workspaces
One account, one set of projects. No shared seats or roles.
Posting to social
Download and post yourself. We do not publish on your behalf.
Generated b-roll or voices
Deliberate, not missing. Your clip is your footage - no stock inserts, no synthetic voice-over, no invented speaker. What comes out is what you shot.
The words that go with it
A clip still needs something written next to it when you post. Cutlist writes that too, in a voice you pick, describing what is genuinely in the clip rather than inventing a hook for it.
Relaxed
Conversational and understated, like telling a friend about it.
Punchy
Short and declarative. Often one line, front-loading the surprise.
Fun
Playful and quick, with an emoji or two where they earn their place.
Corporate
Measured and credible. No slang, no emoji, no exclamation marks.
Analytical
Leads with the number or the specific claim, then why it matters.
Storyteller
Sets the scene in a sentence, then lands the moment.
Alt text and hashtags
Every caption comes with a plain description of the visuals for people who cannot see them, plus a few specific hashtags - never #viral or #fyp.
Copy or export
Copy from the clip card, or find it as a text file next to the video in the zip.
Captions describe what is in the clip. They will not invent a statistic, a name or a claim that is not in the footage, which means they are less breathless than a copywriter would be.