It watches the
footage.Everyone else just reads it.
Clipping tools are transcript engines with a cropper bolted on. They score a wall of text, pick timestamps, and never once look at the video they just cut. That works until the best moment is silent, or the words are great and the picture is a slate.
- 05%
Read the file
Probe: dimensions, rotation, frame rate, whether there is audio at all.
- 30%
Transcribe
Word-level timings. This is what lets a cut land on a sentence rather than mid-breath.
- 45%
Watch and listen
Scene cuts, motion, audience reaction, spectral texture — swept across the whole span.
- 55%
Pick moments
A shortlist three times the size of your order, each window looked at frame by frame.
- 95%
Render
Reframe, caption, brand, and cut. One pass, no re-encode of anything untouched.
The percentages are the ones your job reports while it runs
See
Frames from inside the moment
Every candidate clip is sampled and looked at before it ships. The model reads what is physically happening, catches on-screen text a transcript never contains - a scoreboard, a lower third, a slide title - and throws out windows that turn out to be a commercial, a slate, a dissolve, or ten seconds of black.
A transcript-only tool cannot tell you the clip it just picked is a mid-roll ad.
Listen
The texture of the sound, not just its volume
Loudness alone cannot separate a crowd erupting from a music sting from someone shouting into a mic. Cutlist measures spectral centre, flux and flatness alongside level, so a broadband roar reads differently from a sustained tone and a sharp transient reads differently from both.
This is what finds the moment when nobody is narrating it.
Read
Word-level transcript with real timings
Speech is transcribed with per-word timestamps, which is what makes clip edges land on sentence boundaries instead of mid-thought, and what drives captions that highlight the word actually being said.
Everyone does this part. It is the floor, not the product.
One score,three witnesses.
The senses are not three separate features on a pricing page. They are combined into a single judgment on every candidate: the transcript proposes, the audio corroborates, and the frames decide. A moment the crowd reacted to outranks one it did not. A moment whose frames turn out to be an ad does not ship at all.
- hook
- 0.34
- Does the first frame stop a thumb?
- standalone
- 0.22
- Does it make sense with nothing before it?
- payoff
- 0.20
- Does it land, or end on a setup?
- value
- 0.14
- Is there something in it worth the watch?
- flow
- 0.10
- Does it hold together start to end?
The number on a clip is 70% these five axes, weighted as above, and 30% the model’s holistic read — which catches what no axis names. It is a judgment about the footage, answerable from the footage. It is not a view count and does not pretend to be one.
{
"rank": 1, "score": 94,
"title": "He took it eighty yards",
"seen": "Receiver breaks two tackles at midfield and runs it in
untouched; sideline empties onto the field.",
"heard": "loud and broadband, consistent with crowd noise,
applause or cheering",
"onScreenText": "4TH QTR 2:11 21-17",
"reason": "Self-contained scoring play with a visible lead change
on the scoreboard and a crowd reaction to match."
}Every clip explains itself. You can filter on the score without watching all of them, and when one is wrong you can see which sense misread it.
What changeswhen it can see.
| Capability | Transcript alone | Cutlist |
|---|---|---|
| Finds moments in footage with no dialogue | Needs a separate product | The frames are the signal |
| Knows a clip is a commercial before shipping it | No | Rejected at review |
| Reads on-screen text (scores, lower thirds, slides) | No | Yes, and it informs the title |
| Tells you why a clip scored what it did | A number | What it saw, heard and read |
| Takes timestamps you already have | No | Paste them at upload, anchored rate |
Based on published capabilities of transcript-based clippers as of July 2026.
The honestlimits.
Every clipping tool on the market says it finds viral moments. None of them can, including this one - virality is a property of an audience, not of a file.
The score is a judgment, not a forecast
It rates hook, flow, value, standalone and payoff against what the model can see, hear and read. Every one of those is answerable from the footage itself - it does not know your audience and cannot predict views.
The audio layer is DSP, not a classifier
It describes acoustics honestly - “broadband and loud, consistent with crowd noise” - rather than asserting an event it cannot actually identify.
You will still discard some clips
Fewer than with a transcript-only tool, because the obvious failures are caught before render. Not zero. Anyone promising zero is selling.
Rights are yours to hold
Cutlist will clip whatever you point it at. Whether you may publish it is a question the tool cannot answer for you.