CUTLIST
N°00Three senses

It watches the
footage.
Everyone else just reads it.

Clipping tools are transcript engines with a cropper bolted on. They score a wall of text, pick timestamps, and never once look at the video they just cut. That works until the best moment is silent, or the words are great and the picture is a slate.

The passfive stages, one file
  1. 05%

    Read the file

    Probe: dimensions, rotation, frame rate, whether there is audio at all.

  2. 30%

    Transcribe

    Word-level timings. This is what lets a cut land on a sentence rather than mid-breath.

  3. 45%

    Watch and listen

    Scene cuts, motion, audience reaction, spectral texture — swept across the whole span.

  4. 55%

    Pick moments

    A shortlist three times the size of your order, each window looked at frame by frame.

  5. 95%

    Render

    Reframe, caption, brand, and cut. One pass, no re-encode of anything untouched.

The percentages are the ones your job reports while it runs

01

See

Frames from inside the moment

Every candidate clip is sampled and looked at before it ships. The model reads what is physically happening, catches on-screen text a transcript never contains - a scoreboard, a lower third, a slide title - and throws out windows that turn out to be a commercial, a slate, a dissolve, or ten seconds of black.

A transcript-only tool cannot tell you the clip it just picked is a mid-roll ad.

02

Listen

The texture of the sound, not just its volume

Loudness alone cannot separate a crowd erupting from a music sting from someone shouting into a mic. Cutlist measures spectral centre, flux and flatness alongside level, so a broadband roar reads differently from a sustained tone and a sharp transient reads differently from both.

This is what finds the moment when nobody is narrating it.

03

Read

Word-level transcript with real timings

Speech is transcribed with per-word timestamps, which is what makes clip edges land on sentence boundaries instead of mid-thought, and what drives captions that highlight the word actually being said.

Everyone does this part. It is the floor, not the product.

01Why three beats one

One score,three witnesses.

The senses are not three separate features on a pricing page. They are combined into a single judgment on every candidate: the transcript proposes, the audio corroborates, and the frames decide. A moment the crowd reacted to outranks one it did not. A moment whose frames turn out to be an ad does not ship at all.

The rubricweights, not vibes
hook
0.34
Does the first frame stop a thumb?
standalone
0.22
Does it make sense with nothing before it?
payoff
0.20
Does it land, or end on a setup?
value
0.14
Is there something in it worth the watch?
flow
0.10
Does it hold together start to end?

The number on a clip is 70% these five axes, weighted as above, and 30% the model’s holistic read — which catches what no axis names. It is a judgment about the footage, answerable from the footage. It is not a view count and does not pretend to be one.

What a scored clip carries
{
  "rank": 1, "score": 94,
  "title": "He took it eighty yards",
  "seen": "Receiver breaks two tackles at midfield and runs it in
           untouched; sideline empties onto the field.",
  "heard": "loud and broadband, consistent with crowd noise,
           applause or cheering",
  "onScreenText": "4TH QTR  2:11  21-17",
  "reason": "Self-contained scoring play with a visible lead change
           on the scoreboard and a crowd reaction to match."
}

Every clip explains itself. You can filter on the score without watching all of them, and when one is wrong you can see which sense misread it.

02Why the extra senses matter

What changeswhen it can see.

CapabilityTranscript aloneCutlist
Finds moments in footage with no dialogueNeeds a separate productThe frames are the signal
Knows a clip is a commercial before shipping itNoRejected at review
Reads on-screen text (scores, lower thirds, slides)NoYes, and it informs the title
Tells you why a clip scored what it didA numberWhat it saw, heard and read
Takes timestamps you already haveNoPaste them at upload, anchored rate

Based on published capabilities of transcript-based clippers as of July 2026.

03What we will not claim

The honestlimits.

Every clipping tool on the market says it finds viral moments. None of them can, including this one - virality is a property of an audience, not of a file.

The score is a judgment, not a forecast

It rates hook, flow, value, standalone and payoff against what the model can see, hear and read. Every one of those is answerable from the footage itself - it does not know your audience and cannot predict views.

The audio layer is DSP, not a classifier

It describes acoustics honestly - “broadband and loud, consistent with crowd noise” - rather than asserting an event it cannot actually identify.

You will still discard some clips

Fewer than with a transcript-only tool, because the obvious failures are caught before render. Not zero. Anyone promising zero is selling.

Rights are yours to hold

Cutlist will clip whatever you point it at. Whether you may publish it is a question the tool cannot answer for you.