CUTLIST
N°00The alternative

A transcript is half.It watches the other half.

Most AI clippers score one input: the words. It is cheap to read and it works right up until nobody is talking - then a silent wide shot, a scoreboard, a crowd going up all register as nothing. Cutlist watches the sampled frames and measures the audio as signals in their own right, so the moment is ranked on what it looks and sounds like, not only on what was said over it.

01What the words leave out

Four things a transcriptcannot contain.

This is not about one tool reading better than another. It is about what the input can physically hold. A transcript of a silent second is an empty string, and there is nothing in an empty string to rank.

A stretch with no speech
Transcribes to an empty string, and an empty string cannot be scored. The moment the room goes quiet is the moment a transcript-only pipeline goes blind.
Text burned into the picture
A scoreboard, a lower third, a slide. It is on screen the whole time and never in the transcript, because it was never spoken.
Where the motion went
A play develops across a wide frame. Words cannot tell you where it moved, which is exactly what a vertical crop has to follow.
The crowd
A stadium erupting is the loudest thing in the file and produces no words at all. It is a signal you can hear and cannot read.
02Three witnesses, one score

Read, watchedand heard.

Every moment is scored on three inputs at once. The frames still matter even when the transcript is full - they are what catches the lower third, the slide, and the fact that the best line on the page is a sponsor read.

The frames

Sampled across the whole span and read by a frontier model with the pictures actually in hand - scene cuts, motion, what is on screen, whether the shot is worth a vertical crop.

The sound

Audio energy and spectral texture across the file, so a reaction, a whistle or a drop registers as the peak it is - not as the silence a transcript records.

The words

Where there is speech it is transcribed to the word, used to pick moments that stand alone and to cut on a sentence rather than mid-breath. The transcript is one witness of three, not the only one.

What it sees in an interviewone pass
  • motion
  • audio
  • scene cut
  • nominated

Almost no motion, and that is fine — here the work is in the words and in the two places the room actually reacts. The frames still matter for the other half: they are what catches the lower third, the slide, and the fact that the second-best moment on the transcript is a sponsor read.

Illustrative shape · your own curves are measured, not drawn

03The honest half

When reading the wordsis the right call.

If it were always wrong, this would be advertising. It is not always wrong.

A conversation shot as a single locked-off frame lives entirely in its transcript. There is nothing for the frames to add, and paying a model to watch a static shot is money spent on a picture that never changes. For that footage, reading the words is enough and it costs less.

So Cutlist lets you turn the watching off. Tell it the footage is talk and it leads with the transcript; the pricing page prices that path, and the comparison names the exact footage where you should. The difference is that it is a choice you make, not the only thing the tool was ever able to do.