CUTLIST
N°00Two mechanisms

Some footage
never says it.
Text alone has nothing to score.

This page compares two ways of finding a clip, not two companies. One reads a transcript. The other opens the video, listens to it, and reads the transcript as well. The difference is not effort or taste - it is what the input can hold.

01You give up nothing

Everything the light tool has.It is all here too.

Before the part only one of these can do, the part where they are the same. This is the editor checklist a buyer runs first - text editing, a timeline and shortcuts, filler and pause removal, audio cleanup, reframe, captions, a brand kit, a clean export. Switching does not cost you a single one of them.

Edit by text

Click a word in the transcript to cut it; the footage and its silence go with it, rejoined seamlessly.

Timeline and keyboard

Drag the handles, nudge in tenths with the arrow keys, and the preview scrubs with you.

Filler and pause removal

Tighten strikes the ums and closes the dead air in one click, by cutting your own footage.

Audio cleanup

Even the loudness and pull down the background noise on the audio you recorded.

Reframe to any shape

Vertical, square or wide, the crop following the action rather than centring and hoping.

Word-timed captions

Six styles in your accent colour, burned in with an SRT sidecar alongside.

Brand kit

Logo, colour, caption style and post voice, set once and inherited by every clip.

Export without a watermark

Every clip and caption file in one zip, clean on every plan including the trial.

None of this is where the two mechanisms differ - it is the price of entry, and it is paid. The difference starts at the next section, with the inputs a transcript can and cannot hold.

02The mechanism

One reads the file.The other opens it.

A transcript-only pipeline runs speech to text, ranks the resulting wall of text for hooks and topic shifts, and returns timestamps. It never decodes a frame, because a frame is not part of the judgement. Cutlist runs that same transcript, then samples frames from inside each candidate window and measures the acoustics across it. Three inputs, one score.

A transcript-only pipeline
  • Audio decoded to words with timings
  • A model ranking that text
  • Nothing else enters the score

Fast, cheap, and complete for footage where everything worth clipping is said out loud.

Watch, listen, read
  • Frames sampled from inside the window
  • Level, spectral centre, flux and flatness
  • The same word-level transcript

Slower per minute, and billed higher, because looking at pictures costs image tokens.

Both are settings here, not two companies. The extra senses are a toggle in the studio, and the section on when the cheap way wins is about the footage where you should switch them off.

03What each input holds

A transcript isa record of speech.

That is the whole of it. Everything below is either in that record or it is not, and no amount of engineering puts a scoreboard into a string of words.

Signal in the footageIn a transcriptIn all three inputs
Words that were spokenPresent, with per-word timings.Present, with per-word timings. Same transcript.
A stretch where nobody speaksEmpty. Nothing to transcribe, so nothing to score.Frames and sound texture carry it.
Text burned into the pictureAbsent. A scoreboard is not audio.Read off sampled frames.
Sound that is not speechDropped before scoring. A crowd erupting leaves no words behind.Level, spectral centre, flux and flatness, measured across the window.
A commercial, a slate, ten seconds of blackReads as ordinary speech, or as a gap between sentences.Visible in the frames, and rejected at review.
Where the subject is in the frameNot a property of audio.Tracked, which is what the crop follows.
Whether the picture agrees with the wordsOnly one of the two is in the file.Both are, so they are allowed to disagree.

Claims about what an input can physically contain, not about anyone’s product.

04Where it decides the outcome

Three filesthat settle it.

When everything worth clipping is said out loud, both mechanisms are reading the same evidence and they will land in much the same place. These are the three shapes of footage where they cannot, because one of them is working from an input that does not contain the answer.

A

Nobody is talking

A game filmed on one camera. No commentary bed, no microphone near the pitch, ninety minutes of ambient noise.

Transcript only

The transcript is close to empty for the whole file. There is no text to rank, so a text-ranking pipeline either returns nothing or spreads clips evenly and calls it a result.

Watch, listen, read

The frames become the detector. A run across the frame, a shot, the touchline reacting - all of it is in the picture, and the picture is an input.

B

The context is drawn on screen

Fourth quarter, 2:11 on the clock, 21-17 on the graphic. Nobody reads any of that out loud, because everyone watching can see it.

Transcript only

Score, clock and period are pixels. They were never spoken, so they are not in the transcript at any point in the file.

Watch, listen, read

On-screen text is read from the sampled frames, used to judge the moment and to title the clip that comes out.

C

The payoff is visual

Someone says two words - "oh no" - and the next four seconds of picture explain exactly why.

Transcript only

Two flat words. They can score low and lose a great clip, or score high for a reason that has nothing to do with what happened. Text alone cannot tell those apart.

Watch, listen, read

The frames say what happened. The score moves because of the event, not because of the phrasing.

Case C runs the other way too. A great line delivered over a title card, a slate or a dissolve scores well on text alone and ships as ten seconds of nothing.

05When the cheap way wins

Sometimes the wordsare the whole file.

If every moment worth clipping is announced out loud, the frames only confirm what you already knew and you paid image tokens for the confirmation. A comparison that concedes nothing is an advertisement. Here is the concession, with the bill attached.

A studio podcast

Two people, one framing, nothing moving. The entire value is in the sentence, and the sentence is in the transcript.

A talk or a webinar

The argument is spoken from end to end. Slides add a little; the words carry almost all of it.

A recorded meeting

Decisions get stated out loud. Very little happens in the picture that was not also said.

Anything already timestamped

Chapters, a run of show, a play log. No detector needs to run at all, by either mechanism.

220credits

Watch + listen on

The default. Frames, sound texture and transcript on every candidate.

96credits

Watch + listen off

No frames sampled, no image tokens. The words carry selection.

60credits

Your own timestamps

No detection runs at all. Only the transcript captions need is billed.

One hour of source, captions on. Quoted by the code that bills the job.

Talk leads with the transcript

Tell it the footage is talk and the words drive selection. Action leads with the pictures. Auto reads the footage and tells you which way it went.

The review is a toggle, not a tax

Watch + listen is a switch in the upload panel, and the credit estimate next to it is computed by the same code that bills you. Turn it off and watch the number fall before you spend anything.

Timestamps beat both detectors

Paste chapters or a play log and no detection runs at all - no sweep, no review, no rank pass, so the start-up charge drops to the delivery and the renders and the source bills at the anchored rate. Captions still need a transcript, and that part is billed when you ask for it.

Neither mechanism predicts what an audience will share. The score rates whether a moment stands alone. Anything beyond that is a forecast, and nobody can make one from a file.

06Straight about the gaps

Everything it does.And what it does not.

A mechanism argument is easy to win on paper, so here is the unflattering half: the full capability list, including the rows where the honest answer is that it is not built.

14shipped today
7signature - the reason it exists
3not built, said plainly
  • Finds moments in any footage

    Frames, audio texture and transcript judged together, so it works with or without dialogue

    Sports, podcasts, streams, talks

    Signature
  • Reframing that follows the action

    Motion tracking for anything that moves, face-centring for talking heads, or no crop at all

    Vertical, square and wide

    Signature
  • Explains every score

    Each clip reports what was seen, heard and read to arrive at its number

    Auditable, not a black box

    Signature
  • Timestamps you already have

    Paste chapters or notes and detection is skipped entirely

    Billed at the anchored rate

    Signature
  • Word-timed captions

    Six styles, adjustable size and position, your accent colour, SRT sidecar

    Burned in or off

    Included
  • Writes the post caption

    Six voices, with alt text and specific hashtags, grounded in what the clip shows

    Corporate through to fun

    Signature
  • Editor

    Trim on a timeline, change shape, reframe, caption treatment, then re-cut

    Included on every plan

    Included
  • Filler and pause removal

    One click strikes the ums and uhs and closes the long silences - by cutting your own footage through the editor's existing span mechanism, so nothing is added or generated

    Tighten, in the editor

    Included
  • Audio cleanup

    Loudness evened and background noise reduced on the audio you recorded, so a clip sounds finished without ever leaving your footage

    Your sound, cleaned not faked

    Included
  • Export

    One clip, or every clip and caption file in a single zip

    Never watermarked

    Included
  • Languages

    Transcription auto-detects the language rather than asking you to pick

    Broad language coverage

    Included
  • Your footage is real, and stays real

    No generated b-roll, no stock inserts, no synthetic voice-over, no invented speaker. Every second that comes out is a second you shot - cut, reframed and captioned, never fabricated

    The line we will not cross

    Signature
  • Brand kit

    Logo on every clip, accent colour, default caption style and post voice

    Set once, inherited

    Included
  • Agents can drive it

    An MCP server with four tools, so a workflow can cut video without a person

    cut, poll, estimate, list

    Signature
And what it does not

Three things this category is expected to have that Cutlist does not do yet. If a capability is not on this page, we do not have it.

  • Intro and outro cards

    Topping and tailing a clip with your own frames

    Not built

    Not built
  • Team workspace

    Shared projects and per-seat roles

    Not built

    Not built
  • Posting to social directly

    Download and post yourself for now

    Not built

    Not built
The line we will not cross

We do not fake your footage.

No generated b-roll. No synthetic voice-over. No invented speaker. Cutlist finds the moment, cuts it, reframes it and captions it - and every second that comes out is a second you shot. A pipeline that fabricates footage can pad any file, which is exactly why we will not. Your footage is real, and it stays real.