Some footage
never says it.Text alone has nothing to score.
This page compares two ways of finding a clip, not two companies. One reads a transcript. The other opens the video, listens to it, and reads the transcript as well. The difference is not effort or taste - it is what the input can hold.
Everything the light tool has.It is all here too.
Before the part only one of these can do, the part where they are the same. This is the editor checklist a buyer runs first - text editing, a timeline and shortcuts, filler and pause removal, audio cleanup, reframe, captions, a brand kit, a clean export. Switching does not cost you a single one of them.
Edit by text
Click a word in the transcript to cut it; the footage and its silence go with it, rejoined seamlessly.
Timeline and keyboard
Drag the handles, nudge in tenths with the arrow keys, and the preview scrubs with you.
Filler and pause removal
Tighten strikes the ums and closes the dead air in one click, by cutting your own footage.
Audio cleanup
Even the loudness and pull down the background noise on the audio you recorded.
Reframe to any shape
Vertical, square or wide, the crop following the action rather than centring and hoping.
Word-timed captions
Six styles in your accent colour, burned in with an SRT sidecar alongside.
Brand kit
Logo, colour, caption style and post voice, set once and inherited by every clip.
Export without a watermark
Every clip and caption file in one zip, clean on every plan including the trial.
None of this is where the two mechanisms differ - it is the price of entry, and it is paid. The difference starts at the next section, with the inputs a transcript can and cannot hold.
One reads the file.The other opens it.
A transcript-only pipeline runs speech to text, ranks the resulting wall of text for hooks and topic shifts, and returns timestamps. It never decodes a frame, because a frame is not part of the judgement. Cutlist runs that same transcript, then samples frames from inside each candidate window and measures the acoustics across it. Three inputs, one score.
- Audio decoded to words with timings
- A model ranking that text
- Nothing else enters the score
Fast, cheap, and complete for footage where everything worth clipping is said out loud.
- Frames sampled from inside the window
- Level, spectral centre, flux and flatness
- The same word-level transcript
Slower per minute, and billed higher, because looking at pictures costs image tokens.
Both are settings here, not two companies. The extra senses are a toggle in the studio, and the section on when the cheap way wins is about the footage where you should switch them off.
A transcript isa record of speech.
That is the whole of it. Everything below is either in that record or it is not, and no amount of engineering puts a scoreboard into a string of words.
| Signal in the footage | In a transcript | In all three inputs |
|---|---|---|
| Words that were spoken | Present, with per-word timings. | Present, with per-word timings. Same transcript. |
| A stretch where nobody speaks | Empty. Nothing to transcribe, so nothing to score. | Frames and sound texture carry it. |
| Text burned into the picture | Absent. A scoreboard is not audio. | Read off sampled frames. |
| Sound that is not speech | Dropped before scoring. A crowd erupting leaves no words behind. | Level, spectral centre, flux and flatness, measured across the window. |
| A commercial, a slate, ten seconds of black | Reads as ordinary speech, or as a gap between sentences. | Visible in the frames, and rejected at review. |
| Where the subject is in the frame | Not a property of audio. | Tracked, which is what the crop follows. |
| Whether the picture agrees with the words | Only one of the two is in the file. | Both are, so they are allowed to disagree. |
Claims about what an input can physically contain, not about anyone’s product.
Three filesthat settle it.
When everything worth clipping is said out loud, both mechanisms are reading the same evidence and they will land in much the same place. These are the three shapes of footage where they cannot, because one of them is working from an input that does not contain the answer.
Nobody is talking
A game filmed on one camera. No commentary bed, no microphone near the pitch, ninety minutes of ambient noise.
Transcript only
The transcript is close to empty for the whole file. There is no text to rank, so a text-ranking pipeline either returns nothing or spreads clips evenly and calls it a result.
Watch, listen, read
The frames become the detector. A run across the frame, a shot, the touchline reacting - all of it is in the picture, and the picture is an input.
The context is drawn on screen
Fourth quarter, 2:11 on the clock, 21-17 on the graphic. Nobody reads any of that out loud, because everyone watching can see it.
Transcript only
Score, clock and period are pixels. They were never spoken, so they are not in the transcript at any point in the file.
Watch, listen, read
On-screen text is read from the sampled frames, used to judge the moment and to title the clip that comes out.
The payoff is visual
Someone says two words - "oh no" - and the next four seconds of picture explain exactly why.
Transcript only
Two flat words. They can score low and lose a great clip, or score high for a reason that has nothing to do with what happened. Text alone cannot tell those apart.
Watch, listen, read
The frames say what happened. The score moves because of the event, not because of the phrasing.
Case C runs the other way too. A great line delivered over a title card, a slate or a dissolve scores well on text alone and ships as ten seconds of nothing.
Sometimes the wordsare the whole file.
If every moment worth clipping is announced out loud, the frames only confirm what you already knew and you paid image tokens for the confirmation. A comparison that concedes nothing is an advertisement. Here is the concession, with the bill attached.
A studio podcast
Two people, one framing, nothing moving. The entire value is in the sentence, and the sentence is in the transcript.
A talk or a webinar
The argument is spoken from end to end. Slides add a little; the words carry almost all of it.
A recorded meeting
Decisions get stated out loud. Very little happens in the picture that was not also said.
Anything already timestamped
Chapters, a run of show, a play log. No detector needs to run at all, by either mechanism.
Watch + listen on
The default. Frames, sound texture and transcript on every candidate.
Watch + listen off
No frames sampled, no image tokens. The words carry selection.
Your own timestamps
No detection runs at all. Only the transcript captions need is billed.
One hour of source, captions on. Quoted by the code that bills the job.
Talk leads with the transcript
Tell it the footage is talk and the words drive selection. Action leads with the pictures. Auto reads the footage and tells you which way it went.
The review is a toggle, not a tax
Watch + listen is a switch in the upload panel, and the credit estimate next to it is computed by the same code that bills you. Turn it off and watch the number fall before you spend anything.
Timestamps beat both detectors
Paste chapters or a play log and no detection runs at all - no sweep, no review, no rank pass, so the start-up charge drops to the delivery and the renders and the source bills at the anchored rate. Captions still need a transcript, and that part is billed when you ask for it.
Neither mechanism predicts what an audience will share. The score rates whether a moment stands alone. Anything beyond that is a forecast, and nobody can make one from a file.
Everything it does.And what it does not.
A mechanism argument is easy to win on paper, so here is the unflattering half: the full capability list, including the rows where the honest answer is that it is not built.
Finds moments in any footage
Frames, audio texture and transcript judged together, so it works with or without dialogue
Sports, podcasts, streams, talks
SignatureReframing that follows the action
Motion tracking for anything that moves, face-centring for talking heads, or no crop at all
Vertical, square and wide
SignatureExplains every score
Each clip reports what was seen, heard and read to arrive at its number
Auditable, not a black box
SignatureTimestamps you already have
Paste chapters or notes and detection is skipped entirely
Billed at the anchored rate
SignatureWord-timed captions
Six styles, adjustable size and position, your accent colour, SRT sidecar
Burned in or off
IncludedWrites the post caption
Six voices, with alt text and specific hashtags, grounded in what the clip shows
Corporate through to fun
SignatureEditor
Trim on a timeline, change shape, reframe, caption treatment, then re-cut
Included on every plan
IncludedFiller and pause removal
One click strikes the ums and uhs and closes the long silences - by cutting your own footage through the editor's existing span mechanism, so nothing is added or generated
Tighten, in the editor
IncludedAudio cleanup
Loudness evened and background noise reduced on the audio you recorded, so a clip sounds finished without ever leaving your footage
Your sound, cleaned not faked
IncludedExport
One clip, or every clip and caption file in a single zip
Never watermarked
IncludedLanguages
Transcription auto-detects the language rather than asking you to pick
Broad language coverage
IncludedYour footage is real, and stays real
No generated b-roll, no stock inserts, no synthetic voice-over, no invented speaker. Every second that comes out is a second you shot - cut, reframed and captioned, never fabricated
The line we will not cross
SignatureBrand kit
Logo on every clip, accent colour, default caption style and post voice
Set once, inherited
IncludedAgents can drive it
An MCP server with four tools, so a workflow can cut video without a person
cut, poll, estimate, list
Signature
Three things this category is expected to have that Cutlist does not do yet. If a capability is not on this page, we do not have it.
Intro and outro cards
Topping and tailing a clip with your own frames
Not built
Not builtTeam workspace
Shared projects and per-seat roles
Not built
Not builtPosting to social directly
Download and post yourself for now
Not built
Not built
We do not fake your footage.
No generated b-roll. No synthetic voice-over. No invented speaker. Cutlist finds the moment, cuts it, reframes it and captions it - and every second that comes out is a second you shot. A pipeline that fabricates footage can pad any file, which is exactly why we will not. Your footage is real, and it stays real.