Six hours.
Four good minutes.You cannot remember where.
A long session is not a video with some slow parts. It is mostly nothing, with four or five moments buried in it that you half remember and would have to scrub an hour to find. Cutlist is built for that shape: say roughly where you were good, get billed for only that stretch, and let it do the listening.
Clip the partyou remember.
Two timecode fields in the Direct the cut panel, marked Only. Set them and the pipeline never touches the rest of the file - no scan, no transcript, no model call on the ninety minutes of queue simulator you opened with. Billing follows the same window, because the studio and the meter call the same function.
You already know roughly when
Your VOD has chapters, your chat clipped something, or you just remember it was after the break. "About the second hour" is precise enough - the detection still runs inside the window.
The estimate moves as you type
Set the range and the credit figure on the button drops immediately - straight away for a dropped file, once a pasted link has been probed for its length. It is the same function that places the hold, so the number you see is the number you pay.
Two ranges, two runs
There is one window per job. A session with a hot start and a hot finish is two jobs, and the two together still cost a fraction of scanning the middle.
Or skip detection entirely
If you wrote timestamps down while you streamed, paste them. Nothing is detected, so the start-up charge falls to the delivery and the renders and the meter drops to the anchored rate - but the whole source is still billed, so on a six-hour VOD the range is the bigger lever. The arithmetic is below.
A range is a guess and a guess can be wrong. Cut it too tight and the moment you wanted starts thirty seconds before your window - the run comes back without it, and you paid for the window either way.
One session,three ways to pay for it.
A six-hour VOD, the full engine, captions on. One credit is one minute of source video at the base rate, rounded up; the see-and-listen review carries a per-minute surcharge, and there is a start-up charge for the half of the work that happens once however long the file is. The estimate itemises all three before you spend anything. Here is the same session billed three ways.
| The job | Billed | Credits |
|---|---|---|
| The whole VOD | 360 min x 1.8 + 219 start-up | 868 |
| Only 1:30:00 to 2:10:00 | 40 min x 1.8 + 105 start-up | 178 |
| Pasted timestamps, whole VOD | 360 source minutes x 0.98 | 354 |
The 1.8 is the per-minute rate with the see-and-listen review on, which costs more because the sweep looks at frames as the source gets longer. The 0.98 is the anchored rate multiplied back up by the transcript your captions still need. Neither figure includes the start-up charge, which is the other half of the bill and very nearly the half anchoring removes: 219.05 credits on the whole VOD against 1.15 when you bring the timestamps. Not zero - the clips still have to be encoded and delivered, and that is what the 1.15 is.
Forty minutes of range on a six-hour VOD is 690 credits back. Same engine, same clips out of the stretch that mattered. You decided the other five hours and twenty minutes were not worth reading, and the bill agrees with you.
Scanning a six-hour VOD end to end is the most expensive thing you can ask this product to do, so it wants Studio at 2,400 credits a month: 2 of them, or 13 ranged sessions. On Creator at 700 credits, a whole VOD does not fit at all, but 3 ranged sessions do. If you stream four nights a week, the range is not an optimisation, it is the difference between affordable and not.
Pasted timestamps bill the whole source. Anchors beat a range in the engine, so the file is opened end to end and the range you also set is ignored. On a six-hour VOD that lands above a tight window. Paste timestamps when your moments are scattered across the session, set a range when they are clustered.
A reaction is loudbefore it is articulate.
The moment you clip is usually a noise. A yell, a laugh that runs too long, the half-second of silence before it. A tool that reads only the transcript sees a line of dots there and moves on. Cutlist judges every candidate on the frames, the audio itself and the words together, so the noise is evidence rather than a gap.
- motion
- audio
- scene cut
- nominated
A long session is mostly level, with the reaction sitting on top of it. Dense scene changes are normal here rather than a signal, so cut density is penalised, not rewarded — otherwise every alt-tab in a six-hour stream would outrank the moment you actually want.
Illustrative shape · your own curves are measured, not drawn
Energy, not vocabulary
A sudden jump in level, a shout held past where a sentence would end, a room that goes quiet and then does not. None of that survives transcription, and all of it survives listening.
The frames back it up
Sampled frames go to the model with the audio. A kill feed, a boss health bar hitting zero, a face doing something a caption never would - context the transcript does not contain.
On-screen text is read
A scoreboard, a death banner, a donation overlay, a slide title. It is used both to judge the moment and to name the clip.
Ask for something specific
Type what you want in plain words: "the parts where I lose it", "every time the raid goes wrong". Selection narrows to that instead of to whatever scored highest.
It shows its working
Every clip comes back with a score from 0 to 99 and what it saw and heard to get there. You can disagree with a number you can read.
Loud is not the same as good, and the model is judging the moment, not forecasting the post. It rates whether a clip stands alone to somebody who was not there - nobody can predict a view count from a file.
A face detectorloses a screen share.
Your layout is a game filling the frame and a webcam box in one corner. A reframer that hunts for faces finds the small box, crops to it, and throws the game away - or finds a face inside the game and swings across the frame chasing it. That is not a tuning problem, it is the wrong signal for this footage.
Manual
Place the crop yourself with a slider and it stays there for the whole clip. For a layout that never moves - and yours never moves - this is one drag and it is correct on every clip after it.
Follow
Motion tracking, smoothed so it pans like an operator rather than snapping. The right answer for full-screen gameplay where the action crosses the frame.
Fit
The whole picture on a blurred bed. Nothing is cropped out, which is what you want when the HUD, the chat overlay or the inventory is the point.
Face
Centres on the speaker, and falls back to motion when there is no face to find. For the just-chatting hours, not the gameplay ones.
Every mode renders 9:16, 1:1 or 16:9 at full resolution rather than upscaling from a smaller cut, and the mode is switchable on a finished clip in the editor - so being wrong about it costs a re-cut, not a re-run.
There is no split layout: it will not stack your webcam above the gameplay in one vertical frame. One crop, one source rectangle, and Fit when you need to keep everything.
Ask for twenty,keep the four.
One run returns between 1 and 20 clips, five by default. Asking for more does cost more - the grounded review looks at three candidate windows per clip, so twenty clips is sixty windows of judgement rather than fifteen - but the six hours are still only read once, so twenty clips is 1071 credits against 868 for five. That is 23% more for four times the output, and you can bin sixteen of them. Ask for twenty.
Length bands
Under 30 for a single reaction, 30-60 for a bit that needs its setup, 60-90 when the build-up earns it. Or Auto, and each moment takes what it needs.
Everything arrives together
One zip with every clip and every caption file, named from the clip titles, or download them one at a time.
Post captions come with them
A written caption per clip in a voice you pick, plus alt text and a few specific hashtags. It describes what is actually in the clip - it will not invent a claim to make it sound bigger.
No watermark, ever
Nothing we render is stamped, on any plan, including the free trial.
Stream momentsalways run long.
You talk your way into a joke and then talk past the end of it. An automatic cut errs wide: it keeps the run-up you would have thrown away and lets the laugh run out. The editor is where that gets fixed, and a re-cut is billed at the anchored rate on the length you keep, because the moment is already found.
Trim to the frame
Drag the handles, nudge in tenths of a second, or use the arrow keys. The preview scrubs with you, so you set the in point by watching rather than by guessing.
Cut by deleting words
Click a word in the transcript to strike it. The footage goes with it, the silence goes with it, and the remainder rejoins seamlessly. Double-click any word to jump the preview there.
Reshape without re-running
Change aspect ratio or reframing mode on a finished clip. The source is not touched again and the detection is not paid for again.
Restyle the captions
Six styles, size from 0.6x to 1.6x, height anywhere in frame, previewed live before it renders.
What it will not do
Publish anywhere
No platform is connected and nothing is posted on your behalf. You download the clip and you post it, on your account, on your schedule.
Read your chat
No chat log is ingested and message spikes are not a signal. If your chat popped and the room stayed quiet, the moment can be missed - paste the timestamp and it will not be.
Understand your game
There is no per-title event detection. It has no idea what a clutch is in your game specifically. It reads the screen, hears the room and judges the moment like a stranger would.
Store your VODs
100 GB per account, clips live 30 days and source uploads are removed within 2 days of a finished render. It is working output, not an archive. A six-hour upload is a big file - download what you want to keep.