7 min read

How to clip a video: from in point to export

The workshop manual for what happens after you pick the moment: in point, reframe, captions, loudness, export. Eight checks, and one file every platform accepts.

Every tool page answers this question the same way: upload the video, press the button, export. That answer skips about a dozen decisions between a good moment and a postable clip, and those decisions get made either way, by you or by something making them on your behalf.

Choosing which moment to cut is a different job, and it is the one covered in what to cut. This piece starts from a moment already chosen and ends with a file ready to post. All of it can be done by hand, in any editor you already have.

Set the in point one beat before the first word

The cut does not land on the first syllable. It lands about half a second before it, on the breath, and never on the tail of the previous sentence. An in point that tight reads as a mistake, as though the clip started late. An in point that wide reads as an introduction, and an introduction is the thing viewers leave during.

The out point is the same rule reversed: the last consonant plus two or three tenths of a second, not the long silence that follows it. Silence at the end of a vertical clip is an invitation to scroll, and the loop back to the first frame is what you want instead.

The cut that clicks

Cut on a zero crossing of the audio waveform, or put a 20 ms fade on each end of the clip. A click at the join is small, and it is also the single thing that makes an otherwise correct clip sound amateur. The fade takes two seconds if you would rather not zoom in that far.

The first second and the first line are one decision

The opening line of the caption and the opening second of the video answer the same question: why should a stranger stay. Write them together, before you cut, not after you export.

This is also a test. If you cannot write a hook for a moment, the moment does not have one. Change the clip rather than writing a better caption around a weak one, because a caption that oversells the first second is how an audience learns not to trust the next post.

One operational rule worth carrying across every caption you write: on TikTok, five hashtags is the ceiling. Past that they stop working for you and start reading as noise.

Reframe: keyframes at the cuts, nowhere else

A 16:9 frame cropped down the middle gives you a vertical clip of the gap between two heads. The reframe has to follow whoever is speaking.

The rule that makes hand reframing tractable: the crop only moves at real scene cuts. Mark the shot changes first, set one position per shot, and never ease between them inside a sentence. Motion during speech reads as an error even when the framing it arrives at is correct, while a snap at a camera change reads as intentional, because it is.

Stacking both speakers, one above the other, is right in one case and wrong in most. It works when two people are talking over each other and the reaction is half the moment. It drains a monologue, because half the frame is then a person waiting.

Captions: one style, three hundred clips

Captions want to be readable at arm’s length, positioned clear of the bottom third where every platform puts its own interface, and identical across every clip you ever post. Consistency is the underrated part: the same typeface, the same position and the same colours across three hundred clips is what makes a feed look like a channel rather than a folder.

The figure usually quoted here, that 85% of social video is watched without sound, has no primary source. It was reported by publishers about Facebook in 2016 and never verified. The measured version is better anyway, and it is older than it looks. In an April 2019 survey of 5,616 US adults for Verizon Media and Publicis Media, 92% said they watch mobile video with the sound off and 50% said captions matter for that reason, while 80% said they are more likely to finish a video that has them. Seven years old, and US only, so treat it as the direction rather than today’s percentage.

Word by word highlighting helps on fast speech and on lines carrying numbers. It distracts on a calm sentence, where the eye is already ahead of the voice.

Check the detected language before you trust anything downstream

Language should be decided from the characters in the transcript, not from the label the transcription service attached to it. Those labels are wrong often enough to matter, particularly with dialect and code switching.

The cost of getting it wrong is visible and total: an Arabic clip routed down an English path comes back romanised, with the letters unjoined and the punctuation on the wrong side of the line. Ten seconds of checking at the top removes the entire category. The full test is in the four places Arabic tools break.

Loudness and length, the two settings people guess at

Normalise to a consistent loudness across clips rather than pushing the peak on each one. A feed where every clip arrives at a different volume makes the viewer reach for the phone, and reaching for the phone is where scrolling starts. We have not measured the platforms’ own loudness targets ourselves, so we are not going to quote you numbers for them. Pick one level and apply it to everything.

For length, 30 to 60 seconds is the reasonable default. Wistia’s 2026 report puts engagement on sub-minute video at 52%, and Loopex Digital’s May 2026 roundup puts Shorts above 40 seconds at 33% more engagement than shorter ones, which is why the floor of that window sits where it does rather than at 15 seconds.

There are two reasons to break it. A story with a turn in it sometimes needs 90 seconds to reach the turn, and cutting it to 45 leaves you with a setup and no payoff. And a long exchange worth keeping whole is better served as consecutive parts of one series than as a single clip nobody finishes. The reflex that shorter is always better costs you both.

Export once, post three times

One file satisfies all three platforms: H.264 video, the moov atom written at the front of the file, no edit list, AAC audio at 48 kHz, 9:16, under 300 MB. Export that and stop.

Keeping three export presets, one per platform, is the fastest known method of posting the wrong file to the wrong place at the end of a long day. The differences between the platforms are not in the container. They are in the publishing, which is genuinely awkward: TikTok’s own documentation states that content posted through unaudited clients stays private until the client passes an audit. That single sentence is why finished clips arrive on a phone and get shared out by hand, and the workflow around it is three taps per post.

The eight-point check before you post

Ninety seconds, in this order, every time.

  • In point on the breath, not on the first syllable or the previous sentence.
  • Out point two or three tenths after the last consonant, with no trailing silence.
  • No click at either join.
  • The speaker in frame for the whole duration, with the crop moving only at cuts.
  • Captions inside the safe area, clear of the bottom third.
  • Caption text spell checked and in the language actually spoken.
  • The hook written, and matching the first second.
  • File within spec and under 300 MB.

Nothing on that list makes a weak moment strong. It stops a strong one being thrown away by something mechanical, which is a smaller claim and a more useful one.

What a score can and cannot tell you

Plenty of tools hand back a 0 to 100 number and name it after an outcome they do not control. OpusClip, for example, scores every clip and lets you sort on it. A score is worth having. It ranks candidates inside one episode, which saves you reading time. It does not predict distribution, because distribution is decided by recommendation systems that change their behaviour without notice.

A number on its own is also unusable. What has to come back with it is the start, the end and the sentence that earned the cut, so that when a choice is wrong you can see why in two seconds instead of re-reading the transcript. That is the shape of the output on our own pipeline, and it is a fair thing to demand of any other.

Treat counts the same way. Competing guides advertise 5 to 10 usable clips per episode. We measure 4.3 clips per hour of source across 548 real projects, which on a two-hour show is nine, and we publish the average rather than our best day. The same applies to results: 13.5M views in 28 days means five client accounts between 14 August and 10 September 2026, and nothing about what your show will do.

Three hours, and where they go

Done by hand and done well, an episode costs roughly three hours: reading the transcript, choosing, cutting, reframing, captioning, writing the hooks. That is every episode, and it does not get much faster after the first month, because none of those steps has a shortcut that does not cost quality.

The moment to hand it over is not when it becomes boring. It is when the backlog stops shrinking, when new recordings arrive faster than you clear the old ones. At that point the arithmetic changes, and it is worth knowing that the unit being priced is recording time rather than clip count: our plans start at $19 a month for 1,200 minutes of source, which works out around 85 finished clips. If you have an archive and no intention of touching an editor at all, that is the other service.

Until then, a month of doing this by hand is not wasted. It is how you learn what a good clip looks like in your own material, and that knowledge is exactly what turns a score from a number into something you can argue with.

Clipping this by hand takes about three hours an episode. That is the job we took.

Start clipping