Caption best practices

Short-form caption best practices

Most short-form is watched with the sound off, which makes the captions the video. These are the choices that decide whether they are read or ignored.

Show a few words at a time, not a paragraph

Three to six words on screen at once is the range that reads at short-form pace. Much more and the eye has to scan while the video moves, which is when people stop reading and then stop watching. Much less — a single word at a time — is a style rather than a default, and it costs you the ability to see a phrase as a phrase. Break lines where the sentence breaks, not where the character count runs out.

Time to the word, not to the sentence

Captions that appear a beat after the words are spoken feel broken even when the viewer cannot say why. Word-level timing fixes this, and it is what makes highlight-as-spoken styles work at all. If your captions are drifting, the usual cause is that they were timed to sentence boundaries and the sentence was long.

One face, one size, high contrast

Pick a heavy sans-serif and keep it. Add an outline or a shadow — a caption sitting on a bright background with no separation is unreadable for exactly the frames where it matters. Do not shrink text to fit a long line; break the line instead. And check the size on an actual phone rather than in the editor, because a caption that is comfortable at 40% zoom on a laptop can be too small in the hand.

The first three seconds are the whole game

Whatever you do with the rest of the clip, the opening line has to make sense with no context and no sound. That usually means starting on the sentence that sets up the point rather than on the point itself, and it means never opening on a subordinate clause. If the first caption reads as the middle of something, the clip is cut a few seconds too late.

Read them before you publish

Automatic captions mishear names, technical terms and anything in a second language, and they do it confidently. The cost of an unread caption is a clip that says something you did not say. Read the lines on the finished frame — not in the transcript, where a mistake looks like a typo instead of like the thing a viewer will screenshot.

Questions

How many words should be on screen at once?

Three to six for most short-form. That is enough to read as a phrase and few enough to take in without scanning. Fewer than three can work as a deliberate style but costs you phrasing; more than six is where people stop reading.

Are word-by-word captions better?

They are better at holding attention and worse at conveying a phrase. They suit punchy, quotable lines and suit dense explanation poorly. The bigger risk is practical: a word-by-word style redraws the whole group every frame, so if the line is too wide to fit it is cut off for the entire clip rather than momentarily.

Where should captions sit on the frame?

In the middle third, not at the bottom. The bottom band belongs to the platform’s own interface and, increasingly, to its automatic captions. See the safe zones guide for the exact edges.

Do burned-in captions hurt accessibility?

They are not a substitute for a caption track a screen reader can use, so where a platform accepts an upload with a separate caption file, provide one as well. On short-form platforms that do not, burned-in captions are the only thing standing between a muted viewer and nothing at all.

Basic includes the whole workflow.

Import a source, generate clips, review every one and publish to your own connected channels. Upgrade when you need more.

Caption a clip free