Most short-form video is watched without sound, at least for the first second or two. Captions are not an accessibility afterthought in that context — they are the primary channel. Here are the rules that hold up across platforms, and the specific numbers behind them.
Vertical video is 1080×1920, and both ends of it are contested. The top ~250 pixels are covered by platform chrome. The bottom ~450 pixels carry the caption, the handle, the audio ticker and the action buttons — and on TikTok and Reels that region is taller than people expect.
That leaves roughly the middle 60% as reliably visible. Put your captions in the lower third of that safe area — around 55–65% of the way down the frame. This is low enough to sit under the speaker's face and high enough to clear the interface.
The common error is centering captions vertically. Dead center puts text across the speaker's mouth and eyes, which is the part of the frame the viewer is reading for expression.
On a 1080-wide frame, caption text should be 60–90 pixels tall for the body, which is roughly 6–8% of frame height. Emphasized words go larger, up to about 110 pixels.
Viewed on a desktop monitor while you edit, this looks absurd. On a phone, held at arm's length, in daylight, it is correct. Always check a caption at actual phone size before deciding it is too big — almost nobody errs on the side of too large.
One to four words at a time, swapped in rhythm with the speech. Full sentences require reading; two-word chunks are absorbed without reading, which is the point.
The exception is a deliberate pause. When the speaker stops, holding a complete short phrase on screen through the silence works well, because the viewer's eye has time to take it in.
Emphasis — color, size, a highlight box — should land on the word that carries the meaning, not on every noun. In practice that is:
Two emphasized words in a thirty-second video is about right. Five is too many; at that point nothing is emphasized.
A heavy sans-serif, 700 weight or above. Condensed faces fit more per line but lose legibility at speed; avoid anything below a semi-bold. Avoid true black or pure white as the only tool — text needs separation from the frame behind it, and weight alone will not provide it against a busy street.
Separation options, in order of how well they survive platform re-encoding:
White body text with one accent color for emphasis is the default for good reason. If you use brand colors, check them against the actual footage: a mid-green emphasis word disappears against foliage, and a warm orange vanishes against a sunset.
Whatever you choose, keep it identical across every video. Caption style is one of the few brand signals that survives the feed, and viewers recognize a consistent caption look faster than they recognize a logo.
Captions should appear with the word, not before it and not after. A frame or two early is better than a frame late — early reads as tight, late reads as broken.
Word-level timing beats line-level timing noticeably. When captions swap in sync with each word, the viewer's attention stays locked to the speech rhythm; when a whole line sits static for four seconds, the eye wanders.
Placing caption text behind the speaker — so they occlude it as they move — is a strong effect. It reads as deliberate design rather than an automatic subtitle track, and it creates depth in an otherwise flat frame.
It has two requirements. The segmentation has to be clean, because a halo around the speaker's shoulder is worse than no effect at all. And the text has to remain readable when partially occluded, which means shorter phrases and larger type than you would otherwise use.
Sparingly, and tied to a specific word. One emoji on the emphasized word of a video reads as native to the platform. Three or more reads as someone trying to look native, which is worse than plain text.
Burn captions into the video rather than relying on platform auto-captions. Auto-captions are inconsistently styled, occasionally wrong, sometimes off by default, and they sit wherever the platform decides — usually in the region you are trying to keep clear.
It is still worth uploading a caption file alongside, where the platform accepts one, for accessibility and for search.
Before posting, view the video on a phone, at full brightness and at low brightness, with the platform's interface overlaid. Most caption problems — text under the handle, an emphasis color that disappears, type that is fine on a laptop and unreadable in sunlight — are obvious in ten seconds of that check and invisible in an editor.