Key takeaways
- Most of the people who see your video are not listening to it. A 2019 Verizon Media and Publicis Media study found 92% of US respondents view mobile video with the sound off, and a 2026 survey put subtitle use among UK viewers at 79%.
- YouTube's automatic captions are a draft, not a deliverable. They are generated only in the video's default language, and they fail predictably on names, jargon, accents and anything with music under it.
- Correcting the auto track in Studio, or uploading an .srt you already have from your edit, takes under ten minutes per video and is the single highest-value accessibility job on the list.
- On Shorts, captions are not a settings toggle — they are design. Burned-in text is the norm, which means caption placement now competes with YouTube's own interface for the same pixels.
- The honest evidence for captions is about comprehension and completion, not rankings. Nobody at YouTube has ever described a caption file as a ranking input.
- Captions travel. A clean caption track is what auto-translated subtitles and auto-dubbing build on — and the one part of your packaging that never translates itself is the thumbnail.
There is a version of your video that most of your audience actually receives. It is not the one you exported. It is the one playing on a phone held at a bad angle in a kitchen, with the volume at zero, while someone else in the room is talking. The picture is doing all the work and the audio mix you spent forty minutes on is not in the room at all.
Captions are how that version of the video stays watchable. Most creators treat them as an accessibility chore filed somewhere after the thumbnail, the title and the end screen — something you get to when there is time, which there never is. That ordering is backwards, and it is backwards for a reason that has nothing to do with doing the right thing: the sound-off viewer is not a minority case any more, and a video that stops making sense without audio loses them in the first few seconds, in exactly the window where YouTube is deciding what your video is worth.
What follows is the whole job, worked out end to end. What YouTube's automatic captions actually do and where they break, how to fix them in a few minutes rather than an afternoon, how caption craft differs between a sixteen-minute upload and a forty-second Short, what the evidence for captions genuinely supports and what it does not, and where all of this meets the part of packaging that no transcription tool can help with.
The sound-off default is not a niche
The most-quoted number in this area comes from a 2019 study by Verizon Media and Publicis Media, an online survey of 5,616 US adults aged 18 to 54. It reported that 92% of respondents view videos with the sound off on mobile, that 80% were more likely to watch a video to the end when captions were available, and that half named sound-off viewing as the reason captions mattered to them. The reasons people gave were mundane: no headphones, a quiet space, waiting in a queue, doing something else at the same time.
That study is old now, and it is worth saying so plainly rather than laundering a seven-year-old figure into a fact about 2026. What has happened since is that the behaviour it measured stopped being remarkable. A survey published in May 2026 by XR Extreme Reach, covering 3,000 consumers across the UK, US, France, Germany and Spain, reported that 79% of UK viewers use subtitles at least sometimes, and that 59% of 18- to 24-year-olds use them always or often. The youngest viewers are the heaviest users, which is the reverse of what you would predict if subtitles were primarily a hearing-loss accommodation.
They are not primarily that, though hearing loss is the reason the feature exists and the reason it must be good rather than merely present. The World Health Organization puts the number of people living with some degree of hearing loss at around 1.5 billion, with 430 million needing rehabilitation for disabling hearing loss, and projects some 2.5 billion people with some degree of hearing loss by 2050. On a platform with billions of users, the share of your audience who cannot rely on your audio track is not a rounding error at any channel size.
Put the two groups together — people who cannot hear your audio and people who have chosen not to — and the sensible planning assumption is that a large minority of every video's audience is reading it. Not most videos. Every video.
Captions, subtitles, and the words YouTube uses for them
The vocabulary is muddled everywhere, including inside YouTube's own interface, so it is worth fixing the terms before the workflow.
| Term | What it actually means |
|---|---|
| Captions | Same-language text of the speech, plus the non-speech audio that carries meaning: a door slamming, laughter, the music sting. Written for someone who cannot hear the track. |
| Subtitles | Text in a different language from the audio. Assumes the viewer can hear, and only needs the words translated. |
| Automatic captions | YouTube's speech recognition output, generated on upload in the video's default language when the system can produce them. Published automatically unless you turn them off. |
| Closed captions | A separate track the viewer switches on or off. On YouTube this is the CC button, and it is what lives in the Subtitles tab in Studio. |
| Open or burned-in captions | Text rendered into the picture at export. Cannot be turned off, cannot be translated, cannot be restyled by the viewer. The default on short-form video. |
| Auto-translate | A viewer-side control that machine-translates whatever caption track exists into their language. It needs a caption track to exist first. |
| Transcript | The full text of the video, timed, exposed in the watch page's transcript panel and searchable there by viewers. |
The distinction that matters most for planning is the last-but-one. Auto-translate is not a separate thing you enable. It runs on the caption track you already have, which means the quality of your English captions silently sets the ceiling for every other language a viewer might read you in. A garbled auto-caption in the source language becomes a confidently garbled translation in thirty others.
What automatic captions get right, and where they fail
YouTube's speech recognition has improved enormously and is genuinely usable for clean, single-speaker audio. YouTube's own documentation carries the warning that matters: automatic captions may misrepresent what was said because of mispronunciation, accent, dialect or background noise, and creators are advised to review and correct them.
The failures are not random, which is good news, because predictable failures can be checked in a fixed amount of time. The recurring ones:
- Proper nouns. Your channel name, your guest's name, product names, game titles. The words most likely to be searched, and the words the model is least likely to know.
- Domain jargon. Every niche has fifty terms that a general speech model has barely seen. Finance, tabletop gaming, mechanical keyboards, obstetrics — it does not matter which, it happens in all of them.
- Music beds and effects. Continuous music under a voiceover degrades recognition and produces the run-on caption that has no punctuation for nine seconds.
- Overlapping speech. Two people talking over each other in a podcast produces one merged, unattributed line, with no indication of who said what.
- Accents and code-switching. Recognition quality is uneven across accents within a language and degrades further when a speaker switches languages mid-sentence.
- Numbers and units. "Fifteen hundred" becomes "1500" or "15 hundred" inconsistently within the same video, which reads badly and ruins a transcript search.
There is also a structural limit that surprises people: automatic captions are generated in the video's default language only. There is no automatic Spanish caption track on an English video. What a Spanish viewer gets is auto-translate running over the English track — a translation of the errors included — unless you upload a real Spanish subtitle file.
The default-language trap
Set your channel's and your video's default language deliberately in upload defaults. Get it wrong and YouTube generates the automatic captions with the wrong recogniser, which produces a track that is not merely inaccurate but nonsensical — and that nonsense is what every translated subtitle and every auto-dubbed audio track will be built from.
The fix that takes eight minutes
There are three routes to a correct caption track, and the right one depends on what your edit already produced.
Route one: correct the automatic track. Open the video in Studio, go to Subtitles, duplicate the automatic captions and edit. The timing is already done; you are only fixing words. On a fifteen-minute talking-head video with decent audio this is a five-to-ten-minute job, and most of it is one pass on names and jargon. The corrected version replaces the automatic track as a proper caption track.
Route two: upload a file. If your editor already generated captions — most modern editors and every short-form captioning tool will export one — you have the file. YouTube accepts several caption file types. The simple ones, .srt and .sbv, carry only timings and text and can be edited in any plain text editor. The richer ones, .vtt and .ttml/.dfxp, carry formatting and positioning as well. For nearly every creator, .srt is the correct answer: it is universally supported, trivially editable, and reusable on every other platform you post to.
Route three: paste a transcript. If you scripted the video, you already wrote the captions. Paste the script into Studio's transcript box and let auto-sync align it to the audio. This produces the best result of the three, because the text is correct by construction — provided you actually said what you wrote, and provided you go back and fix the places where you improvised.
None of these is a large job. The reason captions go undone is not cost, it is sequencing: they sit at the end of a publishing checklist that is already running late, and they are the one item with no visible consequence if skipped. Move them earlier. A useful rule is that the caption file gets created in the edit, at export, alongside the thumbnail still — not in Studio after the upload bar fills.
Writing captions somebody can actually read
A correct transcript is not automatically a good caption track. Broadcast subtitling practice has spent decades on this and converges on a small set of conventions that transfer directly to YouTube.
- Around 42 characters a line, two lines maximum. Longer lines force horizontal eye movement that costs more than it saves, and a third line starts eating the picture.
- Break lines at grammatical joints. Split after a clause, not between an adjective and its noun. A badly broken line reads as a stumble even when every word is right.
- Respect reading speed. A caption that flashes for half a second is decoration. If speech is faster than a viewer can read, lightly condense — drop the "sort of", keep the meaning.
- Label speakers when there is more than one. In interviews and podcasts, an unattributed caption track is genuinely confusing; a two-word label fixes it.
- Caption the audio that carries meaning. If the joke is a sound effect, the sound effect is part of the joke. If a music cue signals a section change, note it.
- Punctuate properly. Automatic captions are famously light on commas and full stops, and punctuation is most of what makes a wall of text parse.
One thing worth not doing: keyword-stuffing the caption track. It reads badly to humans, and the theory it rests on — that the caption file is a ranking lever — is not one YouTube has ever endorsed. More on that below.
Shorts: where captions stop being a setting and become design
Everything above assumes a caption track that the viewer can toggle. On short-form, that model mostly breaks down, and the practical standard across Shorts, Reels and TikTok is burned-in text: captions rendered into the frame at export, styled, often animated word by word.
There are good reasons for that. Short-form plays muted by default in a feed, so the captions have to be there before any viewer decision is made. The platforms render their own interface over the video, so a toggleable track's position is not fully in your control. And the styled, high-contrast caption has become a format convention in its own right — viewers read short-form the way they read a poster.
The cost is that burned-in captions cannot be translated, cannot be turned off by someone who finds them distracting, and cannot be restyled by a viewer who needs larger text. The pragmatic answer is both: burn in the captions for the format, and still let YouTube generate or accept a caption track on the upload so that assistive technology and auto-translate have something to work with.
The placement problem is the interesting one, because it is a packaging problem rather than a text problem. The vertical frame is not yours end to end: the top strip carries the title and channel treatment, the right-hand column carries the like, comment and share controls, and the bottom-left carries the channel name, description line and subscribe affordance. That interface has grown over the last two years rather than shrunk. Captions placed low and centre — the instinct from widescreen — collide with it. Captions placed in the middle third, sized large, with real contrast behind them, survive.
The same reasoning governs any text you put on a vertical frame, which is why the detail lives in the Shorts thumbnail guide: the safe area is smaller than the video, and everything that matters has to live inside it. If you want to check a frame before you commit to an export, the thumbnail preview tool will show you how a composition holds up at the sizes people actually see.
The evidence, honestly
There is a widely circulated case study from Discovery Digital Networks, run with the captioning vendor 3Play Media, which found a 7.32% increase in views on captioned videos, with the largest effect — 13.48% — in the first fourteen days after captions were added. It is the number every captioning article quotes, this one included.
It is also a single vendor case study from over a decade ago, on one publisher's catalogue, in a very different YouTube. Treat it as suggestive rather than as a rate card. If someone promises you 7% more views for adding an .srt, they are selling captioning services.
The stronger evidence sits in a less commercial place. Meta-analyses of captioned video in language learning — including work published in Language Learning and System — consistently find positive effects on listening comprehension and vocabulary acquisition, with same-language captions performing best. That literature also carries an important caveat from cognitive load theory: reading text while watching pictures and listening to speech splits attention, and past a certain density it makes comprehension worse rather than better. Captions help because they add a redundant channel, not because more text is always better. This is exactly why line length and reading speed matter, and why animated word-by-word captions on an already busy Short can work against you.
As for search: YouTube has never described a caption file as a ranking input, and the honest account of what does drive discovery is in our YouTube SEO guide. What captions demonstrably do is populate the transcript panel, which viewers can search within a video — the same "find the bit I want" behaviour that chapters and key moments serve. The defensible claim is that captions make a video usable by more people in more contexts, and that this shows up in completion rather than in rankings. That is a smaller claim than the industry usually makes. It is also the one that survives scrutiny.
What happened to community captions
Anyone who remembers YouTube before 2020 remembers community contributions: viewers could submit caption and subtitle tracks for a creator to approve. For channels with deaf audiences or large international followings, volunteers produced dozens of language tracks for free.
YouTube switched the feature off on 28 September 2020, saying it was used on a vanishingly small share of channels — fewer than 0.001% — and had become a persistent source of spam and abuse. The response was loud: a petition gathered over half a million signatures, and deaf creators and accessibility advocates argued that a feature used by few channels was load-bearing for exactly the channels that needed it. YouTube offered affected creators a six-month subscription to the captioning service Amara as a consolation.
The reason to retell this is not nostalgia. It is that the replacement for volunteer labour turned out to be automation, and automation has a specific failure profile: it is fast, it is free, it is available on every video, and it is confidently wrong in the places that matter most to you. The correcting pass that community contributors used to do is now yours. It is eight minutes. Do it on the videos that will still be getting views in two years, and let the rest ride on the automatic track.
Translation, dubbing, and the part that does not translate
Captions are the foundation of everything YouTube now does with language. Auto-translate lets a viewer read your caption track in their own language. Auto-dubbing goes further and replaces the audio: YouTube has expanded auto-dubbing to all creators, covering 27 languages, with an "Expressive Speech" layer that carries tone and is available first in a smaller set of eight languages including English, French, German, Hindi, Indonesian, Italian, Portuguese and Spanish. YouTube has said that in December it averaged more than six million daily viewers watching at least ten minutes of auto-dubbed content.
All of that machinery reads from your speech. A clean caption track and a clearly set default language are what stop the chain compounding errors — mis-transcribed name, mis-translated subtitle, mis-pronounced dub.
And then there is the one surface the machinery cannot touch. A viewer in São Paulo who is served your video with Portuguese audio still sees the thumbnail you designed, with the English words you set in it, at the size of a postage stamp. Translated audio expands your reach; untranslated packaging caps it. That trade-off, and what to do about it, is the subject of the post on auto-dubbing and localised thumbnails.
Accessibility is a floor, not a growth tactic
It is tempting to sell captions entirely on performance, because performance is what creators buy. It is worth being straight that the case does not depend on it.
Regulation has been moving in one direction. The European Accessibility Act became applicable on 28 June 2025, setting accessibility requirements for a defined set of products and services sold in the EU, including elements of e-commerce and audiovisual media service access features. It does not impose a captioning duty on an individual YouTuber — services provided by microenterprises, under ten staff and under €2 million turnover, are exempt from the service requirements, and most creators are nowhere near the covered categories. But it does change the baseline for the businesses creators work with. A brand commissioning a video for its own site in the EU increasingly has an accessibility requirement attached, and "we supply a caption file with delivery" is turning into a line on media kits rather than a favour. If you take brand work, treat captions as a deliverable and say so.
The rest of the argument needs no regulation. Somebody who cannot hear your video is a viewer. The work to include them is measured in minutes.
A per-upload checklist
- Set the video's default language correctly at upload. Check your channel's upload defaults once so you stop doing it per video.
- Export a caption file from your editor if it can produce one. If not, let YouTube generate automatic captions.
- Open Subtitles in Studio and do one correcting pass, hunting names, jargon, numbers and anywhere the punctuation collapsed.
- Check two or three places where music runs under the voice — that is where the worst lines hide.
- For Shorts, burn in captions within the safe area, sized for a phone at arm's length, with contrast behind the text.
- Watch the first thirty seconds muted, at phone size. If the video is incomprehensible without audio, the problem is the edit, not the captions.
- On evergreen videos, add a real subtitle file for your largest non-native audience rather than relying on auto-translate.
Step six is the one that changes how people make videos. A muted first thirty seconds exposes every place where the picture is carrying nothing — the talking head with no visual support, the claim made only in voiceover, the joke that lands only in the delivery. Fixing those improves the video for everyone, not only the sound-off viewer, which is the same argument as the one made in the post on the first thirty seconds.
Where this meets the thumbnail
Every decision in this article happens after the click. Captions are the reason a muted viewer stays. They have no bearing whatsoever on whether that viewer arrives, because before the click there is no audio, no captions and no video — only a still image and a line of text, at the size of a stamp, in a feed of competitors.
That is worth sitting with, because it is the same skill on both sides of the click. The sound-off viewer and the pre-click browser are the same person a few seconds apart, and both are reading rather than listening. The craft of making a frame legible at speed — real contrast, few words, one focal point, text large enough to survive a phone — is what a thumbnail is, and it is what a caption on a Short is. The habits transfer, which is the argument made at length in the post on thumbnail composition.
Caption work is unglamorous and finite: a few minutes per upload, a fixed list of predictable errors, a real gain in who can watch you. Packaging work is neither finite nor predictable, which is why it is where most of a creator's time goes and where the tooling earns its place. Thumblore exists for that half — generating thumbnail options fast enough that testing a second idea costs you a minute instead of an evening — and it is deliberately no help at all with your caption file. Do that one yourself, in Studio, before you close the tab.
The through-line is simple enough to state in one sentence. Your audience is reading your video, before the click and during it. Design for the reader.