Glossary · Accessibility term
Captions
Also called: Closed captions, subtitles for the deaf and hard of hearing
Captions are a synchronized text version of everything you would hear in a video, covering speech and the non-speech sound that matters as well. Sound effects, music, laughter, and who is speaking. Speaker identification is the one to check for, because a conversation between three people is unfollowable without it and a caption file can look complete while carrying none. Closed captions are the kind the viewer can turn off, which usually means a separate file the player reads, and that separation is what lets a player offer resizing, restyling or another language where it supports them. Open captions are burned into the picture and can do none of that. The standard accepts either. Prerecorded video needs them at Level A, and live video needs them at Level AA, where there is no exception clause at all.
In practice
Automatic captioning has improved a great deal and still is not enough on its own. W3C's position is that automatically generated captions do not meet accessibility requirements unless they are confirmed to be fully accurate, and that they usually need significant editing. The failure mode is not clumsy wording. Drop one small word and the captions say the opposite of the audio, and the words that go wrong most are the ones carrying your meaning. Names, product names, prices and numbers. So generate them automatically, then read them against the audio and fix them. That second step is the requirement.
Two scope details are worth having. The rule is written about audio content in synchronized media, so a live audio-only stream with no picture is not covered by the Level AA live rule at all and lands at AAA instead. And the prerecorded rule carries an exception the live one does not. Where the video is itself an alternative to text already on the page, and you say so clearly, that rule does not apply to it.
Placement is a real production task rather than a convention. Captions should not obscure anything in the picture that matters, which is a note attached to the definition rather than a rule of its own, and it is why captions move rather than sitting at the bottom by default. And one question that comes up on every media project has a flat answer. An audio description track does not need captioning, because it is describing something already on screen.
Why it matters
The standard asks for captions that are accurate, which is a much higher bar than captions that are present. Captions garbling your product name and your prices satisfy a checkbox and fail the viewer, and the viewer here is not only deaf and hard of hearing people. W3C names a second audience explicitly, which is people who take written information in better than spoken. Both groups are reading the same file, and both of them are reading it instead of hearing you.
The cheapest test
Watch your own video with the sound off and only the captions to go on. Two things to check for. Can you tell who is speaking when the picture does not show it? And does any line say the opposite of what you meant, because a single dropped negative does exactly that and reads perfectly well on the page.
Where this shows up on the site
Related terms
- TranscriptA transcript is the text version of audio or video, and two different things go by that one word.
- Audio descriptionAudio description is narration added to a video's soundtrack that describes the visual detail the soundtrack alone does not carry.
- Assistive technologyAssistive technology is hardware or software that gives a person with a disability more than a mainstream browser offers on its own.
- CARTCART is live captioning written by a person rather than a machine, and the letters stand for Communication Access Realtime Translation.
Knowing the word is the easy part.
Find out where your own site stands. The free scan checks 10 pages in a real browser against all 90 supported automated rules, separates 27 best-practice checks from its WCAG findings, and names the rule behind every result.