Video accessibility gets flattened into add captions, and captions are one of three obligations. They serve different people, they are not interchangeable, and which ones apply to you depends entirely on what kind of media you actually have. That last part is the piece nobody publishes, so start there.
First, Name What Kind of Media You Have
W3C built a routing rule into the criterion names themselves, and once you know it the whole guideline collapses into one decision. Every criterion under time-based media carries the media type in its own title. If your media is audio-only or video-only, the only criteria that apply are the ones with audio-only or video-only in the name, and the single exception is at Level AAA, where the media alternative rule reaches video-only content too. If your media is neither of those, every other criterion in the group applies to you.
| What you have | What you owe at A | What Level AA adds |
|---|---|---|
| Audio only, prerecorded. A podcast, an interview recording | A transcript, under 1.2.1 | Nothing. Audio description does not apply, because there is nothing to see |
| Video only, prerecorded. A silent screen recording, a soundless loop | A description or a transcript, under 1.2.1 | Nothing. Captions and audio description are both scoped to media that has both tracks |
| Video with sound, prerecorded | Captions under 1.2.2, plus either audio description or a full text alternative under 1.2.3 | 1.2.5 removes the text-alternative option, so audio description itself is required |
| Video with sound, live | Nothing at A for the live stream itself | Live captions, under 1.2.4 |
Notice where 1.2.3 and 1.2.5 stack, because that is the decision with money attached. At Level A you can discharge it with a full text alternative. At AA the audio description itself is required and the text route disappears. Level AA is what nearly every law names, so if the visuals in your videos carry meaning, plan for description from the start rather than writing a transcript twice.
One trap sits underneath the table. Audio-only and video-only are defined terms, and both of them exclude interaction. A podcast with no interaction is audio-only. A podcast with a prompt in it that says click now is not, because the interaction has to happen at a particular moment and a transcript reading click now tells nobody when. The same goes for a silent training video with questions in it. Add interaction and you have synchronized media, and the whole set of criteria applies.
The Exception That Answers What Clients Actually Ask
The question is always some version of this. We already published the help article, then we made a video of it, do we now have to caption and transcribe the video as well. The answer is no, on two conditions. If the media presents no more information than is already in the text on the page, and it is clearly labeled as an alternative to that text, then it is a media alternative for text and three criteria step back from it.
Two failure techniques bracket that exception from opposite sides, and both are easy to walk into. One catches you when you claim the exception and the video actually adds something the article does not have, which is most reshoots. The other catches you when the video genuinely is an alternative and nobody ever labeled it as one. The label is not decoration. It is half the exception.
What Accurate Captions Actually Means
The standard people assume is word for word, and it is not. W3C's own failure technique asks for the dialogue verbatim or in essence, plus every sound that matters, and it says plainly that simplifying spoken text is standard procedure and does not invalidate a caption. That is a materially easier bar than transcription, and it is the bar a professional captioner is already working to. The failure is dropping dialogue or dropping important sounds, not condensing a rambling sentence.
Worth knowing that transcripts are held to the opposite standard. W3C's worked example of a conforming transcript is a verbatim record of everything the speakers say. So captions may be condensed and transcripts should not be, and that difference decides which one you can safely generate from the other.
Auto-Captions Are a Draft, Not a Deliverable
W3C settles this in one sentence. Automatically generated captions do not meet user needs or accessibility requirements unless they are confirmed to be fully accurate, and usually they need significant editing. The exit condition is the important half. Confirmed fully accurate, by somebody, not generated and left.
W3C's own example is better than the ones people usually reach for. Spoken, the recipe says broil on high for 4 to 5 minutes, you should not preheat the oven. Captioned automatically, it came out as broil on high for 45 minutes, you should know to preheat the oven. Not a garbled name. A dropped not, which reverses the instruction, and a lost hyphen that multiplies the cooking time by ten.
The predictable failures are the words carrying your meaning. Proper nouns, which means your company name, your product names and the speakers. Technical vocabulary. Numbers, prices and dates. Accents, and any moment where two people overlap. Punctuation, which changes meaning more than people expect. And non-speech sound that matters, such as a doorbell, an alarm, or laughter.
So generate automatically and then edit, which is the workable process. What we would not do is repeat the industry line about how much faster correcting is than starting from scratch, because W3C's own note runs the other way. Correcting an automatic caption file takes quite a bit of time for anybody who does not do it regularly, and W3C says that is why many organizations outsource captions instead. Budget for a real pass, read it out loud against the video once, and then publish. Across a library rather than a single file, that pass turns into a sampling job, and how to check caption quality at scale sets out the acceptance test to write first.
The cheapest test there is
Watch your own video with the sound off, using only the captions. If you would be lost, confused, or mildly annoyed, so is everyone who relies on them. It takes as long as the video does, and it finds the missing speaker labels and the missing doorbell that no tool will.
Who Captions Are For, and What Subtitles Are Not
Captions serve people who are deaf or hard of hearing, and W3C names one more audience that is worth quoting because it is a disability argument rather than a convenience one. People who process written information better than audio. The other thing you often hear, that captions help anybody in a meeting or on a train or without headphones, is a real benefit and it is a marketing argument. Use it if it helps you sell the work, and do not put it in a conformance report.
There is a distinction underneath this that costs clients real money. W3C splits the two words by language. Captions are the same language as the spoken audio. Subtitles are a translation into another language. And W3C says outright that subtitles in other languages are not directly an accessibility accommodation. So a Spanish subtitle track on an English video is a fine thing to have and it does not discharge the captions rule. We have watched that mistake sit unnoticed on a video library for a year.
Open, Closed, and Burned-In
Closed captions can be turned on and off, come from a separate file, and are the right default. Three formats are in use. WebVTT is the most common on the web, with SRT and TTML behind it. On an HTML5 player the file arrives through a track element, and the same element carries an audio description track.
Open captions are burned into the picture. They cannot be turned off and they are pixels rather than text. None of that makes them a failure. W3C publishes always-visible open captions as a sufficient technique in its own right, and lists them ahead of closed captions. The resize rule and the text spacing rule both except captions outright, and the spacing rule names captions embedded in video frames specifically, so the case against open captions is user need rather than conformance. They are reasonable on social platforms that autoplay silently, and a poor choice on a site you control.
One claim about closed captions needs qualifying, because we used to make it flatly. They are restylable in principle and support for caption styling across browsers and players is inconsistent and sometimes unreliable, so most web video ends up in the player's default presentation, which is white text in a black box. Treat restyling as a capability rather than a promise.
What Goes in a Caption Beyond the Words
Captions carry sound, not speech, and that is where most hand-written caption files fall short. Identify who is speaking when it is not obvious. Put non-speech sound that matters in brackets, so a doorbell, an alarm, or a laugh reaches somebody who cannot hear it.
Music is the case people freeze on, and W3C has worked it through. Its example captions non-vocal music by title, movement, composer, and anything else that helps somebody understand the nature of the audio, so a bracketed suite and movement, a bracketed composer credit, and then a plain description of the music itself wrapped in music-note glyphs. That is more useful than the word music in brackets, and it is not more work once you know the convention.
Audio Description, and the Escape Most Videos Already Take
The test to internalize is to close your eyes and listen. If you miss something that matters, you need description. A talking head needs none. A product demo where somebody says you just click here needs it badly.
There is a rule behind that instinct and stating it plainly is worth money. If all the important information in the video track is already in the audio track, no additional audio description is necessary. The same note sits on the Level A criterion, the Level AA one, and the AAA one. That is the difference between telling a client they need description on two hundred videos and telling them they need it on the handful where the picture says something the narration does not.
Where you do need it, there are three ways to deliver it, and they are not equally expensive. The description can be integrated, meaning the speakers say the visual information out loud as part of the script. It can be a separate timed text file. Or it can be a separate audio track, or a whole second version of the video. W3C is direct that the first one is the cheap one, because a video planned with description in mind from the start often needs no separate description at all, and therefore costs nothing extra. That is the argument for putting this in the brief rather than in the remediation budget.
Two boundaries save arguments. Description does not apply to audio-only content, because there is nothing to see. And description of live video is a genuine user need that WCAG does not require, which is the line a live-video business most needs drawn for them.
Two details people miss. Every piece of text shown on screen has to reach the description somehow, whether that is a title card, a lower third, a speaker's name, or a statistic that only ever appears as a graphic. And a described audio track does not need captioning of its own, so solving the description rule does not quietly create a second captioning job.
What Goes in a Transcript
A transcript is not just the words. The standard's own definition asks for correctly sequenced text descriptions of the visual and auditory information, and then adds a clause almost nobody knows about. It has to provide a means of achieving the outcomes of any time-based interaction. So if the video says click here now, the transcript owes the reader a working route to the same outcome, not a note that a button appeared at four minutes.
Beyond that, W3C's worked example asks for speaker identification and the significant non-speech sounds, naming applause, laughter, and questions from the audience. And for somebody who is both deaf and blind, reaching your video through a braille display, the required form is a descriptive transcript, which is the speech, the non-speech audio, and a text description of the visual information together in one document. That is the only route into a video for that reader, and it is why a transcript is not an optional extra on top of captions.
The practical argument for writing one is that you have most of it already. Captions and transcripts hold the same text, so one can be built from the other, and most caption-editing tools will export a plain text transcript for you. A caption file also gives you something clients like for reasons nothing to do with compliance, because some players turn it into an interactive transcript that highlights each phrase as it is spoken and lets somebody click a line to jump to that point in the video.
Put the transcript on the same page as the video rather than behind a download. A transcript nobody can find is a transcript nobody uses. It is also the most direct overlap with accessibility and SEO, since text is what makes an hour of video findable at all.
Live Video Runs on Two Clocks
A live event creates two obligations with two different deadlines. The stream itself needs live captions at Level AA. The recording you publish afterwards needs captions at Level A, which is the baseline every AA audit already includes, so the archive is the one you cannot skip. W3C names the handover between them as well. Live captions carried straight over to the recording will usually need minor editing for accuracy before you publish.
What you buy for the live half has a name. Real-time captioning is done by professional captioners, often called CART providers, who can work from the audio over a phone or internet connection rather than being in the room. Knowing the term is most of the procurement problem solved.
W3C is unusually honest about the hard case here, and it is worth reading to any newsroom. It acknowledges there may be real difficulty captioning time-sensitive material, and that this can leave an author choosing between delaying the information until captions exist and publishing something inaccessible to deaf viewers. It does not excuse the choice. It just declines to pretend the choice is easy.
The Player Itself Has to Work
Easy to forget, because the content gets all the attention and the player is somebody else's code. It is still an interface on your page, and it has its own guide. The short version is that every control has to be reachable and operable by keyboard, every button needs a real name rather than announcing as button, focus has to be visible on the controls, controls must not vanish on a timer while a keyboard user is reaching for them, and audio must not start on its own and run for more than three seconds without a way to stop it, which is 1.4.2 at Level A.
One honest limit
Four of the nine criteria in this group have no automated rule written for them at all, and every rule that does exist is still at proposed status rather than approved. So a scanner can tell you a caption track is attached. It cannot tell you the captions are accurate, in sync, or that they name the speakers, and it cannot tell you your audio description missed the thing that mattered. Media accessibility is judged by watching and listening, which is why our video and media review is a human pass.