Skip to main content
WCAGrules
Quick navigation

WCAG 2.2 · Guideline 1.2 · Perceivable

Time-based Media

Audio and video need captions, transcripts, and description, because sound and picture do not reach everyone.

A video plays for everyone and what it says does not reach everyone. Deaf and hard-of-hearing people need captions, blind people need the visuals described, and deafblind people need a transcript they can read in braille. This guideline covers all three, live and recorded. What it asks of you depends on what your media actually holds, because every criterion here carries its media type inside its own name. A podcast answers the audio-only rules and nothing else. A silent demo video needs no captions at any level.

Five A/AA success criteria live here, out of nine. They cover captions for recorded and live video with sound, a transcript or an alternative for audio-only and video-only recordings, and description of what happens on screen while nobody is talking about it. The four above them, at Level AAA, add sign language and extended description. Guideline 1.2 is not a thing you pass or fail on its own, though. An audit runs against those five, and which of the five reach you is decided by the media you publish rather than by the guideline.

Test rules for this guideline

The W3C files these ACT rules against guideline 1.2 as a whole rather than any single criterion under it.

Rules
5
Level A
3
Level AA
2
Human testing only
4

All 5 Time-based Media rules

Level AAA in this guideline

4 enhanced criteria sit under 1.2. They are outside the level almost every law names, so our audits treat them as reference rather than scope. Some are still worth adopting, and the Level AAA hub says which.

What goes wrong here

These are the failures we find repeatedly under 1.2, across sites of every size.

Who it affects

  • Deaf and hard-of-hearing people, for whom captions are the content rather than a convenience. Meaningful non-speech sound is part of that content, not a nice extra.
  • Blind people, who need description when something happens on screen in silence. Where the narration already says it, no extra description is owed.
  • Deafblind people reading through braille, who need a text alternative rather than either channel. This is the audience a transcript genuinely rescues.
  • People who take in written information more reliably than spoken, which is a wider group than deafness alone.
  • Audiences reading a translated track. Subtitles translate; captions carry the sound. The two overlap where a translated track also carries speaker changes and non-speech information, and where it does not, translation is not an accessibility deliverable.

How to work through it

  1. 1Inventory every piece of audio and video, then mark each one prerecorded or live, and with sound or without. Those two answers decide which criteria apply, and getting them wrong is how teams audit the wrong rule.
  2. 2Play each one with the sound off. Write down what you lost. That list is what the captions have to carry.
  3. 3Play it again with the screen off. What you lose this time is what description has to carry, and if the soundtrack already said it, you owe nothing further.
  4. 4For prerecorded video with sound, check the A and AA obligation properly: captions at A, and at AA either audio description or a full text alternative for the visuals. A separate transcript is a real kindness and it is not what these two levels ask for.
  5. 5Read any transcript you do publish on its own, with the media closed. It has to stand up alone, including the outcomes of anything interactive inside the video, or it is not the alternative it claims to be.

How the levels build

Which criteria apply depends entirely on what kind of media you have, so read this as a routing table rather than a ladder. Level A wants a text alternative for audio-only recordings and a text or audio alternative for silent video. It also wants captions on prerecorded video with a soundtrack, plus either description or a full text alternative for visuals the soundtrack never states. Silent video needs no captions, because the captions rule is scoped to media where sound and picture run together. There is also an exception worth knowing at every level: where the media is itself a clearly labelled alternative to text already on the page, it is not asked to repeat the work. Level AA adds live captions for streamed events with sound, and audio description for prerecorded video. Level AAA is four more criteria, not two. Sign language interpretation, extended audio description for when the natural pauses are too short, a full media alternative for prerecorded synchronized media, and text for live audio-only. That last one is why a live podcast without a real-time transcript is an AAA gap rather than an AA failure.

Other Perceivable guidelines

Part of the Perceivable principle · browse by level: Level A · Level AA · or the full 55-rule library.

See how your site does against these rules.

An expert review plus a real blind screen-reader user, on up to 10 pages, every finding with its screenshot, criterion, and fix. $499, report in 5 business days.

Order your audit

Go somewhere useful

Find tools, resources and your workspace.

29 destinations