Audio Description (Prerecorded)
At Level AA the text alternative stops being an option. If a prerecorded video shows information the soundtrack does not carry, the soundtrack has to carry it. That is the entire difference between this rule and 1.2.3, and it is why the choice you made at Level A decides whether this one costs anything. Describe the visuals in audio back then and you are already finished. Write them out in text and you have a new job. The escape is the same as before. Where the audio already conveys everything important in the picture, no description is needed and the video passes as it is.
Why it matters
A described video is one a blind viewer can watch alongside everyone else, in the same room, at the same time, with no separate document open. That is the thing a transcript cannot do, and it is why the standard tightens here. Planned early, description costs nothing extra. A script that names what it shows records in the same session as a script that does not. Patched in later it means a second edit, a second voice session and a second file to host. Same fix, several times the price.
Who this rule protects
Blind and low-vision viewers depend on the soundtrack to carry everything the video communicates, because nothing on screen reaches them any other way. Viewers with cognitive disabilities who find it hard to interpret what is happening on screen get the same benefit from hearing it named.
How to check it yourself
- Listen to each video without watching the screen and confirm everything a sighted viewer receives also reaches a listener through the audio.
- Where the narration does not cover the visuals, look for a described version of the video or a player control that offers one.
- Check whether the dialogue has natural pauses. Where it does and nothing has been added to them, that is the documented failure. Where it does not, the answer is a rewritten soundtrack or an extended description that pauses the video to make room.
- For a video that is one person talking to camera against a background that does not change, look for a static text description. That is enough on its own at this level, as long as it carries the on-screen title, name strip and credits and is tied to the video in code.
- Confirm the description is in the same language as the video, or as the page around it. Either satisfies the rule.
- Do not add a captioning job for the description track. A description narrates what is already visible, so it carries no captioning obligation of its own.
Failures we see most often
- A demo says "as you can see, it is simple" over a screen recording nobody ever explains in words.
- On-screen statistics and titles appear throughout a video and are never spoken aloud.
- A full transcript is offered and no described audio exists, which cleared Level A and does not clear AA.
- The dialogue has long gaps and not one of them is used for description.
- A described version exists and the player offers no way to switch to it.
Who this one is for
Read from this rule's own note above, so the grouping and the note cannot disagree.
- Blind and screen reader userspeople who cannot see the screen
- Low visionpeople who can see the screen but not easily
- Cognitive and learningpeople for whom the difficulty is understanding, remembering, or staying with it
How this one is tested
We list 1 ACT rule against 1.2.5. Each one defines exactly what a checker looks at, which is how automated tools decide what to flag. Each one also checks a slice, so passing every rule here is not the same as meeting the criterion, and a rule can be proposed rather than approved or need a person to finish it. The note beside each says which.
- Video element visual content has strict accessible alternativeA tool finds candidates, you decide
How to fix it
- Write the description into the script. "On the cart page, choose Checkout, then Apple Pay" is a describing sentence, and it costs nothing extra to record.
- For one person talking to camera against an unchanging background, publish a static text description tied to the video in code, carrying the title card, the name strip and the credits. That is a documented route and it clears AA for the most common corporate video format there is. It does not stretch to a multi-speaker reel where an on-screen caption is the only thing naming who is talking.
- Where the pictures move faster than the pauses allow, produce an extended described version, which pauses the video to make room for the narration, and offer it alongside the original.
- Where a video already narrates everything on screen, write that down in your conformance notes rather than commissioning a track you do not need.
Step-by-step fix guides (11)
- H96: Add audio description tracks to videoAdvisory
- SM1: Add extended audio description in SMIL 1.0
- SM2: Add extended audio description in SMIL 2.0
- SM6: Add audio description in SMIL 1.0
- SM7: Add audio description in SMIL 2.0
- G8: Provide a version with extended audio descriptions
- G78: Offer a selectable audio description track
- G173: Offer a described version of every video
- G203: Describe talking-head video with static text
- G226: Narrate visual content within the soundtrack itself
Passes vs. fails
Passes
The narrator says "On the cart page, choose Checkout, then Apple Pay. Two taps and you are done", so the steps live inside the spoken script itself.
Fails
The narrator says "Here is how easy checkout is" over twenty silent seconds of screen recording, and never describes what the viewer is actually looking at.
In audits and lawsuits
AA is the level laws and contracts name when they name one, so a conformance claim usually goes wrong on details like this one. When we scope an audit we sort your video into three piles. The ones whose narration already covers the visuals, which pass untouched. The single-speaker ones shot against a background that does not change, which a static text description satisfies. And the ones that genuinely need a described cut, which are the only ones with real money attached. There is one numbered failure here, and it fires only where pauses in the dialogue existed and went unused. A wall-to-wall narration video still fails. It just fails against the rule itself rather than against a named technique.