Glossary · Accessibility term
Transcript
A transcript is the text version of audio or video, and two different things go by that one word. A basic transcript carries the speech and the non-speech sound you need to follow the content. A descriptive transcript adds the visual information as well. That second kind is what somebody who is both deaf and blind actually needs, because they reach your video as text on a braille display. The distinction decides what you have satisfied. For a podcast or any audio-only recording, a transcript is the Level A answer, and the only content excused is a recording that is itself a labelled alternative to text already on the page. For video with sound, a descriptive transcript is one of the two ways to meet the Level A description rule, the other being audio description itself. It stops being a route at Level AA, where audio description is the only answer, and it never replaces captions.
In practice
What goes into one is more than a text dump. Everything said, speaker identification where the content needs it and not where a single narrator makes it noise, and the non-speech sound that carries meaning. For video, add the text that appears on screen and the visuals somebody would otherwise miss. Then add headings and links so it can be navigated, because a wall of unbroken text is technically a transcript and is miserable to read. If anybody is selling you the search benefit, structured text on the page is the version of it worth having.
Build it from your captions. If a video already has a caption file, that is the speech and the sound done, and the work left is adding the visual information a caption track never carried. It is the cheapest route to a descriptive transcript and almost nobody takes it.
Put it on the same page as the media. A separate page is fine when the media itself is hosted elsewhere. An untagged PDF is the version to avoid, because it looks like compliance from the outside and hands a screen reader user a document with no structure to move through.
One nuance that saves real money, and one that costs it back. At Level A, prerecorded video offers a choice between audio description and a full media alternative, so a descriptive transcript is one of two routes rather than an extra, and commissioning description is usually the more expensive of the two. Level AA closes that choice. There the rule asks for audio description and nothing else will do, so a transcript that satisfied you at A does not carry you to AA.
Why it matters
Captions and transcripts are not interchangeable. Captions are timed, visual and on the screen. A transcript is a document you can read at your own pace or send to a braille display, which is why a descriptive transcript is what W3C names as the way to reach somebody who is both deaf and blind. Subtitles are where the words themselves get slippery. W3C says plainly that the two terms mean the same thing in different parts of the world, and then picks one convention for its own pages, using captions for the same language as the audio and subtitles for a translation. So an agency promising subtitles may be promising translation or may be promising captions, and the only way to know is to ask which one they meant.
Which kind you have
If your transcript would let somebody follow a video with the sound off and the screen off, it is a descriptive transcript. If it only carries what was said, it is a basic one, which is the whole job for a podcast and half the job for a video.
Where this shows up on the site
Related terms
- CaptionsCaptions are a synchronized text version of everything you would hear in a video, covering speech and the non-speech sound that matters as well.
- Audio descriptionAudio description is narration added to a video's soundtrack that describes the visual detail the soundtrack alone does not carry.
- Braille displayA braille display is a hardware device with cells whose pins rise and fall to form braille characters, refreshing as the reader moves through a page.
- Screen readerA screen reader is software that reads the screen out loud, or sends it to a braille display, for people who cannot see it.
- Sign language interpretationSign language interpretation is a signed version of spoken content, usually a video track running alongside the main one.
- Text to speechText to speech is the conversion of written text into a synthesized voice.
- CARTCART is live captioning written by a person rather than a machine, and the letters stand for Communication Access Realtime Translation.
Knowing the word is the easy part.
Find out where your own site stands. The free scan checks 10 pages in a real browser against all 90 supported automated rules, separates 27 best-practice checks from its WCAG findings, and names the rule behind every result.