Start by giving up on the number. There is no accuracy percentage anywhere in WCAG, no threshold to clear and no score to report, and any figure you have been quoted came from a captioning contract rather than from the standard. What exists instead is a definition, and the definition is a far better test than a percentage would be, because it tells you what to listen for.
So the deliverable changes shape. Not a grade out of a hundred. A list of moments, each one with a timestamp, what the audio says, what the caption says, and which kind of failure it is. That list is what a caption vendor can act on, what a rework budget can be built from, and what an argument about quality actually needs. A percentage settles nothing, because nobody can tell you which two percent went missing.
What the Standard Says a Caption Has to Carry
Captions are defined as a synchronized alternative for both speech and non-speech audio information needed to understand the content. The definition then spells out what non-speech means, and the list runs to sound effects, music, laughter, speaker identification and location. So a caption file that transcribes every word perfectly and names nobody is not a good caption file with a small gap in it. It is missing one of the two things the definition asks for.
The accuracy standard underneath that is looser than people fear and stricter where it counts, and the captions and transcripts guide works it through. Dialogue verbatim or in essence, plus the sounds that matter. Condensing a rambling sentence is allowed. Dropping the word that reverses the meaning is not. Hold both halves at once and the review gets much easier to run, because you stop marking paraphrase as error and start hearing the things that actually break comprehension.
Write the Acceptance Test Before You Watch Anything
Six failure classes cover nearly everything a library throws up, and naming them in advance is what turns a reviewer's opinion into a finding. Agree these with whoever will do the rework, because a class nobody will fix is a class not worth logging.
- Meaning changed. The caption says something the audio did not. A dropped negation, a misheard number, a product name turned into a different word. The most serious class and the least visible, because the caption reads perfectly.
- Speaker not identified. More than one voice and no way to tell which is which. It matters most in interviews, panels and anything with a question from the room.
- Non-speech sound missing. The alarm, the doorbell, the laugh, the music cue that tells you the tone changed. If the sound carried information, its absence is a gap in the content.
- Out of sync. The words are right and they arrive at the wrong time. Drift accumulates across long recordings, so check the end as well as the start.
- Unreadable on screen. Lines too long to take in, three lines where two would do, captions sitting over burned-in text they obscure.
- Wrong track entirely. A translation filed as captions. Subtitles in another language serve people who can hear, and they do not discharge the captions requirement.
That last one sounds like a filing error and it is a real finding with money attached. A platform whose metadata has one column for both will mark an English video as captioned because a Spanish subtitle track exists, and nothing in the library view tells the two apart. Check the language of the track against the language of the audio before you review a single second of content.
A Timestamped Excerpt Is the Deliverable
Here is the shape of what a reviewer hands back. The example below is illustrative rather than a real client's library, and it is written out in full because the format is the point. Each row is something somebody can go and fix without watching the video again.
| Time | What the audio says | What the caption says | Class and rework |
|---|---|---|---|
| 00:41 | Do not switch the isolator back on until the panel is closed | Do switch the isolator back on until the panel is closed | Meaning changed. Re-cut this caption cue against the audio |
| 02:15 | Second speaker begins, unnamed on screen | Continuous text, no speaker label | Speaker not identified. Add labels from the shot list for every voice |
| 04:02 | Alarm sounds for six seconds under the narration | Nothing | Non-speech sound missing. Add a bracketed cue |
| 07:30 to end | Narration continues | Captions run roughly two seconds early | Out of sync. Retime the last segment rather than the whole file |
Four rows, one video, and every one of them is a specific edit. Compare that against a single accuracy score somewhere in the nineties and you can see why the percentage is the weaker deliverable. Nobody can pay somebody to fix a missing few percent.
How to Choose Which Videos to Watch
There is no published sampling rate for media, so borrow the shape of one from the method auditors use on web pages and say plainly that you have borrowed it. That method builds a structured selection covering the things that differ, then adds a random tenth on top as a check on the structured half. If the random files turn up problems the structured ones did not, the structured selection was not representative and you go back and choose again.
That last move is the useful part, and it is what makes a sample defensible rather than convenient. It gives you a stopping rule. For a video library, the axes worth structuring across are these.
- Where the captions came from. Automatic, vendor-produced, in-house, or carried over from a live event. These fail differently and mixing them hides the pattern.
- How old the file is. Captioning practice in your organization changed at some point, and the sample should straddle that date.
- How many people speak. One presenter is the easy case. A panel is where speaker labels live or die.
- How specialist the vocabulary is. Product names, drug names, place names, anything an automatic system has never heard.
- Whether the video sits inside a process. A training module somebody has to pass is not the same risk as a marketing clip, whatever the runtime says.
- How long it runs. Sync drift and reviewer fatigue both scale with length, and short files hide both.
One more thing about the automatic ones, because it decides your budget. Automatically generated captions do not meet the requirement unless they are confirmed to be fully accurate, and confirming means somebody watched. Correcting an automatic file also takes real time for anybody who does not do it daily, which is why plenty of organizations end up outsourcing captions rather than fixing them in-house. Price the confirmation, not just the generation.
What a Sample Can and Cannot Say
A sample describes the library. It tells you the shape of what is wrong, which classes dominate, and roughly what the rework looks like. It does not certify the files nobody watched, and a report that lets a reader believe otherwise has done harm rather than good. Write the untested remainder into the report as its own line, in the same place every time, so nobody has to hunt for it.
Keep two other things out of the caption findings as well. Whether the player can be operated by keyboard and screen reader is a separate review with a separate owner, and it lives in accessible video players. And whether a video needs audio description at all is a different question from whether its captions are good, answered by closing your eyes rather than by muting the sound.
Live recordings deserve their own line in the plan. Captions produced live and carried straight into the published recording usually need a light editing pass for accuracy before that recording goes up, so an archive built from live sessions has a known and predictable rework queue rather than an unknown one.
One honest limit
We review media and we never produce it. A caption quality review hands you the timestamped list and the classes, and the re-cutting belongs to your team or your caption vendor. Language expertise and library scale both change what is possible, so tell us what the library looks like before anybody quotes. The matched door is the video and media review, with a blind screen-reader user at the play button.