What scanners do well
Automated tools are excellent at facts a machine can read straight out of the code. An image with no alt attribute. A page that declares no language. Two solid colors whose contrast ratio measures 2.8:1. They check every page in seconds and they never get bored on page four hundred, which is why every WCAGrules audit starts with a full automated pass. It is the fastest way to clear the mechanical failures before a person spends an hour on them.
Even here the boundary is narrower than it sounds. W3C's own contrast test rule records that passing it does not mean the criterion is met, because there is no settled method for judging text that sits over a photograph or a gradient. And a field labelled by nothing but a placeholder is handed an accessible name by that placeholder, so it passes the label check and fails a different rule the tool cannot reach. That is the mechanical reason a scanner report and an auditor disagree about the same page.
Where scanners stop
A scanner can confirm alt text exists, not that it says anything useful. It can find a button, not whether the checkout can actually be completed. Most of what makes a site usable is a judgment or a journey, and machines have neither. Across the 55 rules we audit, 52 of 55 need human review for some or all of their checks, and 25 of them are entirely invisible to tools.
What human testing actually means
It means a real blind screen-reader user, keyboard only, working through your actual tasks. Tab to the navigation, hear what the screen reader announces, find the product, reach the checkout, submit the form. Focus order, announcement quality, keyboard traps and missing feedback reveal themselves only in that journey, because each of them is a property of the route rather than of any single page. It is the difference between checking a page and using one, and it is why every audit pairs the automated pass with a hands-on human session. The full process is on How we test.
Rule by rule, for the 55 that laws name
This is not a judgement about any particular tool. It grades how far the rule itself can be decided by machine, which is a property of the rule rather than of the software pointed at it.
That distinction is also why our figure sits so far below the coverage percentages scanner vendors publish. They count defects found against defects present, and we count rules. Both numbers can be honest at once, and only one of them tells you what a clean report proves. The arithmetic is worked through further down this page, because the gap between the two looks like somebody is lying and nobody is.
Mostly detectable by automated tools.3 of 55
Partially detectable, so human review is needed to catch the rest.27 of 55
- 1.1.1 Non-text Content
- 1.2.2 Captions (Prerecorded)
- 1.3.1 Info and Relationships
- 1.3.4 Orientation
- 1.4.2 Audio Control
- 1.4.4 Resize Text
- 1.4.10 Reflow
- 1.4.11 Non-text Contrast
- 1.4.12 Text Spacing
- 2.1.1 Keyboard
- 2.1.2 No Keyboard Trap
- 2.2.1 Timing Adjustable
- 2.2.2 Pause, Stop, Hide
- 2.3.1 Three Flashes or Below Threshold
- 2.4.1 Bypass Blocks
- 2.4.2 Page Titled
- 2.4.4 Link Purpose (In Context)
- 2.4.6 Headings and Labels
- 2.4.7 Focus Visible
- 2.5.3 Label in Name
- 2.5.8 Target Size (Minimum)
- 3.1.2 Language of Parts
- 3.3.1 Error Identification
- 3.3.2 Labels or Instructions
- 3.3.8 Accessible Authentication (Minimum)
- 4.1.2 Name, Role, Value
- 4.1.3 Status Messages
Not detectable by automated tools at all. Human testing is required to check it.25 of 55
- 1.2.1 Audio-only and Video-only (Prerecorded)
- 1.2.3 Audio Description or Media Alternative (Prerecorded)
- 1.2.4 Captions (Live)
- 1.2.5 Audio Description (Prerecorded)
- 1.3.2 Meaningful Sequence
- 1.3.3 Sensory Characteristics
- 1.4.1 Use of Color
- 1.4.5 Images of Text
- 1.4.13 Content on Hover or Focus
- 2.1.4 Character Key Shortcuts
- 2.4.3 Focus Order
- 2.4.5 Multiple Ways
- 2.4.11 Focus Not Obscured (Minimum)
- 2.5.1 Pointer Gestures
- 2.5.2 Pointer Cancellation
- 2.5.4 Motion Actuation
- 2.5.7 Dragging Movements
- 3.2.1 On Focus
- 3.2.2 On Input
- 3.2.3 Consistent Navigation
- 3.2.4 Consistent Identification
- 3.2.6 Consistent Help
- 3.3.3 Error Suggestion
- 3.3.4 Error Prevention (Legal, Financial, Data)
- 3.3.7 Redundant Entry
The middle group is the one that misleads. Partly detectable means a tool finds some instances and cannot find the rest, so a clean report on those rules says a machine found nothing, not that there is nothing.
You do not have to take our grading for that, because the standards body publishes the same conclusion from the other side. W3C's own test rules each state what their result means, and there are 87 live ones. On 59 of them a clean pass is recorded as meaning the criterion still needs further testing. Exactly four rules in the entire set can clear a criterion outright, and the other 24 name no success criterion at all.
Then there is the coverage question underneath that one. Those 87 rules reach 37 of WCAG's 86 criteria between them, which leaves 49 criteria with no published test rule of any kind. Narrow it to the 55 rules an audit actually tests and 24 of those have none. So more than half the standard is out of reach of any conformance-tested automated rule, and that is before anybody argues about how good a particular scanner is.
Why Deque says 57% and we say 2.8%
Both figures are right. They are answering different questions, and once you see which, the gap stops looking like somebody exaggerating. Deque publishes 57% in the documentation for automated accessibility testing, and the figure comes from anonymised audit data across more than 13,000 pages and page states holding close to 300,000 issues between them. Ours is 10 fully machine-checkable techniques out of the 356 we graded, which is 2.8%.
The difference is the denominator. Deque counts issue instances. Picture a page holding 191 problems, where 190 are low-contrast elements and the last one is a heading that is really a pull quote. Catch the 190 easy ones and you have scored 99.5% while missing the only thing on the page a machine was never going to see. We count the standard. A technique counts once whether it fails on one element or on a thousand, which is the only way to answer the question an owner is actually asking, because nobody conforms to Level AA by instance.
A third dataset explains why instance counts run so high. WebAIM's annual scan of a million home pages found that 96% of everything detected falls into six categories, and all six are the easy ones. Low contrast, missing alt text, missing form labels, empty links, empty buttons and missing page language. A handful of high-frequency failure types dominate any instance count, so instance counts flatter automation by construction and criterion counts do not.
There is an independent benchmark as well, and it is the one worth knowing, because nobody in it was selling a tool. In 2017 the UK's Government Digital Service built a page carrying 143 known failures across 19 categories and ran ten tools at it. The best single tool found 41%, and that number counts prompts telling a human to go and look as finds. All ten together reached 71%. Nobody runs ten scanners, and even the aggregate missed almost a third of a page built specifically to be caught. If somebody quotes you the 71%, quote them the other half of the sentence.
Two things to carry away about the 57%. It is a vendor's study of its own tool, and the page it sits on has no date on it, so nobody can say how old the measurement is. And a third figure travels beside it in the same marketing, up to 80% with Deque's tools, which counts a person answering questions inside a guided test. That one is not an automation figure at all, whatever it is standing next to.
What our own free scan reaches, and what it misses
It would be an odd page that held every other tool to this standard and skipped its own, so here is ours. Our free scan drives real Chrome and runs all 90 supported automated rules: 63 mapped to WCAG and 27 additional best-practice checks that the report labels separately. The WCAG-mapped checks touch 21 success criteria, and 20 of those sit at Level A or AA. That leaves 35 of the 55 rules we audit untouched by the free scan entirely, along with every journey that runs across more than one page.
It also renders at a single desktop size, so nothing about how your site behaves on a narrow screen is tested at all, and it sees each page once in the state it arrives in, so a dialog nobody opened is a dialog nobody checked. The whole list is written out on the free scan page, in one place rather than a sentence at a time.
Run it anyway. A free scan finding a failure is a real failure worth fixing today, and it is the fastest way to clear the mechanical failures before anyone bills you an hour. Then read this page again and decide what the gap is worth to you.
Why this matters commercially
Overlay widgets and scan-only services do repair a real set of problems at runtime, and the practitioner reference document on overlays is equally explicit that full compliance cannot be reached with one. The Federal Trade Commission put the same point in a filed complaint, stating that no automated testing tool alone can determine whether a website meets accessibility standards and that manual human testing is required instead. So if a report never involved a person, it verified the part of the standard that was cheapest to check and left the rest unanswered. That is the part a customer walks into. The documented record on overlays sets out the rest.
How we graded every technique
We grade the 55 rules an audit tests for how far a machine can decide each one. A technique then takes the grade of the hardest rule it serves, because a technique is only settled when everything it serves is settled. Explicit overrides sit on top of that, for the cases where a technique is plainly more or less machine-checkable than its rules. A missing alt attribute is one. Judging whether alt text is any good is another, in the other direction. Detection here means finding the problem, and verifying a fix is correct usually takes a person either way.
Taking the hardest rule is a deliberate choice and it is worth saying which question it answers. It answers whether a scanner can settle a technique on its own. Grading on the easiest rule instead would answer whether a scanner could ever settle any part of it, which is a different and much friendlier question, and it is the one tool marketing tends to answer. G87 is the clearest case. It covers captions, it serves 1.2.2 and 1.2.4, and 1.2.4 is one no tool reaches. A scanner cannot settle captions, so G87 sits with the rules that need a person.
That leaves 76 techniques we do not grade at all. Most serve only Level AAA rules, which sit outside the 55 an audit tests. Nine are techniques W3C attaches to no rule at all. Two more are the interesting case: their hardest graded rule is one a machine settles, but they also serve a rule we never graded, so we cannot honestly call the technique settled. Scoring any of these as needing a person would turn a gap in our data into a verdict about tooling, so they are counted separately and every figure on this page is stated against the 356 that remain.
Two more things follow. The overrides are judgements, so a fair reader could grade a handful of these differently and land a few points either side of our numbers. And the grading has nothing to say about any particular scanner, because a tool that misses something the rule permits it to find is a tool problem, and this table is about the rules.