There is a published method for auditing a website against WCAG, W3C wrote it, and it is free to read. It is called WCAG-EM, it runs in five steps, and its most valuable sentence is the one that costs an audit firm money to repeat. A WCAG conformance claim cannot be made for an entire website on the strength of a sample of its pages. The reason given is not that samples are sloppy. It is that there is always a chance a page nobody opened carries an error, and no sample size closes that gap.
That sentence is why this page exists, and it is worth holding on to before the rest. Almost everything sold as an accessibility audit is a sample. Ours is. The method W3C publishes is itself a sampling method, it says so, and then it tells you plainly what a sample can and cannot support afterwards. Everything else in the document is machinery for making a sample as honest as a sample can be.
That machinery is worth having, because it hands you a checklist to hold any audit proposal against. Five steps, each producing something written down. A baseline agreed before anybody tests anything. A structured sample plus a random tenth on top, then compared against each other. And a minimum report with four items in it that commercial audits routinely leave out.
What a Group Note Is, and Why That Matters Here
WCAG-EM is not a standard, and the difference is not a technicality. It is a W3C Group Note. Reading the status line at the top of any W3C document is the fastest way to tell which kind you are holding. A Recommendation creates requirements, and content can conform to it, which is what WCAG is. A Group Note explains, illustrates or advises, and nothing conforms to it.
The note says as much about itself, in about the plainest terms W3C uses. It is endorsed by the working group that wrote it, and it is not endorsed by W3C itself nor its Members. WCAG2ICT carries the identical disclaimer for the identical reason, which is worth knowing if a contract ever names either document as though it were a standard you could be measured against.
None of that makes it weak. It makes it the wrong shape for a compliance clause and exactly the right shape for a quality test. Nobody can require you to conform to WCAG-EM. Anybody can hold an audit report up against it and see what is missing. For a buyer, the second of those is far more useful than the first.
One version note, because it changes what you will find elsewhere. The current edition arrived in July 2026 and it is a rewrite rather than a reprint. The 2014 original was written about websites and web pages. This one is written about digital products and views, which stretches it to cover mobile apps, kiosks, e-books and documents alongside websites. Guidance that describes WCAG-EM in terms of websites only is describing the old one.
The Five Steps, and What Each One Produces
Each step carries a numbered requirement of its own, and each produces something that has to be written down. The written record is the point. An evaluation nobody documented cannot be checked by anyone, including the person who ran it.
| Step | What you do | What goes in the report |
|---|---|---|
| 1. Define the scope | Say exactly what is being evaluated, pick the target level, and agree the accessibility support baseline with whoever is paying | The scope, the level, the baseline, and any extra requirements agreed |
| 2. Explore the product | Find the common views, the tasks that matter, the kinds of page, the technologies the site depends on, and the pages that matter specially to disabled visitors | The technologies relied upon. The four exploration lists are optional |
| 3. Select the sample | Build a structured sample, add a random set at a tenth of its size, then pull in every step of every process the sample touches | The structured sample, the random sample and how it was picked, and the processes |
| 4. Evaluate | Test every sample against all five conformance requirements at the target level, then compare the structured set against the random one | Findings from the initial samples, from the processes, and from the comparison |
| 5. Report | Document the outcome of the four steps above | The report. A public statement, a score and a machine-readable version are all optional |
The order can vary, and evaluators are expected to go backwards. Something learned at step four can send you back to step three, and the method says so outright. That is a more honest description of how an audit actually runs than a straight line would be, and it is the reason the sample is described as being refined during evaluation rather than fixed before it.
The Baseline That Explains Why Two Honest Audits Disagree
Here is the part almost nobody publishes, and almost every disagreement traces back to it. Before any testing starts, the method requires the evaluator to write down which browsers, assistive technologies and other user agents the site is expected to work with. That list is the accessibility support baseline, and it is settled with whoever commissioned the audit rather than chosen quietly by whoever runs it.
WCAG needs this because accessibility supported is one of the five conformance requirements, and the standard deliberately refuses to define it for you. There is no list of blessed software anywhere in WCAG. What counts as supported depends on the site's audience, its language, the technologies it uses, and which user agents people in that place actually run. So the baseline is not paperwork. It is the yardstick, and two audits holding different yardsticks will return different numbers for the same site while both of them are right.
A worked example makes it concrete. A custom combobox may behave correctly in one screen reader and browser pair and fall apart in another. Test it against a baseline that names the second pair and you have a finding. Test it against a baseline that does not, and the same widget passes. Neither tester was careless. They measured different things, and unless both baselines were published, the argument that follows has nothing in it to settle the question.
Three details follow that a buyer can use directly. The baseline can be widened partway through, and when it is, the extra tools get added to the record rather than left off it. For an intranet, where you know every machine in the building, a narrow baseline is honest rather than lazy. And your own developers may already keep a supported-browser list from the build, which the method names as a fair starting point rather than making you invent one.
Ask for the baseline before you sign anything
The baseline is a required part of the report and a required part of any statement published afterwards. So a proposal that cannot tell you which screen reader and browser pairs it will test against is proposing to leave out something the method treats as mandatory. It is a fair question on a first call, it takes ten seconds to ask, and the answer tells you most of what you need to know about the rest of the work.
How the Sample Gets Drawn, and Why a Random Tenth Goes On Top
Sampling is where an evaluation usually goes soft, and it is the part the method works hardest on. There are two sample sets rather than one, and the second one exists to check the first.
The structured set is chosen on purpose. It has to reflect the common views, the essential functionality, the different kinds of page, the technologies the site relies on, and any pages that matter specially to disabled visitors. That last group is the one teams forget, and W3C lists it separately for exactly that reason. The accessibility statement itself, the help pages, the pages explaining settings and shortcuts, contact and support, and anything handling sign-in, personal data or payment.
Then comes the random set, and its size is fixed. It is ten percent of the structured set, and it is added on top rather than carved out of it. Eighty structured pages means eight random ones and eighty-eight in total. They are drawn from anywhere inside the scope, by a method that has to be recorded, and any page already in the structured set gets swapped for another.
The reason for the random set is the good part, and it is not extra coverage. It is a test of your own sampling, with a defined failure condition attached. If the random pages turn up kinds of content or kinds of finding that the chosen pages did not, then the chosen pages were not representative, and the method sends you back to step three to select more. You repeat that loop until the random slice stops surprising you.
So a report showing a structured sample, a random sample and a comparison between them is telling you something a bare page list cannot. It is showing its working on the one decision that most affects the result. A vendor who picked ten pages and cannot say how has asked you to trust the choice rather than check it.
Complete Processes Come In Whole, Branches and All
A checkout is not a page and cannot be sampled like one. Complete processes are one of the five conformance requirements, and the rule there is unforgiving. Every page in the process conforms at the level claimed, or none of them does, including the pages that would sail through on their own. So the method pulls whole processes into the sample rather than pages from inside them.
The instruction is specific and it is worth copying. For any sampled page sitting inside a process, find the start of that process and include it. Record the default route to completion, which assumes no mistakes and no unusual choices, so a checkout paid with the stored card and the saved address. Then record the branch routes that people commonly take and that matter to finishing, such as adding a new delivery address, and include those too.
One practical detail there is worth stealing for your own testing whatever method you use. A URL is usually not enough to identify a page inside a process, so the method asks you to write down the actions needed to get from one step to the next. Fill in name and address, press Submit. Without that, nobody can reproduce your finding, and that includes you in three months when the developer disputes it.
The Sentence That Costs Us a Selling Point
Now the part we would quietly leave out if we were being commercial about it. WCAG-EM states that a WCAG conformance claim cannot be made for an entire website on the strength of an evaluation of a selected sub-set of its pages and functionality. The reason it gives is stronger than the one people assume. Not that a sample is unrepresentative, but that it is always possible an unexamined page carries an error, which is a problem no sample size solves.
It then goes further, and the second half is the part nobody quotes. Because most uses of the method evaluate a sample, most uses of the method do not produce a WCAG conformance claim at all. That is W3C describing the ordinary outcome of its own procedure, in its own document, without apology.
This lands on top of something conformance already establishes. WCAG defines conformance for a web page and never for a website. A claim can cover many pages, and it is assembled out of per-page conformance, so it holds only where every page inside it was itself evaluated or produced by a process that guarantees each one passes. A ten-page sample does not meet that description, and neither does a two-hundred-page one.
We sample, and so does everybody honest. What you get from us, and what you should expect from anyone, is a true account of the pages tested plus an unusually good guide to the templates sitting behind them. What nobody can hand you is a certificate covering the URLs they never opened. Our pricing page says the same thing in nearly the same words, and our audit method lists it first among the three things an audit here does not cover.
What the Report Has to Contain, and the Four Parts Usually Missing
The method sets a minimum report, and it is short enough to check a proposal against in a couple of minutes. Required, not optional, are all of the following.
- Who and when. The evaluator's name, the commissioner's name, and the date the evaluation was completed or the period it ran over.
- The scope, the level and the baseline. What was evaluated, which conformance level was targeted, the accessibility support baseline, and any extra requirements that were agreed.
- The technologies relied upon. Which technologies the site depends on to conform, identified during the exploration step.
- The samples. The structured sample, the random sample together with the method used to select it, and the complete processes that were included.
- The findings. Outcomes from the initial samples, from the processes, and from the comparison between the structured and random sets.
Four of those are the ones to look for, because they are the four a commercial audit report most often leaves out. The accessibility support baseline. The technologies relied upon. The random sample together with how it was picked. And the complete processes. None of the four is optional, and a report missing them is incomplete against the method whatever else it happens to contain.
Two smaller requirements are worth knowing before you read your next report. It should carry at least one example for every conformance requirement and every success criterion not met, so a bare count of failures does not satisfy it. And the documentation does not have to be public, because confidentiality is normally the buyer's call. That last one cuts against a vendor citing confidentiality as the reason they cannot show you a sample report. It is their choice, not a rule they are stuck with.
An Evaluation Statement Is Not a Conformance Claim
The method has its own answer to the problem its sampling sentence creates, and the answer is a different document with a different name. An evaluation statement describes the outcome of the evaluation. It is not a WCAG conformance claim, and it does not pretend to be one.
Three conditions hold before you can publish one. Every non-optional requirement of the method was satisfied. Every sample that was evaluated met the target level. And the product owner commits to keeping the statement accurate, which is the condition doing the most quiet work, because a statement about a site that has changed since is worth nothing to the person reading it.
Six things go in it, and the list is short enough to audit at a glance.
- The date the statement was issued.
- The guidelines title, version and address, written out in full.
- The conformance level that was evaluated, so A, AA or AAA.
- The definition of what was evaluated, carried over from the scope.
- The technologies relied upon.
- The accessibility support baseline.
It handles the partial case too. Where only part of the product reached the level, the statement names the areas that did not and gives the reason, and the method allows exactly two reasons. Third-party content, or a lack of accessibility support for languages. And the statement itself has to be published in an accessible format, which is a requirement that catches more accessibility statements than anyone would like.
One more thing the method says out loud, and it is unusual for a standards body. On aggregated scores, it states that no single metric currently offers the reliability, accuracy and practicality needed, that scores can mislead, and that this is among the reasons WCAG provides no rating scheme of its own. If you have ever been handed an accessibility score out of a hundred, that is the document to read next.
How to Read an Audit Proposal Against This
You do not need to run the method to use it. Five questions, asked before you sign, and each one is answerable in a sentence by anybody who intends to do the work properly.
- Which browsers and assistive technologies will you test against? That is the accessibility support baseline. A vendor with no answer is testing against whatever is installed on their laptop.
- How will you choose the pages, and will any of them be chosen at random? The structured plus random split is the method's own answer to cherry-picking, and it costs a vendor almost nothing to adopt.
- Will you test whole processes end to end, including the branches? Checkout, sign-up, booking. A page-by-page audit that stops at the cart is not testing the thing that earns the money.
- Will the report name the technologies the site relies on? It is a required item and it is the one that tells you whether anybody looked at how the site is built rather than only at how it renders.
- What exactly will the report let me claim afterwards? The honest answer is a true account of what was tested. Anybody promising a site-wide certificate from a sample is promising something the method says cannot be produced.
One honest limit on this page
Following WCAG-EM does not make an audit correct, and no method can. It makes the audit checkable, which is a different and more modest thing. A tester who applies it perfectly and misreads a success criterion produces a well-documented wrong answer. What the method removes is the ability to hide a weak sample, an undeclared baseline and a missing process behind a tidy-looking findings list. That is worth having, and it is not the same as a guarantee. Our own audit method borrows its structure and says where it departs from it.