There is no official list of browsers and screen readers an audit has to cover. W3C says so on the record, in a note attached to the definition that would have carried the list if one existed. The working group and W3C do not specify which or how much assistive technology support a use of a technology needs before it counts as supported. So a supplier who hands you a required matrix has written that matrix themselves, and the honest ones will tell you so.
What exists in place of a list is better for a buyer, because you get a say in it. It is called the accessibility support baseline, it is agreed before any testing starts, and the methodology makes it a required part of both the evaluation and the report. Asking for it on a first call takes ten seconds and tells you most of what you need to know about the rest of the proposal.
The short version
No standard names combinations. The evaluator and the buyer agree a baseline before testing, and it goes in the report. Your own developers' supported-browser list is a legitimate starting point. A closed network may narrow it. It can be widened partway through, and when it is, the extra tools get added to the record rather than left off. And the combinations you tested are optional in a conformance claim, which is exactly why a report that omits them is hard to check.
Why Support Is a Property of a Combination Rather Than of Your Page
The reason no list exists starts with how the standard defines support, and the definition has two halves. The first says a use of a technology counts as supported when it has been tested for interoperability with users' assistive technology in the human languages of the content. The second says accessibility-supported user agents have to be available to users, by one of four named routes.
Read the first half slowly, because it contains something people miss. Support is established by testing, against actual assistive technology, in the language the content is written in. It is not a property your markup has. It is a property of markup plus browser plus assistive technology plus language, all four at once. That is why the same custom combobox can behave correctly in one pairing and fall apart in another, with nobody having made a mistake.
A further note makes the point sharper still. When a technology is used in an accessibility-supported way, that does not imply the whole technology or all of its uses are supported. Most technologies, the standard adds, including HTML, lack support for at least one feature or use. So there is no version of this where you check a box for a technology once and move on.
What a Baseline Is, and Who Gets to Set It
W3C's evaluation methodology turns all of that into a requirement with a number on it. Before any testing happens, the evaluator has to define the browsers, assistive technologies and other user agents that the product's features are expected to be supported on. That definition is the baseline, and the methodology says it is carried out in consultation with whoever commissioned the evaluation.
That word, consultation, is the part worth holding on to. The baseline is not chosen quietly by whoever runs the tests and revealed afterwards in an appendix. It is a decision you are entitled to be in, and it decides the result, because two competent auditors holding different baselines will return different findings on the same page while both of them are right.
You may already have most of it written down. The methodology names the product owner's or developer's existing list of supported combinations as a fair starting point, and adds that it may need updating, for example to reflect how the product behaves with more current browsers. If your engineering team keeps a browser support policy, that document is where this conversation starts rather than a blank page.
Building the Matrix, One Column at a Time
A workable baseline has four columns rather than two, and the ones people forget are the third and fourth. Operating system belongs in it because screen reader behaviour is partly the platform's. Version belongs in it because a finding without a version is not reproducible six months later. Build the grid below for your own product, and the rows will come from who your users actually are rather than from a template.
| Operating system | Browser | Assistive technology or setting | Why this row is in your baseline |
|---|---|---|---|
| Windows | Chrome, current | NVDA, current | A free screen reader anyone can install, so nothing about it is gated behind a licence |
| Windows | Chrome or Edge, current | JAWS, current | A licensed screen reader, so include it when your audience or workforce is issued one |
| macOS | Safari, current | VoiceOver | The pairing VoiceOver is built and tuned for |
| Windows | Chrome, current | Browser zoom to 400 percent | Reflow and text spacing are AA criteria and are tested by setting, not by screen reader |
| Windows | Chrome or Edge, current | Forced colors mode | A rendering mode your CSS can break independently of everything above |
Two of those rows are not screen readers at all, and leaving them out is a common way a matrix ends up narrower than the standard it is being tested against. Reflow, text spacing, contrast in a forced palette and target size are all judged with settings rather than with speech. A baseline listing only screen readers has quietly dropped several Level AA criteria from the environment they need.
Tested Combinations Are Not Universal Compatibility
Here is the sentence that keeps a report honest, and it is worth writing into your own accessibility statement as well. A pass on the combinations in the baseline is evidence about those combinations. It is not a statement that the product works with every assistive technology in existence. No audit at any price produces that statement, because the set of combinations is open ended and moves with every release of every product in it.
The standard itself is careful here in a way that trips people up. The combinations used for testing are listed among the optional components of a conformance claim, not the required ones. So a claim can be technically complete while telling you nothing about what it was tested on. The methodology closes that hole from its own side by making the baseline a required part of its report, which is a good reason to ask for an evaluation report rather than a claim.
A narrow baseline is not dishonest, and a wide one is not automatically better. What makes either of them honest is publishing it. A proposal that cannot tell you which pairings it will use is proposing to leave out something the methodology treats as mandatory, and that is a fair thing to notice before you sign.
Narrowing It for a Closed Network, and Widening It Mid-Audit
Two adjustments are provided for, and both are worth knowing because they come up in most real engagements. The first is narrowing. For a product in a closed network such as an intranet, where both the users and the machines they use are known, the methodology says the baseline may be limited to what is used inside that network. WCAG's own support definition allows the same thing, treating a closed environment as one of the four routes by which a user agent counts as available.
In every other case the methodology points the other way, saying the baseline is ideally broad enough to cover the majority of current user agents used by people with disabilities in the relevant geographic region and language community. That last phrase does real work if you ship in more than one language, and whether every language version needs testing follows it through.
The second adjustment is widening, and it happens more often than the neat version of this process suggests. An evaluator who meets content the baseline did not anticipate is free to reach for another tool, and the methodology says the baseline is then extended with whatever was used. That is the right behaviour, and the tell for whether it was done properly is whether the extra tool appears in the report. A baseline that grew during the work and shrank again before delivery is the one to worry about.
What to Ask For, and What We Run
Bring three things to the conversation and the matrix builds itself. Who your users are, including anything you know about the assistive technology your audience or your workforce actually runs. What your engineering team already supports, from the browser policy they keep for other reasons. And which of your journeys carries the most consequence if it breaks, because that is where the widest pairing coverage should go.
On our side, the report names the exact screen reader, browser and version behind every finding, so a developer can reproduce it and a sceptic can check it. Which pairings a given audit runs is part of what gets agreed when we confirm scope, and pricing sets out what each package includes. If your audience sits mostly on one screen reader, say so at that point, because it changes what gets set up rather than what it costs.
One limit belongs here rather than in a footnote. Our sessions are desktop-based, so a mobile screen reader on a real handset is outside this engagement, and nothing here should be read as a promise of native-app or hardware testing. If the phone experience is the one you are worried about, mobile accessibility explains what desktop evidence does and does not reach.
Where This Page Stops
A baseline is a technical agreement, and it is not a legal one. Some regulations and many contracts name a standard and a level without saying a word about combinations, and a few procurement documents name specific assistive technology that your baseline then has to cover whatever your own users run. Reading which of those apply to you, and what a particular clause obliges you to test, is a lawyer's job rather than an auditor's.
What an auditor can do is make the technical half unambiguous, so the legal conversation has something exact to work from. Agree the baseline in writing, keep it in the report, and update it when the product or the audience moves.