The tempting move is to take your existing WCAG 2.1 report, test the criteria 2.2 added, staple the results together and call the whole thing a 2.2 audit. It does not hold, and the reason people miss is not the new rules. It is that WCAG 2.2 was republished on December 12, 2024 carrying errata that modified four defined terms, and those terms appear inside ten of the Level A and AA criteria your 2.1 audit already covered.
So the old results were measured against wording that is not the wording in force now. That does not mean they are wrong, and it would be an overstatement to say ten criteria have quietly started failing. It means the evidence behind them was gathered against a slightly different question, which is a reason to re-check rather than a reason to worry.
Add the two ordinary reasons on top, that your site has moved since and that the criteria interlock in ways a partial pass cannot see, and you get three routes rather than one shortcut. Here is what actually changed, why the shortcut fails, and how to pick the route that matches what you need the report for.
What Changed Between 2.1 and 2.2
Nine success criteria were added and one was removed. The rest carry over, which W3C describes with a small hedge that is worth keeping in mind rather than reading past.
| Criterion | Level | What it asks for |
|---|---|---|
| 2.4.11 Focus Not Obscured (Minimum) | AA | When an item gets keyboard focus, it is at least partially visible |
| 2.4.12 Focus Not Obscured (Enhanced) | AAA | When an item gets keyboard focus, it is fully visible |
| 2.4.13 Focus Appearance | AAA | The focus indicator has sufficient size and contrast |
| 2.5.7 Dragging Movements | AA | Any action involving dragging has a simple pointer alternative |
| 2.5.8 Target Size (Minimum) | AA | Targets meet a minimum size or have enough spacing around them |
| 3.2.6 Consistent Help | A | Help sits in the same place when it appears on multiple pages |
| 3.3.7 Redundant Entry | A | The same information is not asked for twice in one process |
| 3.3.8 Accessible Authentication (Minimum) | AA | Signing in does not require solving, recalling or transcribing something |
| 3.3.9 Accessible Authentication (Enhanced) | AAA | The same, without the exceptions the minimum version allows |
Six of those land at Level A and AA, which is the level almost every obligation names. That is 2.4.11, 2.5.7, 2.5.8, 3.2.6, 3.3.7 and 3.3.8. Add them to the 50 criteria a 2.1 audit covers at A and AA, take away the removed one, and you get the 55 that a 2.2 audit tests. Our guide to WCAG versions and levels has the longer comparison, and the version history shows which criterion arrived when.
None of the six is a glance-and-tick check either, which matters if you are budgeting the delta as an afternoon. Target size has five separate exceptions and one of them is a spacing route rather than a size route. Consistent Help is scoped across a set of pages and to four named kinds of help mechanism. Redundant Entry carries three exceptions. Dragging Movements asks about every drag interaction in your product, which on a modern application is usually more of them than anybody remembers building.
The One That Left, and What Your Old Report Says About It
4.1.1 Parsing is obsolete and removed from WCAG 2.2. In the versions that still list it, 2.1 and 2.0, it is treated as always satisfied. So an old report carrying a 4.1.1 failure is carrying a finding that cannot exist under the version you are moving to, and cannot exist under the version it was written for either.
Do not just delete the row. The things that used to get filed there are usually real problems belonging to a different criterion. A duplicate id that makes a label point at the wrong field, say, or a broken relationship a screen reader cannot resolve. W3C's own position is that they should be written up under the criterion they actually break. Our page on what happened to 4.1.1 covers where those findings go and why scanners still raise them.
The Change Underneath the Rest
This is the part that is not in any comparison table, and it is the reason the shortcut fails on its own terms. The December 12, 2024 republication modified the definitions of single pointer, used in an unusual or restricted way, motion animation and programmatically determined. It also reformatted the definitions of several others and removed one that had become defunct.
Those are not obscure terms sitting in a glossary nobody opens. Definitions are what the criteria are built out of, and programmatically determined alone appears inside nine success criteria. Counting only Levels A and AA, and setting aside the new criterion that uses one of them, ten criteria your 2.1 audit already tested contain a term that has since been re-defined.
| Re-defined term | Criteria at A and AA that use it |
|---|---|
| programmatically determined | 1.3.1 Info and Relationships (A), 1.3.2 Meaningful Sequence (A), 1.3.5 Identify Input Purpose (AA), 2.4.4 Link Purpose (In Context) (A), 3.1.1 Language of Page (A), 3.1.2 Language of Parts (AA), 4.1.2 Name, Role, Value (A), 4.1.3 Status Messages (AA) |
| single pointer | 2.5.1 Pointer Gestures (A), 2.5.2 Pointer Cancellation (A), and 2.5.7 Dragging Movements (AA), which is new anyway |
| used in an unusual or restricted way | 3.1.3 Unusual Words, which is Level AAA and outside an A and AA audit |
| motion animation | 2.3.3 Animation from Interactions, which is Level AAA and outside an A and AA audit |
Be careful what you take from that. These were errata, so corrections and clarifications rather than new requirements, and a changed definition does not automatically flip an outcome. Most of those ten will come back the same way. What the table establishes is narrower and still decisive. The wording your evidence was measured against is not the wording in force, so the ten are the shortlist to re-check first, and a report that never re-opened them cannot honestly be labeled a 2.2 evaluation.
Two More Reasons the Delta-Only Pass Fails
The definitions are the reason nobody expects. Two ordinary ones sit behind it.
Your site moved. However old the 2.1 report is, something has shipped since, and the changes with the widest reach are the ones least likely to be remembered as changes. A component updated. A design token moved. A journey gained a step. Every one of those puts pages back in doubt that the old report speaks for, and none of them is a new criterion. Our page on how long a report stays true has the full list of change events and what each one costs.
Conformance does not add up in slices. A page conforms or it does not, and a process conforms only if every page in it does. Test six new criteria on a page and you have learned six things about that page. You have not refreshed the other 49, and you cannot combine an old pass on 49 with a new pass on six to produce a current statement about the page, because the two halves describe different days.
The Upgrade Worksheet
Five questions, answered honestly, and the answers point at one of three routes. This takes about fifteen minutes with the old report and your release notes open.
- When was the old evidence gathered? Use the evaluation date, not the delivery date. If it predates December 12, 2024, the definition question applies in full.
- What has shipped since? Specifically a template, a shared component, a design token, a journey step, a new layout or a third-party embed. Copy edits do not count here.
- Was the old report listed per occurrence or per criterion? A per-criterion report gives you one specimen per failure, so its counts were never a total and cannot be treated as one now.
- Who is going to read the result? Your own team tolerates a patched-together answer. A customer, a procurement questionnaire or a regulator will read the dates.
- Do you need Level AA only, or AAA as well? Two of the four re-defined terms only touch AAA criteria, so an AA-only target narrows the shortlist to the top two rows of the table above.
Three Routes, and What Each One Buys
| Route | What it covers | What you can honestly say | When it is the right buy |
|---|---|---|---|
| A. Delta pass only | The six new A and AA criteria, on the pages the old report sampled | That those six were evaluated on those pages on this date. Nothing about the other 49, which still carry the old date | Internal planning, where you want to know what the new rules cost you and nobody outside will read the answer |
| B. Delta plus targeted re-check | The six new criteria, the ten criteria containing a re-defined term, and anything touched by a change since the old evaluation | That the report was brought up to 2.2 on this date, with the old evidence retained for the criteria nothing has disturbed, and the reasoning written down | Most teams, most of the time. It is the honest middle and it is defensible because you can show the shortlist and why each item is on it |
| C. Full re-evaluation at 2.2 | All 55 criteria at A and AA, on a fresh sample, with new evidence throughout | That these pages were evaluated against WCAG 2.2 at this level on this date, full stop | Anything somebody outside your company reads, anything older than about a year, and any site that has had a redesign or a replatform since |
Route B is the one worth defending, because it is the only one that requires you to think. Its output is not just a report. It is a report plus a written note saying which criteria were re-tested and why, which criteria kept their earlier evidence, and what date each half carries. A reader who disagrees with your shortlist can argue with it, which is exactly what makes it a defensible position rather than a hopeful one.
If you go that way, one extra thing to steal from W3C's methodology. When an evaluation is re-run, it suggests keeping part of the original sample so the two sets of results can be compared. Then replacing another part, typically about half, so the new pass sees pages the old one never opened. A version upgrade is a re-run with a longer criterion list, and the same logic applies.
What to Do With the Old Report
Keep it, and keep it labeled. Three things it goes on doing.
- It is the comparison baseline. A 2.2 pass that can read the 2.1 findings can tell you what got fixed, what regressed and what was never touched, which is a document a fresh audit cannot produce at any price.
- It is still a true statement about its own date. Nothing retroactively invalidates evidence. What changes is what you may label it, and evaluated against WCAG 2.1 on this date stays accurate forever.
- Its component findings mostly still stand. A finding about a shared component is true about that component until somebody changes it, and version numbers do not change components.
What to stop doing is describing it as a 2.2 result, or leaving a badge or a statement up that names a version the evidence never covered. If you publish an accessibility statement, the version and the date in it are the two things a reader checks first, and they are the two easiest to leave stale.
One honest limit
Which version your obligation actually names is a separate question with a different answer in different places, and this page does not answer it. Some rules name 2.0, some name 2.1, some name 2.2 and some name none at all. Our guide to WCAG versions and levels covers how to work out which one applies to you, and the laws by country pages carry the specifics. Upgrading your evidence and being required to upgrade it are two different projects.
The Short Version
Six new rules at A and AA, one removed, and ten criteria you already tested that now sit on re-defined wording. Test the six, re-check the ten, re-check anything your team has changed since, and write down which half of the report carries which date. That is an upgrade somebody can check. Stapling six new results onto an old document is not.