Skip to main content
WCAGrules
Quick navigation

Glossary · Accessibility term

Tagged PDF

A tagged PDF carries an invisible structure layer naming each part of the document. This is a heading, this is a list, this is a table header, this is decoration. The layer is stored separately from the marks on the page, with pointers between the two. Which is why the order of the tags can be right while the order the ink was painted in is not. Tagging is optional in the PDF standard itself. A file with no tags is a completely valid PDF, and the flag recording that a document follows the tagging conventions defaults to false. So a tool that emits an untagged file has broken no PDF rule. That is the whole reason accessibility has to be added on purpose. Where anybody has measured, the picture is poor. A PDF tooling vendor ran an automated check across nearly 70,000 PDFs on German public sector websites in 2026, and 37.4% carried no tags at all while fewer than one in ten passed.

In practice

The tag tree is separate from what you see. Underneath, a page of PDF is instructions for painting marks at coordinates, and it has no idea which of those marks form a heading. The tags are a parallel structure that says so, hanging off a single root in the document catalogue. That root is the thing a checker looks for when it tells you a file is untagged.

Tags come from the source document when you export properly. Word and InDesign both produce them with tagging switched on, and Word's export dialog carries a checkbox for document structure tags. Printing to PDF is the route that usually loses them, flattening everything to ink, though not always any more, since Chrome has produced tagged output from its own Save as PDF destination for several versions now. Which is why the answer is to check the file rather than to trust the workflow. Where the tags did go, a screen reader user can still reach the words and has lost the headings, the alt text and everything else the tags were carrying.

A structure tree is not by itself enough, because the standard sets rules about the page content as well. All text has to be convertible to Unicode, word breaks have to be marked explicitly, and real content has to be told apart from artifacts of layout such as running headers and page numbers. A file that fails those can carry a complete tag tree and still be read aloud as a run of nonsense.

Why it matters

Without tags a screen reader gets the words in whatever order the file happens to store them, with no headings to navigate by and no way to tell a table from a paragraph. It is the difference between a document somebody can move around in and one they have to sit through from the top. The optionality is what makes this everybody's problem rather than one vendor's bug. Nothing in the PDF standard obliges a writer to produce tags, the default is off, and the tools most people use will happily hand you a valid file with nothing in it for assistive technology to navigate by. Some readers will try to infer a structure from the layout. Do not plan around that, because what they infer is a guess and it varies by product.

How bad it gets in one sector

A peer-reviewed study checked nearly 20,000 scholarly PDFs against six accessibility criteria. Only 12.6% passed the tagging one, and 74.9% passed none of the six. That is a narrow population rather than the whole web, and an automated check rather than a full evaluation. It is also the population most people mean when they say academic publishing has a PDF problem.

Where this shows up on the site

Related terms

Knowing the word is the easy part.

Find out where your own site stands. The free scan checks 10 pages in a real browser against all 90 supported automated rules, separates 27 best-practice checks from its WCAG findings, and names the rule behind every result.

Run the free scan

Go somewhere useful

Find tools, resources and your workspace.

29 destinations