Compliance
What Is Actually Inside a Remediated PDF? Tags, Reading Order and Roles
You approved a quote to remediate two hundred documents. The files come back looking exactly like the files you sent — same pages, same layout, same fonts, no visible difference whatsoever. It is a fair moment to ask what you paid for.
The answer is that everything which changed is invisible by design. A remediated PDF looks identical because the visual layer was never the problem. What was added is a parallel description of the document’s meaning, which assistive technology reads and your eyes never see. Understanding that layer is what lets you tell genuine remediation from a file that had a checkbox ticked.
Why does a PDF need this at all?
PDF was built to preserve appearance. At its core it is a set of instructions for placing marks at coordinates: put this glyph here, draw this line there. That is why a PDF looks the same everywhere, and it is exactly why it is hostile to assistive technology.
Nothing in that instruction set says a line of text is a heading. Nothing says twelve numbers arranged in a grid are a table with headers, or that a column of short lines is a list, or that the page number at the bottom is decoration rather than content. A sighted reader infers all of it instantly from size, weight and position. A screen reader has no eyes to infer with.
Tagging supplies what the visual layer leaves out: a structure tree, sitting alongside the page content, that states what each piece of content is and what order it belongs in.
What is in the structure tree?
A properly tagged document contains a hierarchy of standard structure elements. The ones that carry most of the weight:
<Document>— the root container for the whole file.<H1>through<H6>— real headings in a real hierarchy. This is what lets a screen-reader user jump between sections instead of reading 40 pages start to finish, and it is usually the single biggest usability gain in the whole job.<P>— paragraphs, each one a complete unit rather than a fragment per line.<L>,<LI>,<Lbl>,<LBody>— a list, its items, each item’s bullet or number, and the item’s actual content. This is why a real list is announced as “list of seven items” rather than as seven unrelated paragraphs beginning with a dash.<Table>,<TR>,<TH>,<TD>— table, row, header cell, data cell. TheScopeattribute on each<TH>says whether it heads a row or a column; complex tables useHeadersandIDto tie each cell to the right headers. Get this right and a cell is announced as “Residential, Second quarter, $84”. Get it wrong and it is announced as “84”.<Figure>with an/Altentry — an image plus its text alternative.<Link>— a link whose clickable region and its text are associated, so the link is announced with its purpose rather than as a bare URL.<Artifact>— the opposite of content. Page furniture: running headers, footers, page numbers, decorative rules, background watermarks. Marking these as artifacts is what stops a screen reader announcing “Page 14 of 92, County of Example, Annual Report” between every single paragraph.
Beyond the tree, several document-level properties matter just as much:
- Language must be declared, and any passage in another language tagged with its own — WCAG 3.1.1 and 3.1.2.
- The document title must be set in the metadata, and the file must be configured to display that title rather than the filename, so a user with nine tabs open can tell them apart.
- Tab order must follow the structure, so keyboard navigation and the reading order agree.
- Bookmarks for any long document, giving a navigable outline.
- Form fields, if there are any, each need an accessible name and a tooltip, and their tab order must match the visual order.
Reading order is the part that goes wrong
Tags say what things are. Reading order says in what sequence. They are separate, and a file can have perfectly good tags in a nonsensical order.
The order comes from the structure tree, not from position on the page — which means it can disagree with what you see. The classic failures:
Multi-column layouts read across instead of down. A two-column newsletter is read as one line from the left column, then one from the right, alternating into nonsense.
Sidebars and pull quotes interrupt. A callout box lands mid-sentence in the main text because it was placed there in the content stream.
Captions arrive before their figures, or a table’s footnote is read before the table.
Tables read column-first. Every value is announced against the wrong header — the worst case of all, because the output is not obviously broken. It is fluent, confident and wrong. A resident reading a fee schedule that way gets a real number attached to the wrong service.
Reading order is also where right-to-left documents fail hardest, because visual and logical order genuinely differ in Arabic, Hebrew, Farsi and Urdu. That is a large enough problem to have its own post: Arabic and right-to-left PDF accessibility.
How can I tell if a file was really remediated?
You do not need specialist software for a first pass. Four checks, in about two minutes:
1. Select all the text and copy it into a plain text editor. What you get is close to what an assistive technology encounters. If columns interleave, if headings land in odd places, if whole sections are missing, the structure is wrong. If nothing can be selected at all, the file is a scanned image and needs OCR before anything else can happen.
2. Check the document properties. Is there a real title, or is it “Microsoft Word - final_v3_REVISED.docx”? Is the language set, and set to the language actually in the file?
3. Look for the tag tree. Any full PDF editor will show it. An empty tree, or a tree consisting of one <P> per page, is an untagged file regardless of what the delivery note said.
4. Listen to a page. Turn on a screen reader and play sixty seconds of a table or a two-column section. This finds reading-order problems immediately and unambiguously.
Our PDF accessibility checker will report the structural facts for a specific file — tags, language, title, and the signals that indicate reading-order trouble — if you would rather start from something written down.
The tell-tale signs of a cosmetic pass
Some suppliers deliver files that pass an automated checker and remain unusable. What that looks like:
- Autotag and ship. Software applies a tag tree in seconds and nobody checks it. Everything is a
<P>, tables are unstructured, reading order is whatever the content stream happened to be. - Alt text that describes nothing. “image”, “chart”, “logo”, or the filename. Present, so the checker is satisfied; useless, so the reader is not.
- Headings that are only bold. Large bold text tagged
<P>. It looks like a heading and is not one, so the document has no navigable structure at all. - Nothing marked as an artifact. Headers and footers read on every page.
- Language left as English on a translated file. Very common, and it makes the whole document unreadable by ear.
Ask for the checker report, and run the copy-paste test yourself on the two most complex documents in the batch — the ones with tables. That is where corners get cut, because that is where the manual work is.
The multilingual dimension
If your documents exist in several languages, each translated file needs its own full structure. Tags do not survive translation intact: text expands or contracts, tables reflow, page breaks move, and a document translated after remediation frequently arrives with its tag tree scrambled. Translated alt text has to be written, not copied — an English description left inside a Somali document fails in both directions — and a passage quoted in another language needs its own language tag, which is WCAG 3.1.2.
The efficient order is to translate first and remediate the final files, or to have the same team do both. Remediating an English file and then translating it means paying for the structural work twice and hoping it survives.
Want to know what state your documents are in? Send us a representative sample — the messiest ones, with the tables — and request a quote. Tell us the document count, the page counts, and every language each one exists in, and we will tell you which files need tagging, which need reading-order repair, and which need to go back to source. Our document accessibility page covers how the work is done.
Need this done right?
Taika Translations provides certified translation, interpretation, and accessibility services in 300+ languages.