Compliance
Arabic, Hebrew and Farsi PDFs: Why Right-to-Left Breaks Accessibility
A right-to-left document can look perfect on screen and be completely unusable with a screen reader. Not degraded — reversed, interleaved, or read as isolated characters. If your agency publishes in Arabic, Hebrew, Farsi, Urdu, Pashto or Dari, this is the failure mode to know about, because it does not exist in your English or Spanish files and it will not show up in a visual review.
Why is RTL harder than any other language?
Because two different things have to be right, and they can disagree.
Visual order is where marks appear on the page. Logical order is the sequence in which the content is meant to be read. In an English document these are essentially the same, so a PDF’s tag tree usually matches what you see.
In a right-to-left document they are not the same, and PDF is a format that describes positions, not meaning. A generator can place Arabic glyphs so they look correct while recording them in the file in an order that has nothing to do with reading. A sighted reader sees a clean page. A screen reader, which follows the tag tree, reads it in the recorded order — which may be backwards, or fragmented, or start in the middle.
Add bidirectional text and it compounds. Arabic prose containing an English product name, a Latin acronym, or — very commonly — a number, mixes directions inside a single line. Numbers in Arabic script run left to right even though the surrounding text runs right to left. That means a phone number, a case reference, or a dollar figure inside an Arabic sentence is a direction change mid-line, and it is a frequent point of corruption.
What specifically goes wrong
The failures we find most often, roughly in order of frequency:
1. No tag tree at all. The document was exported without structure. For RTL this is worse than for English, because with no tags a reader falls back on the raw content order, and for RTL that order is frequently meaningless.
2. Reversed reading order. Paragraphs or lines read in the wrong sequence. The document is coherent visually and incoherent aloud.
3. Broken character shaping. Arabic letters change form depending on position in a word. Some export paths store the presentation forms as separate characters, so text that looks right cannot be searched, copied, or read correctly — each letter is announced in isolation.
4. Numbers and Latin terms swallowed or relocated. The bidirectional boundaries are not marked, so an embedded reference number lands in the wrong place in the spoken output.
5. Tables mirrored incorrectly. In RTL, the first column is on the right. If the tag tree records headers left to right while the visual layout runs right to left, every cell is associated with the wrong header. A rate table read this way gives a resident actively wrong information.
6. The language is not declared. Even a correctly ordered Arabic document tagged as English gets English pronunciation rules applied to it, which produces nothing usable. This is the same failure we cover in translated documents and WCAG language of parts — RTL just makes the consequences louder.
Which WCAG criteria are we talking about?
Mainly four, all within WCAG 2.1 Level AA — the standard behind the ADA Title II rule, the HHS Section 504 rule, and Section 508 procurement:
- 1.3.2 Meaningful Sequence (A) — the reading order has to be programmatically correct. This is the one RTL documents fail most.
- 3.1.1 Language of Page (A) and 3.1.2 Language of Parts (AA) — declare the document language, and declare any passage in a different language.
- 1.3.1 Info and Relationships (A) — table headers, lists and headings have to be marked as what they are, with the correct associations.
Which regime applies to you and by when is in the ADA Title II deadline post.
How do I check a file I already published?
Three checks, in increasing order of usefulness:
Copy and paste. Select a paragraph of Arabic and paste it into a text editor. If the letters arrive in isolated forms, or the order is scrambled, the file has a character-level problem that no amount of tagging fixes. This takes ten seconds and finds problem #3 immediately.
Read the properties. Check the document language. An Arabic document declaring English is failure #6.
Listen to it. Turn on a screen reader with an Arabic voice and play thirty seconds. This is the only check that finds reading-order problems, and it is unambiguous — either the content makes sense or it does not.
Our PDF accessibility checker will report the structural issues in a specific file if you would rather begin from a written result.
How does it actually get fixed?
There is no button. The work is:
- Recover real text if the file is a scan or has broken shaping. Nothing else can happen first. For Arabic script this needs a person who reads the language to verify the output, because recognition errors in Arabic are not obvious to someone who does not.
- Rebuild the tag tree in logical order, not visual order — paragraph by paragraph, with bidirectional boundaries marked where direction changes.
- Set language at document and passage level, including the Latin-script inclusions.
- Fix table structure with correct header association for RTL orientation.
- Translate the alt text. Arabic alt text, not English alt text in an Arabic document.
- Verify with a screen reader in the target language. By someone who can tell whether what they hear is right.
That last step is the one that cannot be outsourced to a tool, and it is the reason RTL remediation is a translation competency as much as a technical one. An accessibility firm that does not read Arabic can tag a file correctly in structure and have no way to know the content is wrong. A translation firm can produce flawless Arabic and hand back a file with no tags at all.
Where Taika fits
We do both halves. Structure the source, translate with the structure intact, tag in logical order, set language per passage, translate the alt text, and verify with a screen reader in the language — one pass, one vendor, and someone who reads the language looking at the result.
Start at document accessibility, or accessibility and compliance services if you are scoping a full multilingual estate. Our languages page lists what we work in.
Publishing in Arabic, Farsi or Urdu and unsure whether it works? Request a quote and tell us the languages and roughly how many documents. We will tell you honestly what condition the files are in and what it takes to make them usable — starting with whether the text is real text at all.
Need this done right?
Taika Translations provides certified translation, interpretation, and accessibility services in 300+ languages.