Guides
Accessible Video: Captions, Transcripts, and Audio Description
Video is where accessibility programs most often stall. Websites get audited, documents get remediated, and the media library sits untouched because nobody is sure what is actually required.
The requirement is more specific than “add captions,” and auto-generated captions do not satisfy it.
Three separate requirements, not one
WCAG treats video as needing up to three distinct things, depending on the content and the level you are targeting.
| Requirement | Serves | WCAG level |
|---|---|---|
| Captions | Deaf and hard-of-hearing users | A (prerecorded), AA (live) |
| Audio description | Blind and low-vision users | A/AA depending on criterion |
| Transcript | Everyone; required for audio-only | A |
These serve different people. Captions do nothing for a blind user. Audio description does nothing for a deaf user. Doing one and calling video “accessible” leaves an entire audience out, which is the most common mistake we see.
Captions
Captions carry the spoken content and the non-speech audio that matters — a door slamming, a siren, laughter, the fact that music is playing. That last part is what separates captions from subtitles. Subtitles assume you can hear and translate the dialogue; captions assume you cannot hear at all.
Requirements that get missed:
- Speaker identification when more than one person talks.
- Non-speech sounds relevant to understanding.
- Accurate punctuation — it carries tone and sentence boundaries.
- Reading-rate-appropriate timing, synchronized to the audio.
Audio description
A second audio track describing essential visual information that is not conveyed by the existing audio: on-screen text, a demonstration, a chart, someone silently pointing at something.
The test is simple: turn off your monitor and listen. If you lose information, you need audio description. A talking-head video usually does not. A training video demonstrating a procedure almost certainly does.
Transcripts
A text version of the full content. For audio-only content, a transcript is the requirement. For video, a transcript is not strictly required if you have captions and description — but it is inexpensive, helps everyone who would rather read than watch, and is the only version of the content a search engine can fully index.
Why auto-captions do not count
Every platform now offers automatic captions, and using them unedited fails WCAG.
Speech recognition degrades exactly where accuracy matters most:
- Proper nouns — names, agencies, programs, place names.
- Technical and legal vocabulary — the substantive content of most institutional video.
- Accented speech and multiple speakers, especially with overlap.
- Poor source audio — recorded meetings, hybrid rooms, phone audio.
- Punctuation and speaker changes, which ASR handles poorly.
A caption track that is 90% accurate sounds acceptable and is not. The missing 10% is disproportionately the names, numbers, and terms that carry the meaning. A misheard dollar figure or deadline in a public notice is a real harm, not a typo.
The workable approach is ASR as a first draft, then human correction against the source audio. That is materially cheaper than captioning from scratch and reaches the accuracy the standard requires.
Live video
Live content — council meetings, hearings, webinars — needs real-time captions at Level AA. In practice that means a live captioner (often CART) rather than platform auto-captions.
A common and defensible pattern for public meetings: live captioning during the broadcast, then a corrected caption file attached to the archived recording. The archived version is the one that gets cited, searched, and relied on later, so it is worth the correction pass.
Where video accessibility meets translation
If you serve multilingual communities, video is where accessibility and language access requirements collide — and the distinction matters:
- Captions are same-language and include non-speech audio.
- Subtitles are translated dialogue for viewers who can hear.
- A translated caption track does both, and is what you need for a Spanish-speaking deaf viewer.
Agencies subject to language access obligations frequently discover they need both, and that one does not substitute for the other. Related: Title VI language access plans.
A realistic sequence for a large library
You will not caption everything at once. Prioritize:
- Anything legally required or currently being relied on — public notices, benefits explanations, safety information.
- Highest-viewed content, which is usually a small fraction of the library.
- Content that will still matter in a year. Skip the 2019 event recap.
- Everything new, from now on. Make captioning part of publishing, so the backlog stops growing.
Then decide what you will not do, and write that down. A documented decision to leave low-value archived content uncaptioned, with a mechanism to caption on request, is a defensible position. Silence is not.
Getting help
Taika provides accessible video, captioning & subtitling, Section 508 remediation, and accessibility assessments — with captions and subtitles available across 300+ languages.
To scope a media library, request an assessment. Related reading: Section 508 vs ADA vs WCAG and WCAG 2.2: what changed.
Need this done right?
Taika Translations provides certified translation, interpretation, and accessibility services in 300+ languages.