A mid-size software company recorded a product demo video, ran it through an AI captioning tool, watched the captions appear, and moved on. Nobody re-read the text line by line. Why would they? The numbers were right, the sentences made sense, the video shipped.
Months later, someone finally scrubbed through the captions on mute. Buried in the video was the company's own contact web address, and the caption spelled it wrong. Not a typo a viewer would laugh off. A domain name that, if someone actually typed what the caption showed, would take them nowhere near the real site. The video had been live, indexed, and shared for months.
That is the quiet failure mode of AI captioning: it is not usually wrong in ways that are obvious. It is wrong in ways that look finished.
The Stat: In our own real test, an AI captioning engine transcribed every number and technical term correctly, then mis-transcribed "WCAG" as "WKAG" twice, including inside a spoken web address, so the caption literally showed the wrong domain: wkag.world instead of wcag.world. (Source: our own real, reproducible test, detailed in the article below)
What We Actually Tested
Most claims about AI caption accuracy come from either marketing copy or a vague "it's pretty good now." We wanted something more concrete, so we built a small, honest test we could describe exactly.
We wrote a real product-demo voiceover script, about 45 seconds long, packed with the kind of language an accessibility or compliance team actually uses: acronyms like WCAG and ARIA, a testing tool name (axe-core), a legal reference (Section 508), a product brand name (AccessRadar), and several specific issue counts. We generated real speech audio from that exact script using an AI text-to-speech engine. Then we ran that audio through an AI transcription and captioning engine and diffed the output against our original script, word for word.
This was not a formal industry benchmark and we are not naming vendors. It is a small, reproducible test of one real workflow, reported plainly. We are not claiming this represents every AI captioning tool on the market, or that the same script would fail the same way on a different engine. What we can say, because we watched it happen on our own machine with our own audio file, is that this exact failure mode is real, it is repeatable, and it is the kind of thing a busy team is unlikely to catch without deliberately looking for it.
The point of running a test this small and this specific is that it is easy to verify and hard to argue with. There is no sample size to debate, no methodology to pick apart. There is a script, an audio file, and a transcript, and you can read all three yourself.
The Results, Honestly
The good news first, because it matters: the AI did not fall apart. Numbers came through perfectly. Twenty-six, fourteen, and twelve all appeared correctly as 26, 14, and 12, with no rounding, no dropped digits, no confusion between similar-sounding numbers.
Technical jargon also held up. "ARIA attributes," "axe-core," and "Section 508 conformance" all transcribed correctly and legibly, exactly as spoken. If the test had stopped there, the honest headline would be "AI captions are better than people assume."
But the test did not stop there. Two specific terms broke, and they were not random.
"WCAG" was mis-transcribed as "WKAG" in both places it appeared, including inside the spoken contact email domain. The caption did not just misspell an acronym in passing. It rendered a visible, on-screen web address as wkag.world instead of wcag.world. Anyone reading the caption instead of listening to the audio would see the wrong destination for the product being demonstrated.
The product name "AccessRadar" was split into two separate words, "Access Radar," breaking the brand name wherever it appeared in caption text.
Why This Pattern Is the Real Story
The verdict here is not "AI captions are unreliable." Common words and numbers transcribed perfectly, consistently, across the whole clip. The real, specific failure mode was exactly the terms that matter most for a company's own credibility and findability: its acronym and its brand name, both mangled, with one error landing directly inside a spoken web address shown as text on screen.
That is a very particular kind of risk. A generic transcription error might confuse a viewer for a second. A wrong URL, sitting in burned-in or synced caption text, actively misdirects them, and it does so in the exact video where a company is trying to look credible and, often, trying to demonstrate compliance. If that same error showed up in a video making accessibility or conformance claims, mislabeling the very acronym the claim depends on is not just embarrassing. It is the kind of small, un-reviewed AI error that becomes a real legal exposure the moment someone screenshots it.
Proper names, acronyms, and brand terms sit outside the AI's normal vocabulary distribution. The model has seen "twenty-six" a billion times and "AccessRadar" essentially never. That is exactly why these are the words that need a human pass, every time, no exceptions.
A Practical Checklist Before You Publish
- Never publish AI-generated captions unreviewed. Treat the first output as a draft, not a deliverable.
- Search the caption file for every acronym, brand name, and proper noun in your script. Do a literal find-and-check, not a skim.
- Pay special attention to spoken URLs, emails, and domains. These are where a small transcription slip becomes a wrong address shown as text.
- Read the captions with the sound off. Errors that are invisible while you're listening to correct audio become obvious the moment you rely on the text alone, which is exactly what a Deaf or hard-of-hearing viewer is doing.
- Keep a style glossary of your product names and industry acronyms and run it against every new caption file before publishing.
None of this replaces a real accessibility strategy. Captions are one piece of a much larger picture, and if you have not nailed down the basics of what your video actually needs, that groundwork comes first, not after. It is also worth separating caption text accuracy, which is what we tested here, from whether your video player's controls themselves are usable, since captions can be flawless and a video can still fail WCAG if the player controls aren't accessible.
For the legal backdrop on why this matters beyond user experience, the W3C's guidance on captions lays out the accessibility case in detail, and the ADA.gov site is the place to understand the compliance stakes directly. Neither source needed our test to make the point; our test just made it concrete.
If you want a second set of eyes checking exactly this kind of gap across your own video library, before a customer or a regulator finds it first, that is precisely what AccessRadar is built to catch.
The Soft Ask
You do not need to run your own version of this test to know the risk is real; you just watched us find it in under a minute of footage. If you want a team that will actually read your captions instead of trusting the export button, reach out to experts@wcag.world or take a look at AccessRadar. A wrong web address in your own demo video is a fixable, five-minute problem, right up until someone else finds it first.
