For most of the security industry, the audio challenge is a compliance box to check, an accessibility feature bolted onto a visual-first architecture to satisfy WCAG requirements and reduce legal exposure. This framing misunderstands what the audio challenge is becoming.
AI agents acting on behalf of vision-impaired and deaf/blind users are in production traffic now, at volumes that will only increase as AI assistants become the primary interface through which millions of people navigate digital services. At the network layer, these agents look identical to the autonomous systems running credential stuffing campaigns. A security layer that cannot tell them apart is not doing its job. And the architectural choice that determines whether the distinction is possible, speech-based versus sound-based, is one most vendors have gotten wrong.
The Structural Problem With Speech-Based Audio
Most audio challenge solutions are built on the same architectural foundation: spoken words, letters, or numbers that a user must transcribe. This creates a contradiction that is difficult to engineer away. Any audio challenge that can be heard and understood by an AI accessibility agent can, by definition, be solved by that same agent or any comparable speech recognition system.
Making it harder to solve automatically means making it harder to use accessibly. Making it accessible means making it solvable by automation. The attack surface and the accessibility surface are the same surface. No implementation improvement closes this gap without eliminating the accessibility functionality.
The competitive picture makes this concrete. Most legacy CAPTCHA vendors use spoken digits or phrases layered with background noise, trivially defeated by any English speech-to-text system, which is precisely the tool that accessibility users rely on.
Some competitors offer no audio challenge at all. Their primary accessibility accommodation routes users off-site to register an email address in exchange for an encrypted cookie that must be periodically refreshed and is subject to failure in browsers that block cross-site cookies, including Brave and privacy-configured Firefox and Safari. An optional text-based challenge is sometimes offered as an enterprise alternative. Neither path resolves the core problem: both the cookie flow and the text challenge are bypassable without genuine audio understanding, and neither is designed to distinguish an authorized AI accessibility agent from automated abuse.
Other vendors' audio challenges remain speech-based: spoken words that a user transcribes. Whatever languages they support, the structural vulnerability is the same: any speech-to-text system capable of serving a legitimate accessibility user is equally capable of solving the challenge programmatically. Language coverage is irrelevant to the security argument.
Sound-Based, Not Speech-Based
The alternative to speech-based audio is sound-based audio, challenges built on sounds rather than words, where the question itself is language agnostic by design. Arkose Labs' audio challenge is built on exactly this principle.
Non-speech audio — animal sounds, environmental audio, counting footsteps, speaker identification tasks — requires auditory perception and contextual reasoning that humans perform naturally but that speech-to-text systems cannot engage with. There is no speech to transcribe, so the attack vector that defeats speech-based audio challenges does not apply. Arkose Labs' audio challenge is built entirely on this category of sound-based interaction.
Because the format is not speech-based, it requires no translation. The challenges are language and culture agnostic by design, supporting a wide range of languages natively compared to the English-only or severely limited coverage of every competitor's speech-based implementation. That is a structural advantage no speech-based approach can replicate without building separate audio assets for each language it supports.
The entire experience is inline, no off-site redirect, no email requirement, no browser compatibility restrictions. TalkBack on Android and VoiceOver on iOS are fully supported. And Arkose Labs' Allow/Monitor/Challenge/Throttle/Block policy framework can be configured to recognize and appropriately handle authorized AI agents, including those operating as accessibility tools on behalf of identified legitimate users, without routing them through challenge flows designed for unknown traffic.
Built From Real Feedback, Not Assumptions
Most audio challenges are designed by engineers reasoning about what accessibility should feel like. Arkose Labs took a different approach: we conducted testing sessions with users with visual impairments directly, recorded their interactions, and let them tell us where the challenge was alienating and what went wrong.
Those sessions changed the design. Not in abstract ways, in specific, concrete ones. The places where the experience created confusion or friction for someone navigating entirely by sound became the inputs for the next iteration. As documented in our user testing sessions, the result is an audio challenge that was shaped by the people it is supposed to serve, not just evaluated against a standard after the fact.
This matters because a checklist can tell you whether a product meets criteria. It cannot tell you whether a user with visual impairments finds the experience respectful or disorienting. Only the user can tell you that. And most vendors never ask.
Certified, Not Self-Assessed
Accessibility compliance claims in this industry are almost always self-assessed or internally audited, which leaves customers exposed if that claim is ever challenged. Independent, third-party certification closes that gap. Arkose Labs is the only agent trust and control vendor with enforcement challenges certified to WCAG 2.2 Level AA by an independent, accredited third-party assessor, me2 Accessibility. Not self-assessed. Not internally audited. Certified by a named organization whose conformance report is available to customers.
The distinction matters for customers because it determines legal exposure when a compliance claim is challenged.
Audio Sessions Generate Signal Too
The intelligence loop from Blog 2 runs through audio sessions as fully as visual ones. Solve timing distributions, failure patterns, and interaction consistency all flow back into the detection model. Legitimate users navigating by sound, including AI agents acting on their behalf, produce behavioral profiles that are measurably different from automated systems attempting to solve programmatically. Nothing is wasted.
What makes Arkose Labs' audio challenge different:
- Non-speech audio, animal sounds, environmental audio, counting footsteps, speaker identification, breaks the link between accessibility and automated solvability, since there's no speech for a speech-to-text system to transcribe
- Language and culture agnostic by design, with native support across a wide range of languages rather than the English-only or limited coverage typical of speech-based challenges
- Fully inline, with no off-site redirect, no email requirement, and no browser compatibility restrictions, including full TalkBack and VoiceOver support
- Independently certified to WCAG 2.2 Level AA by a third-party assessor, not self-assessed
- Audio session behavior feeds the same detection model as visual challenges, distinguishing legitimate accessibility use from automated abuse
"In the age of agentic AI, your audio challenge is not just an accessibility feature, it is a critical signal about who is in your traffic. Our audio challenge is independently certified to WCAG 2.2 Level AA by me2 Accessibility, built on non-speech audio that defeats speech-to-text attacks, and designed to distinguish legitimate AI accessibility agents from malicious automation." – Gab Fitzgerald
Accessible, defensible challenge design is only as strong as the infrastructure delivering it. Blog 6 covers the engineering layer beneath the challenge: CAPI v4's VM-based obfuscation, per-session encryption, and browser allowlisting, and why vendors whose infrastructure was built for a slower threat are now dangerously exposed.
Continue Reading: The Disrupting Fraud Economics Series
Blog 1: The Economics of Fraud Have Changed. Here's Why.
Blog 2: We Are Not a CAPTCHA — Why the Turing test model is obsolete
Blog 3: What Attackers Taught Us — Proprietary attacker data that shaped MatchKey
Blog 4: Inside MatchKey — Architecture designed to make attacks economically irrational
Blog 5: The Audio Challenge — The only audio challenge that is both accessible and secure (this post)
Blog 6: Challenge Engineering Built to Withstand AI-Powered Attacks — The engineering beneath the challenge
Blog 7 coming soon.




