Audio captcha: transcribed locally in 1.3 seconds

How an AI agent answers the audio captcha of reCAPTCHA v2 on its own machine: a local Whisper tiny model, 1.3 seconds, no speech service and no model shipped.

Davide Grasböck · Founder, Scalebrowser

Published Oct 6, 2026 · 5 min read

Short answer

An AI agent answers an audio captcha by downloading the clip and transcribing it on the same machine: on 14 August 2026 a local Whisper tiny model turned a reCAPTCHA v2 clip into text in 1.3 seconds, model loading included, and the typed answer turned the checkbox green. No speech service was called and no model ships with the product; the operator names the transcription program. The audio path is a fallback for a page that offers nothing else, since image tasks are solvable too.
1.3 s, local Whisper tiny, beside a model box with three speaker dots

Every major captcha keeps a second door next to the picture puzzle: a headphone icon that plays a few spoken words or digits instead. For an agent that cannot see well or keeps misreading tiles, that door looks like the easy way through. We measured it on Google's reCAPTCHA v2 demo on 14 August 2026, end to end and through the agent's own tools, and the numbers turned out smaller than we had assumed.

What is an audio captcha?

An audio captcha is the spoken alternative to a visual challenge: the widget plays a short, distorted clip and asks the visitor to type what was said. It exists for people who cannot solve an image task, above all blind and low-vision users with screen readers. The W3C's note "Inaccessibility of CAPTCHA", a Group Draft Note of 16 December 2021, describes the idea as offering "another non-textual method of using the same content".

The same note records both of its problems. For people, distorted audio is hard: in one test cited there, Hotmail's clip "was unintelligible to all four test subjects, all of whom had good hearing". For machines it is easy: the note cites a University of Maryland study with a "90% success rate cracking Google's audio reCAPTCHA using Google's own speech recognition service". Vendors keep the door open anyway, because closing it would lock out the people it was built for. GeeTest's accessibility page describes its own "audio challenge" as the way its captcha meets WCAG 2.0. What the three most common captchas each ask of an agent is covered in reCAPTCHA vs hCaptcha vs Turnstile.

ProviderAudio alternativeWhat the visitor types
Google reCAPTCHA v2headphone button in the image puzzlethe spoken words
GeeTest v4button at the bottom of the visual challengethe spoken numbers

How does the agent transcribe the clip?

The agent transcribes the clip with a local speech model and types the text, using the same audio alternative a person with a screen reader is offered. On Google's reCAPTCHA v2 demo the clip was a 40 kilobyte MP3 at 22 kilohertz, and no speech service outside the machine was involved.

MeasureValue
ClipMP3, 40 kilobytes, 22 kilohertz
ModelWhisper tiny, int8, on the CPU
Model file39 megabytes
Time to text, model loading included1.3 seconds
Time to text, model already loaded0.2 seconds
Whisper base, same clipsame text in 11 seconds
Run 1 transcript"cleaning request, efficient schedule"
Run 2 transcript, tools only"and very important channel"
Resultgreen checkmark in both runs
Source: Scalebrowser agent on Google's reCAPTCHA v2 demo, production site key, 2 runs, 14 August 2026

The model is small. OpenAI's Whisper repository lists tiny at 39 million parameters and about 1 gigabyte of video memory, with a relative speed of about 10 times the large model, under the MIT licence. Our first estimate, written down before anyone measured, was a model of about 150 megabytes and a noticeable CPU cost per clip. The measurement came in at a quarter of that size and 1.3 seconds, and an estimate that far off would have wrongly ruled the path out.

Scalebrowser

Scalebrowser gives each agent its own isolated browser with a persistent identity, on your own machine, so a run stays signed in, handles the captcha and finishes without anyone watching it.

See how Scalebrowser does it

Why is no speech service involved?

No speech service is involved because the transcription runs as a program the operator names on the same machine, and the product ships neither a model nor a network call for it. That is the same rule behind solving captchas without a service: no paid solver, no key for one, and no code path that sends a challenge anywhere. transcribe_audio takes one command from the configuration, gives it the audio file as its single argument and reads its standard output as the answer; without a configured command it says so instead of guessing. A demo video needs no recognition step of this kind, because its captions come from text the agent already wrote; that path is covered in AI subtitles for demo videos.

Two reasons kept the model out of the product. A 39 megabyte model and its native toolchain would sit in every build for a path that is meant to be used rarely. And the model version decides whether a transcript can be reproduced, so it belongs to the operator, who can pin it. The verification page of the documentation lists the command's contract.

The one party that sees the clip besides you is the captcha provider that served it. The file travels from Google to the browser, from the browser to a local file, and from there to a local program; nothing leaves the machine a second time. One trap showed up on the way: a Windows path handed to a script on the Linux side of the same machine lost its backslashes, and the script reported "file not found". The path now goes out with forward slashes, which Windows accepts everywhere.

How do image tasks compare?

Image tasks are solvable too, so the audio path is a fallback rather than the default. The same demo's image grid was solved at the first attempt with 3 correct tiles, and hCaptcha solved by an agent took 4 of 4 runs of about 35 seconds per set on a third-party deployment. The audio door exists for people who cannot use the image, and we treat it as the path for a page that offers nothing else. A challenge that asks for a held press is covered in PerimeterX press and hold. A challenge that asks for a piece to be slid into a gap is covered in slider captcha.

Run it on your own machine

Seven days to try it with your own agents on your own sites. Starting the trial needs a card.

Written by

Davide Grasböck

Founder, Scalebrowser

Builds Scalebrowser, the browser layer for AI agents that runs on your own machine. Measures every change a web page could observe against a real browser before it ships, and writes up the ones that turned out wrong.

All articles by Davide Grasböck