Back to all posts

Infrastructure · 6 min read

Accessibility tree for LLM agents: 462 tokens per step

What the accessibility tree is, why LLM agents read it instead of HTML, what a page map built from it costs per step, and how it catches a page's error message.

DG
Pixel art of a lone bare larch on a snowy knoll, its branches spreading in a clear orderly pattern, blue hour just before dawn

In short

The accessibility tree lists every element a screen reader can operate, each with a role, a name and a state, and a page map built from it cost 462 tokens per step on Hacker News against 12,076 to 26,438 for five other browser tools. The raw tree is larger than the HTML; the saving comes from keeping only operable elements and sending only what changed. Read with its live regions and a wait of up to 1.2 seconds, the same map carries the error a page shows after a click, for 1 extra line in 1,101 over 8 sites.

An agent that has to click "Sign in" needs to know that a sign-in button exists, what it is called and whether it is disabled. The HTML knows all of that, buried under class names, wrappers and scripts; a screenshot knows it as pixels. Many browser agents today read a third representation, the one screen readers use, and how they cut it down decides what every step costs.

What is the accessibility tree?

The accessibility tree is the browser's own model of a page as a person with a screen reader meets it: every element that means something, reduced to what it is and what it can do. The browser builds it from the DOM and hands it to the operating system's accessibility interface, which is where screen readers and voice control read it. Each node carries four properties:

PropertyWhat it answersExample
Rolewhat kind of thing it isbutton, link, textbox, checkbox
Namehow you refer to it"Sign in", "Search Wikipedia"
Statewhat condition it is indisabled, checked, expanded, invalid
Descriptionwhat else it says about itselfa hint or an error sentence tied to a field

A div styled to look like a button becomes a button in the tree only if the page gives it a role. That makes the tree a reliable list of controls on well built pages and a thinner one on badly built pages. Chrome shows it in the DevTools accessibility pane, and a driver can read the whole tree over the Chrome DevTools Protocol in one call.

Why do agents read it instead of HTML?

Agents read the accessibility tree because it can be reduced to the elements they can act on, each with a name they can refer to. Three properties make it the better input:

  • It names what can be operated. A line such as button "Post" (disabled) carries the role, the accessible name and the state an agent needs to decide; the same button in HTML is a tag with a dozen class names and an icon inside.
  • It can be filtered hard. The raw tree is not small: on the Wikipedia article the full tree is almost four times the size of the HTML. Keeping only operable elements takes the same article from 985,710 tokens of HTML to 2,771.
  • Its nodes keep their identity between readings. Because an element stays the same node while the page changes around it, a map can send only what changed after an action. On Hacker News that made a follow-up step cost 462 tokens with Scalebrowser, against 12,076 to 26,438 with five other browser tools that send the whole page again.
Tokens to see the pageWikipediaHacker NewsGitHub repository
Raw HTML985,71011,831135,129
Raw accessibility tree, full3,753,843154,516172,329
Page map of operable elements2,7714,3093,200
Page map plus the page's wording5,6885,8665,400
Source: Scalebrowser token bench, 3 live pages, cl100k_base tokenizer, 8 September 2026

The first two rows were read with Playwright directly, without an agent layer: the document as page.content() returns it and the tree as Accessibility.getFullAXTree returns it. The last row checks that the map is not small by leaving things out: adding the page's full text through a separate read still keeps every page under 6,000 tokens, and the Hacker News map carries 391 addressable elements, every story and comment link among them.

The free path makes the same choice. Microsoft's Playwright MCP describes itself as using "Playwright's accessibility tree, not pixel-based input" and promises that it "avoids ambiguity common with screenshot-based approaches". The difference between tools is how much of the tree travels with each answer, not whether they use it.

How does an agent notice what the page says after a click?

An agent notices what the page says through the same tree, provided the map reads live regions and field errors and waits up to 1,200 milliseconds (1.2 seconds) for the page to answer. Both halves were missing in our own map until 13 September 2026. An agent typed an address with no Google account behind it, clicked Next, and the screenshot showed "Couldn't find your Google Account". The reply to the click said nothing had changed, and a full map after it did not contain one word of the error.

The first cause was the channel. The W3C's WAI-ARIA 1.2 defines live regions as "perceivable regions of a web page that are typically updated as a result of an external event when user focus may be elsewhere". A plain aria-live region appears in the tree as a generic node, and Chrome gives alert and status an empty name, with their words in the child nodes. A map that filters by role and name therefore dropped the one sentence the page wanted to say.

The second cause was timing: Google rejected the unknown address 582 milliseconds after the mouse button came up, and the reply had already gone out.

The map now reads what a screen reader would announce, and nothing more:

On the pageLine in the map
aria-live="polite" or "assertive" with texta status or alert line with the words
role="alert", "status" or "log"a line with the role and its words
a field marked invalid, by ARIA or by native validationthe state (invalid) on the field
aria-errormessage or aria-describedby on an invalid fieldthe sentence on the field's line
aria-live="off", an empty region, a hint on a valid fieldnothing
red text with no semantics at allnothing, and a screenshot or text read is how you see it

The off row follows MDN's description of aria-live: updates to such a region "should not be presented to the user unless the user is currently focused on that region". The cost of the new lines was measured over 8 real sites (GitHub, Reddit, Wikipedia, Hacker News, Amazon, X, Stack Overflow and Microsoft): 1 additional line in 1,101, 0.2 percent of the characters. The waiting costs little on a page that answers: one reading of a whole page took 26 to 46 milliseconds, so the reply reads the page again until it differs, and only a page that stays silent costs the full wait. How these replies add up over a session is what the comparison of MCP browser token cost measures across six servers.

Would a screenshot-based agent do better?

A screenshot-based agent does better only on what a page paints without saying it, such as red text with no semantics or a chart in a canvas, and worse on everything the tree names. Anthropic's vision documentation prices a 1920 by 1080 screenshot at 1,560 visual tokens on its standard tier and 2,691 on the high resolution tier, whatever the page contains. Visual tokens are not the text tokens counted above, but the order of magnitude holds: far below a full tree on a large page, and above a 462 token diff on every step. The same documentation calls Claude's coordinate outputs "approximate", while a map line addresses the element itself. Keeping one still image for every browser step is covered in AI agent observability.

The comparison of computer use agents with page map agents sets both approaches side by side, and the MCP server documentation shows the map format line by line.

Run it on your own machine

Seven days to try it with your own agents on your own sites. Starting the trial needs a card.

Start the 7-day trial
DG

Davide Grasböck

Founder, Scalebrowser

Builds Scalebrowser, the browser layer for AI agents that runs on your own machine. Measures every change a web page could observe against a real browser before it ships, and writes up the ones that turned out wrong.