---
title: "Accessibility tree for LLM agents: 462 tokens per step"
description: "What the accessibility tree is, why LLM agents read it instead of HTML, what a page map built from it costs per step, and how it catches a page's error message."
canonical: "https://scalebrowser.net/blog/accessibility-tree"
last_modified: "2026-10-04"
published: "2026-10-04"
author: "davide"
category: "infrastructure"
---

> ## Documentation Index
> Fetch the complete documentation index at: https://scalebrowser.net/llms.txt
> Use this file to discover all available pages before exploring further.

# Accessibility tree for LLM agents: 462 tokens per step

> What the accessibility tree is, why LLM agents read it instead of HTML, what a page map built from it costs per step, and how it catches a page's error message.

An agent that has to click "Sign in" needs to know that a sign-in button exists, what it is called and whether it is disabled. The HTML knows all of that, buried under class names, wrappers and scripts; a screenshot knows it as pixels. Many browser agents today read a third representation, the one screen readers use, and how they cut it down decides what every step costs.

## What is the accessibility tree?

The accessibility tree is the browser's own model of a page as a person with a screen reader meets it: every element that means something, reduced to what it is and what it can do. The browser builds it from the DOM and hands it to the operating system's accessibility interface, which is where screen readers and voice control read it. Each node carries four properties:

| Property | What it answers | Example |
| --- | --- | --- |
| Role | what kind of thing it is | `button`, `link`, `textbox`, `checkbox` |
| Name | how you refer to it | "Sign in", "Search Wikipedia" |
| State | what condition it is in | disabled, checked, expanded, invalid |
| Description | what else it says about itself | a hint or an error sentence tied to a field |

A `div` styled to look like a button becomes a button in the tree only if the page gives it a role. That makes the tree a reliable list of controls on well built pages and a thinner one on badly built pages. Chrome shows it in the DevTools accessibility pane, and a driver can read the whole tree over the Chrome DevTools Protocol in one call.

## Why do agents read it instead of HTML?

Agents read the accessibility tree because it can be reduced to the elements they can act on, each with a name they can refer to. Three properties make it the better input:

- **It names what can be operated.** A line such as `button "Post" (disabled)` carries the role, the accessible name and the state an agent needs to decide; the same button in HTML is a tag with a dozen class names and an icon inside.
- **It can be filtered hard.** The raw tree is not small: on the Wikipedia article the full tree is almost four times the size of the HTML. Keeping only operable elements takes the same article from 985,710 tokens of HTML to 2,771.
- **Its nodes keep their identity between readings.** Because an element stays the same node while the page changes around it, a map can send only what changed after an action. On Hacker News that made a follow-up step cost 462 tokens with Scalebrowser, against 12,076 to 26,438 with five other browser tools that send the whole page again.

<Evidence source="Scalebrowser token bench, 3 live pages, cl100k_base tokenizer, 8 September 2026">

| Tokens to see the page | Wikipedia | Hacker News | GitHub repository |
| --- | --- | --- | --- |
| Raw HTML | 985,710 | 11,831 | 135,129 |
| Raw accessibility tree, full | 3,753,843 | 154,516 | 172,329 |
| Page map of operable elements | 2,771 | 4,309 | 3,200 |
| Page map plus the page's wording | 5,688 | 5,866 | 5,400 |

</Evidence>

The first two rows were read with Playwright directly, without an agent layer: the document as `page.content()` returns it and the tree as `Accessibility.getFullAXTree` returns it. The last row checks that the map is not small by leaving things out: adding the page's full text through a separate read still keeps every page under 6,000 tokens, and the Hacker News map carries 391 addressable elements, every story and comment link among them.

The free path makes the same choice. Microsoft's [Playwright MCP](https://github.com/microsoft/playwright-mcp) describes itself as using "Playwright's accessibility tree, not pixel-based input" and promises that it "avoids ambiguity common with screenshot-based approaches". The difference between tools is how much of the tree travels with each answer, not whether they use it.

## How does an agent notice what the page says after a click?

An agent notices what the page says through the same tree, provided the map reads live regions and field errors and waits up to 1,200 milliseconds (1.2 seconds) for the page to answer. Both halves were missing in our own map until 13 September 2026. An agent typed an address with no Google account behind it, clicked Next, and the screenshot showed "Couldn't find your Google Account". The reply to the click said nothing had changed, and a full map after it did not contain one word of the error.

The first cause was the channel. The W3C's [WAI-ARIA 1.2](https://www.w3.org/TR/wai-aria-1.2/) defines live regions as "perceivable regions of a web page that are typically updated as a result of an external event when user focus may be elsewhere". A plain `aria-live` region appears in the tree as a generic node, and Chrome gives `alert` and `status` an empty name, with their words in the child nodes. A map that filters by role and name therefore dropped the one sentence the page wanted to say.

The second cause was timing: Google rejected the unknown address 582 milliseconds after the mouse button came up, and the reply had already gone out.

The map now reads what a screen reader would announce, and nothing more:

| On the page | Line in the map |
| --- | --- |
| `aria-live="polite"` or `"assertive"` with text | a `status` or `alert` line with the words |
| `role="alert"`, `"status"` or `"log"` | a line with the role and its words |
| a field marked invalid, by ARIA or by native validation | the state `(invalid)` on the field |
| `aria-errormessage` or `aria-describedby` on an invalid field | the sentence on the field's line |
| `aria-live="off"`, an empty region, a hint on a valid field | nothing |
| red text with no semantics at all | nothing, and a screenshot or text read is how you see it |

The `off` row follows MDN's [description of aria-live](https://developer.mozilla.org/en-US/docs/Web/Accessibility/ARIA/Reference/Attributes/aria-live): updates to such a region "should not be presented to the user unless the user is currently focused on that region". The cost of the new lines was measured over 8 real sites (GitHub, Reddit, Wikipedia, Hacker News, Amazon, X, Stack Overflow and Microsoft): 1 additional line in 1,101, 0.2 percent of the characters. The waiting costs little on a page that answers: one reading of a whole page took 26 to 46 milliseconds, so the reply reads the page again until it differs, and only a page that stays silent costs the full wait. How these replies add up over a session is what the comparison of [MCP browser token cost](/blog/browser-mcp) measures across six servers.

## Would a screenshot-based agent do better?

A screenshot-based agent does better only on what a page paints without saying it, such as red text with no semantics or a chart in a canvas, and worse on everything the tree names. Anthropic's [vision documentation](https://platform.claude.com/docs/en/build-with-claude/vision) prices a 1920 by 1080 screenshot at 1,560 visual tokens on its standard tier and 2,691 on the high resolution tier, whatever the page contains. Visual tokens are not the text tokens counted above, but the order of magnitude holds: far below a full tree on a large page, and above a 462 token diff on every step. The same documentation calls Claude's coordinate outputs "approximate", while a map line addresses the element itself. Keeping one still image for every browser step is covered in AI agent observability.

The comparison of computer use agents with page map agents sets both approaches side by side, and the [MCP server documentation](/docs/agents/mcp-server) shows the map format line by line.
