Say you’re booking a ticket from Boston to SFO. You just pulled up a flight aggregator website and you want to see what flights will get you there on time for your conference.
A sighted user can skim down and quickly confirm details, see flight options, and navigate each of them. But if the page’s structure isn’t well exposed to assistive technology, that same view can turn into something like this:
Flight page and semantic outline
Flight options
Payment
Choose a flight to continue.
- link logo
- button menu collapsed
- Boston
- graphic airplane
- San Francisco
- October sixth
- eight thirty five AM
- button cheapest
- button fastest
- link details
Flight Booking
Ticket Details
Flight Options
Payment
Choose an outline item to locate it on the page.
A screen reader is a piece of assistive technology that many Blind and Low Vision (BLV) people use to navigate computer applications. Under the hood, screen readers work with accessibility information exposed by the browser or operating system. On the web, the accessibility tree is related to the DOM, but it isn’t the same thing.
When that information doesn’t convey the page’s structure, the consequences go beyond a slower interaction. In the American Foundation for the Blind’s survey, participants reported missing information, losing privacy, and having to wait for a sighted person to help. [1]
This project is about our early progress building a semantic reader—a screen reader that can navigate a page’s semantic outline, with a path back to the underlying content and controls.
A local model
Our initial approach was just directly sticking an LLM into an application that takes the DOM and spits out a curated semantic outline of the page. While this worked well in our early experiments, it has obvious privacy concerns. Imagine sending all of your page data to a separate company. A screen reader is with you for everything you do on your computer: your email, bank statements, medical portal, private messages, and work documents. Routing each of those pages through a remote model would amount to giving another company a running view of your entire computer, just so you can navigate it.
We wanted to experiment with an alternative approach: could we fine-tune a local language model to produce these semantic outlines while keeping page content on device?
Such an approach also carries other constraints. First, a system running in the background needs to work on a wide variety of hardware and take up minimal space. Secondly, it must do an acceptable job at producing semantic outlines. Thirdly… what does it even mean to produce a good semantic outline?
We’ll spend most of this note on points two and three. After some initial testing on outlines and latency (including Qwen3.5-9B, Qwen3.5-4B, Qwen3-4B-Instruct-2507, and Qwen3.5-0.8B), we selected Qwen3.5-2B for fine-tuning.
Its initial results weren’t great. One issue was the amount of content we were asking it to work through: raw DOM often bloated the context window with wrapper nodes, long text, and repeated list structure. Before fine-tuning, we needed to make that input smaller while keeping the parts someone would need to navigate the page.
We tried many variations of ways to construct this tree, including semantic similarity, templates, and hierarchical clustering. In our pilots, the qualitative results weren’t worth the extra latency and cost. We ended up with a pipeline that compacts the source tree and uses pointer-based output rather than full tree generation.
To do that, we created an intermediate representation (IR) that gives the model a more consistent description of the interface:
- Identity
- An ID for each element, so the outline can point back to it.
- Content
- Roles, text, names, and descriptions.
- State
- Checked, selected, expanded, and other current states.
- Relationships
- Labels, descriptions, and errors associated with an element.
- Structure
- Headings, lists, tables, and their relationships.
We compute a shared IR so that the same model can work across platforms: web, macOS, Windows, and others. It is designed to support web accessibility trees, macOS AX, and Windows UIA. The implementation here starts with the web; native platform adapters are a next step.
We also use deterministic compaction to reduce the input: removing wrapper nodes, truncating long text, and summarizing repeated list structure. In this initial experiment, that brought the input down to a median of about 10.6k tokens. The aim is to spend fewer tokens on implementation details while preserving the information needed to navigate.
Outline generation pipeline
Local model
The model learns to produce labels, nesting, and references from a compact input. Supervised fine-tuning teaches this mapping using example outlines.
0|Flight Booking| 1|Ticket Details|n4 1|Flight Options| 2|Compare flights|n12 1|Payment|n24
Pointer-based output gives us a way to link within the semantic outline and back to existing nodes. The model can create useful headings without rewriting the content of the entire page. Our compact output format was about 45% smaller than verbose JSON, while preserving the same outline information. At any point, a user should be able to go down to the actual source nodes for interaction.
So our final task is: given a compacted semantic IR, output a series of pointers with labels that form a navigable semantic outline.
What makes a good outline?
First we developed a rubric to assess whether a semantic outline was good. We conducted an iterative review of 25 generated outlines, critiquing each one and identifying useful navigational cues. A few qualities seemed to lead to more useful outlines.
Nesting adds semantic grouping. Some outlines mirrored the accessibility tree, producing long chains of groups with only one child each. Every level added another step to navigate without offering a useful distinction. Ideally, each choice should cut the remaining search space substantially, getting us closer to logarithmic navigation as the page gets larger.
Different levels should offer different levels of meaning. Moving down the outline should feel like zooming in: from the page’s broad purpose, to meaningful sections, to individual content and controls. Each level should reveal something more specific. However, this is also branch-specific: level three on one branch might contain an individual control, while level three on another might still be organizing a long list of results. What matters is whether that next level gives someone a useful distinction.
Surface the page’s underlying conceptual structure. We wanted the model to identify what someone might want to understand or do on a page. On a recipe page, for example, “Ingredients,” “Directions,” and “Time and yield” offer useful entry points. They let someone find what they need without first working through the page’s implementation details.
Make the outline shorter while keeping the page within reach. A product card or recipe step can become one meaningful reading unit, with its details grouped together. The semantic outline need not recreate all the details for these objects. It should allow the user to navigate to the object in the DOM, from which they can find the details and interact with it.
Reduce hallucinated page content. Early on, a model generated a group called “Tips: ingredients, technique, storage.” This is a sensible section for a recipe, but the label needs to be supported by content on the page. We need to check what the outline points to, as well as whether its organization sounds plausible.
Given these observations, we created a rubric with four categories: Coverage (how many important elements in the source we cover), Grounding (how well nodes in the semantic tree are linked to real source nodes), Grouping (whether the grouping produces an easily navigable tree), and Labels (whether the labels provide information scent for the elements below them).
We also started testing outlines with computer-use agents: if we provide an agent the semantic outline instead of the raw DOM, does it reduce the token cost or improve task accuracy? We’ll be writing more on this later.
Generating training data
Once we had an input representation and a rubric, we needed examples for the model to learn from. We curated a dataset of Gemini 3.1 Pro–labeled semantic outlines and used those examples for supervised fine-tuning (SFT) of the 2B model.
Each example pairs a compacted semantic IR with an outline: its labels, nesting, and pointers back to source nodes. This gives the small model examples of both the output format and how to organize a page. We focused on grounding during SFT so that the outline would point back to actual elements. These are model-generated training labels, rather than human gold standards.
Results
We ran our base Qwen3.5-2B, an SFT version, and a stronger 27B comparison model on 50 pages each. The base model’s results were lackluster: 49 of 50 outputs failed the validator, and the rubric judge rejected the one that passed. The SFT version dramatically improved results, but it still produced 26 validation failures, compared with 16 for the larger model.
When it does produce valid output, it performs quite well: 19 of 24 outlines were accepted by the automated rubric judge (79%), compared with 31 of 34 (91%) for the larger model. But those percentages leave out the failures. Across all 50 pages, acceptance is 38% and 62%, respectively.
Accepted outlines
Every page counts, including validation failures. This is end-to-end acceptance by the judge, not task success or a measure of accessibility benefit.
| Model | Accepted | Revise | Reject / judge error | Failed validator |
|---|---|---|---|---|
| Base Qwen3.5-2B | 0 | 0 | 1 | 49 |
| SFT Qwen3.5-2B | 19 | 4 | 1 | 26 |
| 27B comparison | 31 | 1 | 2 | 16 |
Of the outputs that passed the validator, the fine-tuned model does a good job grounding its outline in the source.
| Model | Coverage | Grounding | Grouping | Labels |
|---|---|---|---|---|
| SFT 2B · 24 pages | 2.08 | 2.79 | 2.33 | 2.42 |
| 27B · 34 pages | 2.81 | 2.97 | 2.41 | 2.84 |
The scale ranges from 0 to 3, where 3 indicates no problems and 0 indicates failure on that dimension.
When SFT works, it’s close. Here’s the same page, Uptime Kuma, from both models. Both outlines were accepted. I’ve put them into trees so it’s easier to see the structure. Select a leaf to follow its pointer, switch to HTML to see how the page elements are represented, or choose Diff to see what changed between the two outlines.
Uptime Kuma: model outputs
SFT 2B
Uptime Kuma group
Header group
Main content group
Footer group
27B comparison
Uptime Kuma group
Header group
Main Content group
Quick Links group
Project Links group
Diff: SFT 2B → 27B comparison
Only in SFT 2BOnly in 27BMoved nodes appear in both places
Uptime Kuma group
Header group
Main Content group
Added: Quick Links new group
Added: Project Links new group
Removed: Footer removed group
Open a group to explore its children. Groups are generated headings; a pointer such as nu refers back to a page element.
With JavaScript enabled, select a pointer in either outline to see its element highlighted in a reconstructed page and its HTML. Both model outlines and their diff are printed here in full.
The same pointer can receive different labels and groupings.
The outlines preserve the labels and pointers from the supplied example. The page and HTML are a minimal reconstruction, not the original evaluated snapshot; the command is abbreviated as in the draft. Choose Diff to compare them: red nodes appear only in the SFT outline and green nodes only in the 27B outline. With nu selected, the same element moves from “Footer” to “Main Content.”
The judge gave SFT 2/3/2/2 (coverage, grounding, grouping, labels) and the 27B model 3/3/2/3. Qualitatively, we can see this difference explicitly in the labels. The larger model invents useful groups—“Quick Links,” “Project Links”—and names the long command “Docker Command.”
The SFT model mostly mirrors the page’s own regions (Header / Main / Footer), uses the raw command as a label, and puts the install command in “Footer” next to the navigation links. It’s grounded and usable, just flatter. That fits part of the broader picture: grounding is close, while coverage and labels still have room to improve.
So, is this learnable?
We began this project with the goal of producing a local model capable of running on many devices in the background. These initial results suggest that a small model can learn to read a compacted page and point back at real source nodes. They also show that getting it to do that reliably is still an open problem.
Along the way we developed a compaction algorithm, curated a dataset of model-labeled semantic outlines, and used those to fine-tune a small 2B model. Looking through the traces helped us identify a set of goals for judging the output. The interface design and the model work kept informing each other.
Our next goal is to get closer to the data. Scaling up means having better labels and a clearer understanding of what useful semantic outlines actually are. We’re creating a labeler for manually making and editing these outlines, involving our BLV collaborator.
For me, that’s the interesting part of working across human–AI interaction and the model layer. A choice about how someone should navigate becomes a choice about the representation, the training examples, and the evaluation. Then the model’s mistakes send us back to the design question.