Dylan Woottondwootton@mit.edu

Building a Semantic Reader

Can a small local model learn to turn a page into something easier to navigate?

Say you’re booking a ticket from Boston to SFO. You just pulled up a flight aggregator website and you want to see what flights will get you there on time for your conference.

A sighted user can skim down and quickly confirm details, see flight options, and navigate each of them. But if the page’s structure isn’t well exposed to assistive technology, that same view can turn into something like this:

Figure 1

Flight page and semantic outline

Your trip
BOSBoston
SFOSan Francisco
October 6 · One way · 1 adult

Flight options

CheapestFastest
8:35 am → 12:10 pm1 stop · 6h 35m
$248Flight details

Payment

Choose a flight to continue.

A screen reader’s navigation representation
  • link logo
  • button menu collapsed
  • Boston
  • graphic airplane
  • San Francisco
  • October sixth
  • eight thirty five AM
  • button cheapest
  • button fastest
  • link details
An authored semantic outline
Flight Booking
  • Ticket Details
  • Flight Options
  • Payment

Choose an outline item to locate it on the page.

A screen reader is a piece of assistive technology that many Blind and Low Vision (BLV) people use to navigate computer applications. Under the hood, screen readers work with accessibility information exposed by the browser or operating system. On the web, the accessibility tree is related to the DOM, but it isn’t the same thing.

When that information doesn’t convey the page’s structure, the consequences go beyond a slower interaction. In the American Foundation for the Blind’s survey, participants reported missing information, losing privacy, and having to wait for a sighted person to help. [1]

This project is about our early progress building a semantic reader—a screen reader that can navigate a page’s semantic outline, with a path back to the underlying content and controls.

A local model

Our initial approach was just directly sticking an LLM into an application that takes the DOM and spits out a curated semantic outline of the page. While this worked well in our early experiments, it has obvious privacy concerns. Imagine sending all of your page data to a separate company. A screen reader is with you for everything you do on your computer: your email, bank statements, medical portal, private messages, and work documents. Routing each of those pages through a remote model would amount to giving another company a running view of your entire computer, just so you can navigate it.

We wanted to experiment with an alternative approach: could we fine-tune a local language model to produce these semantic outlines while keeping page content on device?

Such an approach also carries other constraints. First, a system running in the background needs to work on a wide variety of hardware and take up minimal space. Secondly, it must do an acceptable job at producing semantic outlines. Thirdly… what does it even mean to produce a good semantic outline?

We’ll spend most of this note on points two and three. After some initial testing on outlines and latency (including Qwen3.5-9B, Qwen3.5-4B, Qwen3-4B-Instruct-2507, and Qwen3.5-0.8B), we selected Qwen3.5-2B for fine-tuning.

Its initial results weren’t great. One issue was the amount of content we were asking it to work through: raw DOM often bloated the context window with wrapper nodes, long text, and repeated list structure. Before fine-tuning, we needed to make that input smaller while keeping the parts someone would need to navigate the page.

We tried many variations of ways to construct this tree, including semantic similarity, templates, and hierarchical clustering. In our pilots, the qualitative results weren’t worth the extra latency and cost. We ended up with a pipeline that compacts the source tree and uses pointer-based output rather than full tree generation.

To do that, we created an intermediate representation (IR) that gives the model a more consistent description of the interface:

Identity
An ID for each element, so the outline can point back to it.
Content
Roles, text, names, and descriptions.
State
Checked, selected, expanded, and other current states.
Relationships
Labels, descriptions, and errors associated with an element.
Structure
Headings, lists, tables, and their relationships.

We compute a shared IR so that the same model can work across platforms: web, macOS, Windows, and others. It is designed to support web accessibility trees, macOS AX, and Windows UIA. The implementation here starts with the web; native platform adapters are a next step.

We also use deterministic compaction to reduce the input: removing wrapper nodes, truncating long text, and summarizing repeated list structure. In this initial experiment, that brought the input down to a median of about 10.6k tokens. The aim is to spend fewer tokens on implementation details while preserving the information needed to navigate.

Figure 2

Outline generation pipeline

Local model

The model learns to produce labels, nesting, and references from a compact input. Supervised fine-tuning teaches this mapping using example outlines.

0|Flight Booking|
1|Ticket Details|n4
1|Flight Options|
2|Compare flights|n12
1|Payment|n24
↳ Pointers resolve back to source elements, where interaction happens.

Pointer-based output gives us a way to link within the semantic outline and back to existing nodes. The model can create useful headings without rewriting the content of the entire page. Our compact output format was about 45% smaller than verbose JSON, while preserving the same outline information. At any point, a user should be able to go down to the actual source nodes for interaction.

So our final task is: given a compacted semantic IR, output a series of pointers with labels that form a navigable semantic outline.

What makes a good outline?

First we developed a rubric to assess whether a semantic outline was good. We conducted an iterative review of 25 generated outlines, critiquing each one and identifying useful navigational cues. A few qualities seemed to lead to more useful outlines.

Nesting adds semantic grouping. Some outlines mirrored the accessibility tree, producing long chains of groups with only one child each. Every level added another step to navigate without offering a useful distinction. Ideally, each choice should cut the remaining search space substantially, getting us closer to logarithmic navigation as the page gets larger.

Different levels should offer different levels of meaning. Moving down the outline should feel like zooming in: from the page’s broad purpose, to meaningful sections, to individual content and controls. Each level should reveal something more specific. However, this is also branch-specific: level three on one branch might contain an individual control, while level three on another might still be organizing a long list of results. What matters is whether that next level gives someone a useful distinction.

Surface the page’s underlying conceptual structure. We wanted the model to identify what someone might want to understand or do on a page. On a recipe page, for example, “Ingredients,” “Directions,” and “Time and yield” offer useful entry points. They let someone find what they need without first working through the page’s implementation details.

Make the outline shorter while keeping the page within reach. A product card or recipe step can become one meaningful reading unit, with its details grouped together. The semantic outline need not recreate all the details for these objects. It should allow the user to navigate to the object in the DOM, from which they can find the details and interact with it.

Reduce hallucinated page content. Early on, a model generated a group called “Tips: ingredients, technique, storage.” This is a sensible section for a recipe, but the label needs to be supported by content on the page. We need to check what the outline points to, as well as whether its organization sounds plausible.

Given these observations, we created a rubric with four categories: Coverage (how many important elements in the source we cover), Grounding (how well nodes in the semantic tree are linked to real source nodes), Grouping (whether the grouping produces an easily navigable tree), and Labels (whether the labels provide information scent for the elements below them).

We also started testing outlines with computer-use agents: if we provide an agent the semantic outline instead of the raw DOM, does it reduce the token cost or improve task accuracy? We’ll be writing more on this later.

Generating training data

Once we had an input representation and a rubric, we needed examples for the model to learn from. We curated a dataset of Gemini 3.1 Pro–labeled semantic outlines and used those examples for supervised fine-tuning (SFT) of the 2B model.

Each example pairs a compacted semantic IR with an outline: its labels, nesting, and pointers back to source nodes. This gives the small model examples of both the output format and how to organize a page. We focused on grounding during SFT so that the outline would point back to actual elements. These are model-generated training labels, rather than human gold standards.

Results

We ran our base Qwen3.5-2B, an SFT version, and a stronger 27B comparison model on 50 pages each. The base model’s results were lackluster: 49 of 50 outputs failed the validator, and the rubric judge rejected the one that passed. The SFT version dramatically improved results, but it still produced 26 validation failures, compared with 16 for the larger model.

When it does produce valid output, it performs quite well: 19 of 24 outlines were accepted by the automated rubric judge (79%), compared with 31 of 34 (91%) for the larger model. But those percentages leave out the failures. Across all 50 pages, acceptance is 38% and 62%, respectively.

Figure 3

Accepted outlines

Base · 2B
0%0 / 50
Fine-tuned · 2B
38%19 / 50
Comparison · 27B
62%31 / 50

Every page counts, including validation failures. This is end-to-end acceptance by the judge, not task success or a measure of accessibility benefit.

Full counts · 50 pages per model
ModelAcceptedReviseReject /
judge error
Failed
validator
Base Qwen3.5-2B00149
SFT Qwen3.5-2B194126
27B comparison311216
Early results from this experiment. “Accepted” is the rubric judge’s decision after structural validation. Validation failures are excluded from the conditional percentage, but remain the largest source of failure for the fine-tuned model.

Of the outputs that passed the validator, the fine-tuned model does a good job grounding its outline in the source.

Mean rubric scores among valid outputs · ordinal scale, 0–3
ModelCoverageGroundingGroupingLabels
SFT 2B · 24 pages2.082.792.332.42
27B · 34 pages2.812.972.412.84

The scale ranges from 0 to 3, where 3 indicates no problems and 0 indicates failure on that dimension.

When SFT works, it’s close. Here’s the same page, Uptime Kuma, from both models. Both outlines were accepted. I’ve put them into trees so it’s easier to see the structure. Select a leaf to follow its pointer, switch to HTML to see how the page elements are represented, or choose Diff to see what changed between the two outlines.

Figure 4

Uptime Kuma: model outputs

Visit the live Uptime Kuma page ↗
Produced outline

SFT 2B

  • Uptime Kuma group
    • Header group
    • Main content group
    • Footer group

27B comparison

  • Uptime Kuma group
    • Header group
    • Main Content group
      • Quick Links group
      • Project Links group

Diff: SFT 2B → 27B comparison

Only in SFT 2BOnly in 27BMoved nodes appear in both places

  • Uptime Kuma group
    • Header group
    • Main Content group
      • Added: Quick Links new group
      • Added: Project Links new group
    • Removed: Footer removed group

Open a group to explore its children. Groups are generated headings; a pointer such as nu refers back to a page element.

Linked page element

With JavaScript enabled, select a pointer in either outline to see its element highlighted in a reconstructed page and its HTML. Both model outlines and their diff are printed here in full.

The same pointer can receive different labels and groupings.

The outlines preserve the labels and pointers from the supplied example. The page and HTML are a minimal reconstruction, not the original evaluated snapshot; the command is abbreviated as in the draft. Choose Diff to compare them: red nodes appear only in the SFT outline and green nodes only in the 27B outline. With nu selected, the same element moves from “Footer” to “Main Content.”

The judge gave SFT 2/3/2/2 (coverage, grounding, grouping, labels) and the 27B model 3/3/2/3. Qualitatively, we can see this difference explicitly in the labels. The larger model invents useful groups—“Quick Links,” “Project Links”—and names the long command “Docker Command.”

The SFT model mostly mirrors the page’s own regions (Header / Main / Footer), uses the raw command as a label, and puts the install command in “Footer” next to the navigation links. It’s grounded and usable, just flatter. That fits part of the broader picture: grounding is close, while coverage and labels still have room to improve.

So, is this learnable?

We began this project with the goal of producing a local model capable of running on many devices in the background. These initial results suggest that a small model can learn to read a compacted page and point back at real source nodes. They also show that getting it to do that reliably is still an open problem.

Along the way we developed a compaction algorithm, curated a dataset of model-labeled semantic outlines, and used those to fine-tune a small 2B model. Looking through the traces helped us identify a set of goals for judging the output. The interface design and the model work kept informing each other.

Our next goal is to get closer to the data. Scaling up means having better labels and a clearer understanding of what useful semantic outlines actually are. We’re creating a labeler for manually making and editing these outlines, involving our BLV collaborator.

For me, that’s the interesting part of working across human–AI interaction and the model layer. A choice about how someone should navigate becomes a choice about the representation, the training examples, and the evaluation. Then the model’s mistakes send us back to the design question.