OCR Integration Through read_text for Non-Accessible Text in pi-computer-use

You can extract text from image-based or custom-drawn UI elements by capturing screenshots, applying external OCR, and feeding the results back into the read_text pipeline via synthetic output references.

The pi-computer-use repository provides accessibility-driven automation tools that rely on native OS APIs to read text from UI elements. Because these APIs cannot expose custom-drawn canvas text or image-based dialogs, you must implement OCR integration through read_text for non-accessible text to capture content from inaccessible regions. This approach leverages the existing pagination system in src/output.ts and the executor registration in src/bridge.ts without requiring modifications to the core library.

How read_text Handles Accessible vs. Non-Accessible Text

Native Accessibility Limitations

The read_text tool operates on two reference types: desktop (@e) refs created by inspect_ui or observe_ui, and browser (@o) refs representing immutable session output. While @e refs expose the UI accessibility tree, they cannot surface content from non-reporting elements like canvas-rendered text or image-based CAPTCHAs. These limitations are inherent to the operating system's accessibility APIs, which only report explicitly tagged UI elements.

The read_text Implementation

In src/bridge.ts, the executor registers the read_text handler that processes UTF-8 byte slices from captured outlines. The src/output.ts module manages truncation for large outputs, adding a continue hint when pagination is required. This architecture is defined in src/contract.ts and documented in docs/usage.md and docs/architecture.md.

OCR Integration Workflow for Non-Accessible Elements

When accessibility APIs fail to expose text, follow this three-step pattern:

  1. Capture a screenshot of the target region using observe_ui with screenshot:true or the low-level capture API.
  2. Execute OCR using an external library like Tesseract WASM or Google Vision.
  3. Store the result as a synthetic @o ref using a store_output helper, enabling read_text to consume it with standard pagination.

The following example demonstrates basic read_text usage on an accessible element:

// Obtain a reference to a text-bearing UI element
const { outline } = await pi.runTool("inspect_ui", { selector: "#article" });
const textRef = outline.refs[0];      // e.g. "@e3"

// Read the first 1 KB of text (UTF-8 bytes)
const { text, continue: next } = await pi.runTool("read_text", {
  ref: textRef,
  offset: 0,
  limit: 1024,
});
console.log(text);

// If output was truncated, request the next page
if (next) {
  const more = await pi.runTool("read_text", { ref: textRef, offset: 1024 });
  console.log(more.text);
}

For non-accessible regions, combine screenshot capture with OCR processing:

// 1️⃣ Capture a screenshot of the target area
const { screenshot } = await pi.runTool("observe_ui", {
  selector: "#captcha",      // element that only shows an image
  screenshot: true,
});

// 2️⃣ Send PNG bytes to an OCR service (Tesseract WASM example)
import { recognize } from "tesseract-wasm";
const ocrResult = await recognize(screenshot.base64);

// 3️⃣ Store OCR text as synthetic output ref
const ocrRef = await pi.runTool("store_output", {
  content: ocrResult.text,
});

// 4️⃣ Read via read_text like any other ref
const { text } = await pi.runTool("read_text", { ref: ocrRef, offset: 0 });
console.log("OCR output:", text);

Handling Large OCR Outputs with Pagination

When processing large images or extensive OCR results, implement chunking to respect memory constraints and API limits:

const pageSize = 5000;
let offset = 0;
let fullText = "";

while (offset < screenshot.base64.length) {
  const chunk = screenshot.base64.slice(offset, offset + pageSize);
  const { text: part } = await recognize(chunk);
  fullText += part;
  offset += pageSize;
}

const ocrRef = await pi.runTool("store_output", { content: fullText });
const { text } = await pi.runTool("read_text", { ref: ocrRef, offset: 0 });
console.log(text);

Key Source Files and Implementation Details

Understanding the core architecture helps you extend the library correctly:

  • src/bridge.ts – Registers the read_text executor and defines the tool payload structure.
  • src/output.ts – Handles truncation of large textual outputs and provides the continue hint for pagination.
  • src/contract.ts – Declares the public read_text interface and parameter types.
  • docs/usage.md – Documents pagination semantics and reference handling.
  • docs/architecture.md – Explains the immutable state model for UI roots and outlines.
  • extensions/computer-use.ts – Declares tools for the Pi extension side; useful when adding custom helpers like store_output.

Summary

  • OCR integration through read_text bridges the gap when native accessibility APIs cannot expose visual content.
  • Use observe_ui with screenshot:true to capture non-accessible regions as base64 PNGs.
  • Process images with external OCR libraries, then store results via synthetic @o references.
  • The existing pagination system in src/output.ts handles large OCR outputs without core library modifications.
  • All integration happens through the public tool contract in src/contract.ts, requiring no changes to pi-computer-use internals.

Frequently Asked Questions

Can read_text directly perform OCR on screenshots?

No. The read_text tool only processes UTF-8 text from accessibility trees or stored output references. You must run OCR externally using libraries like Tesseract or Google Vision, then feed the resulting string back into the system via store_output to create a readable @o reference.

What is the difference between @e and @o references?

@e (desktop) references point to live UI elements captured via inspect_ui or observe_ui, reading from the OS accessibility tree. @o (output) references are immutable session objects containing arbitrary text data, including synthetic OCR results you create with store_output. Both support the same read_text pagination semantics defined in src/contract.ts.

How does pagination work for large OCR results?

The read_text tool accepts offset and limit parameters to read specific byte ranges. When output exceeds the limit, src/output.ts returns a continue flag. You increment the offset and call read_text again until the flag is false, as implemented in the executor within src/bridge.ts.

Do I need to modify the core pi-computer-use library to add OCR support?

No. The architecture supports external integration through the public tool contract. You only need to implement a store_output helper in your extension (see extensions/computer-use.ts for patterns) to create synthetic references, then process screenshots with your chosen OCR library.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →