# How to Use computer.vision for Screenshot Analysis and OCR in Open Interpreter

> Unlock screenshot analysis and OCR with Open Interpreter. Learn to use computer.vision.ocr() for text extraction and computer.vision.query() for natural language descriptions.

- Repository: [Open Interpreter/open-interpreter](https://github.com/openinterpreter/open-interpreter)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Use `computer.vision.ocr()` to extract text via EasyOCR and `computer.vision.query()` to generate natural language descriptions using the Moondream2 model, accepting PIL Images, file paths, or base-64 data as input.**

The `computer.vision` module in the [openinterpreter/open-interpreter](https://github.com/openinterpreter/open-interpreter) repository provides a unified interface for screenshot analysis and optical character recognition (OCR). Located under `interpreter.core.computer.vision`, this subsystem integrates EasyOCR for text extraction and the lightweight Moondream2 transformer for visual question answering, enabling automated analysis of desktop environments without heavy startup overhead.

## Understanding the computer.vision API

The vision subsystem exposes two primary methods through the `computer.vision` object:

- **`computer.vision.ocr()`** – Extracts raw text from images using EasyOCR. This method performs standard optical character recognition and returns detected text as a concatenated string.

- **`computer.vision.query()`** – Answers natural language questions about image content using the Moondream2 vision model. This enables high-level descriptions like "What applications are open?" or "Describe the current UI state."

Both methods utilize **lazy loading** to minimize resource consumption. EasyOCR loads on the first `ocr()` call, while the Moondream2 model (approximately 70MB) downloads and initializes on the first `query()` call.

## Prerequisites and Input Formats

The `computer.vision` methods accept four input formats interchangeably:

1. **`pil_image`** – A `PIL.Image` object (returned by `computer.display.screenshot()`)
2. **`path`** – String path to a local image file
3. **`base_64`** – Base-64 encoded image string
4. **`lmc`** – Local Message Content dictionary with `{"format": "...", "content": "..."}` structure

When processing base-64 strings or LMC objects, the system writes temporary files to disk before passing them to the underlying libraries, ensuring compatibility with EasyOCR and Moondream2 file-based APIs.

## Method 1: Extract Text with computer.vision.ocr()

The `ocr()` method in [`interpreter/core/computer/vision/vision.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/vision/vision.py) wraps EasyOCR's `Reader.readtext()` function, returning concatenated text from detected regions.

### Capturing and Analyzing the Active Window

To extract text from the currently focused application:

```python

# Capture only the active window

screenshot = computer.display.screenshot(active_app_only=True, show=False)

# Perform OCR

text = computer.vision.ocr(pil_image=screenshot)

print("Extracted text:", text)

```

This workflow targets UI automation scenarios where you need to read dialog text, menu items, or document content without processing the entire desktop.

### Processing Local Image Files

For existing screenshots or saved images:

```python
image_path = "/path/to/screenshot.png"
extracted_text = computer.vision.ocr(path=image_path)

print(extracted_text)

```

The method handles file I/O internally, converting the image to the format expected by EasyOCR.

### Handling Base-64 Encoded Images

When working with images from web APIs or remote sources:

```python
import base64
import pathlib

# Encode local file to base64 (simulating API response)

raw_bytes = pathlib.Path("image.png").read_bytes()
b64_string = base64.b64encode(raw_bytes).decode()

# OCR from base64

result = computer.vision.ocr(base_64=b64_string)
print(result)

```

The implementation writes the decoded bytes to a temporary file in `Vision.ocr()` before processing, ensuring compatibility with EasyOCR's file-based interface.

## Method 2: Describe Screenshots with computer.vision.query()

The `query()` method leverages the Moondream2 vision model (implemented in [`interpreter/core/computer/vision/vision.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/vision/vision.py)) to generate natural language descriptions or answer specific questions about image content.

### Generating Screen Descriptions

To obtain a high-level analysis of the current desktop state:

```python

# Capture full screen

full_screen = computer.display.screenshot(active_app_only=False, show=False)

# Query the vision model

description = computer.vision.query(
    pil_image=full_screen,
    query="Describe what is currently displayed on the screen. List all visible applications and UI elements."
)

print("Analysis:", description)

```

The first execution downloads the Moondream2 model weights (approximately 70MB) and caches them locally. Subsequent calls use the loaded model without additional network requests.

### Targeted Visual Questions

For specific extraction tasks:

```python
screenshot = computer.display.screenshot()

# Extract specific information

result = computer.vision.query(
    pil_image=screenshot,
    query="What is the value in the 'Total' field?"
)

print(result)

```

This approach combines the spatial understanding of vision models with precise information extraction, useful for form reading and data entry automation.

## Implementation Details and Source Code

The vision functionality is distributed across several key files in the repository:

**[`interpreter/core/computer/vision/vision.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/vision/vision.py)**
Contains the main `Vision` class implementing `ocr()` and `query()`. The `load()` method handles lazy initialization of EasyOCR and Moondream2 models. Temporary file management for base64/LMC inputs occurs in lines 70-99 (OCR) and 122-131 (query).

**[`interpreter/core/computer/display/display.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/display/display.py)**
Provides `Display.screenshot()` (lines 107-115), which returns `PIL.Image` objects compatible with the vision API. Supports `active_app_only` parameter for window-specific capture.

**[`interpreter/core/computer/computer.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/computer.py)**
Instantiates the `Vision` object as `computer.vision` during `Computer.__init__()`, wiring the subsystem into the main computer API.

**[`interpreter/core/computer/utils/computer_vision.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/utils/computer_vision.py)**
Contains legacy OCR utilities including `pytesseract_get_text()` and `find_text_in_image()`, though the primary modern interface resides in the main [`vision.py`](https://github.com/openinterpreter/open-interpreter/blob/main/vision.py) module.

The architecture emphasizes **lazy loading** to maintain fast interpreter startup times. Heavy dependencies (`easyocr`, `transformers`) are only imported when their respective methods are first invoked, making the vision subsystem suitable for workflows where not every session requires image analysis.

## Summary

- **`computer.vision.ocr()`** extracts text from screenshots using EasyOCR, supporting PIL Images, file paths, and base-64 strings.
- **`computer.vision.query()`** answers natural language questions about images using the lightweight Moondream2 vision model.
- Both methods use **lazy loading**—EasyOCR and Moondream2 models load only on first use, keeping startup times minimal.
- Input flexibility includes `pil_image`, `path`, `base_64`, and LMC format dictionaries, with automatic temporary file handling for non-path inputs.
- The implementation resides primarily in [`interpreter/core/computer/vision/vision.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/vision/vision.py), with screenshot capture provided by `computer.display.screenshot()` in [`interpreter/core/computer/display/display.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/display/display.py).

## Frequently Asked Questions

### What is the difference between computer.vision.ocr() and computer.vision.query()?

`computer.vision.ocr()` uses EasyOCR to perform optical character recognition, returning raw text strings detected in the image. `computer.vision.query()` uses the Moondream2 vision-language model to answer natural language questions about image content, enabling descriptive analysis like "What applications are open?" rather than just text extraction.

### Does using computer.vision require downloading large machine learning models?

The vision subsystem uses **lazy loading** to minimize resource usage. EasyOCR loads on the first `ocr()` call, and the Moondream2 model (approximately 70MB) downloads on the first `query()` call. If you only use `ocr()`, the Moondream2 model is never downloaded, and vice versa.

### Can I analyze images that are already saved as files rather than live screenshots?

Yes. Both `computer.vision.ocr()` and `computer.vision.query()` accept a `path` parameter pointing to local image files. They also accept `base_64` strings for API workflows, `pil_image` objects for programmatic image manipulation, and LMC format dictionaries used internally by the interpreter.

### Where is the computer.vision implementation located in the source code?

The main implementation resides in [`interpreter/core/computer/vision/vision.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/vision/vision.py), which defines the `Vision` class with `ocr()` and `query()` methods. Screenshot capture functionality is in [`interpreter/core/computer/display/display.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/display/display.py). The `Vision` object is instantiated as `computer.vision` in [`interpreter/core/computer/computer.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/computer/computer.py).