How to Use computer.vision for Screenshot Analysis and OCR in Open Interpreter
Use computer.vision.ocr() to extract text via EasyOCR and computer.vision.query() to generate natural language descriptions using the Moondream2 model, accepting PIL Images, file paths, or base-64 data as input.
The computer.vision module in the openinterpreter/open-interpreter repository provides a unified interface for screenshot analysis and optical character recognition (OCR). Located under interpreter.core.computer.vision, this subsystem integrates EasyOCR for text extraction and the lightweight Moondream2 transformer for visual question answering, enabling automated analysis of desktop environments without heavy startup overhead.
Understanding the computer.vision API
The vision subsystem exposes two primary methods through the computer.vision object:
-
computer.vision.ocr()– Extracts raw text from images using EasyOCR. This method performs standard optical character recognition and returns detected text as a concatenated string. -
computer.vision.query()– Answers natural language questions about image content using the Moondream2 vision model. This enables high-level descriptions like "What applications are open?" or "Describe the current UI state."
Both methods utilize lazy loading to minimize resource consumption. EasyOCR loads on the first ocr() call, while the Moondream2 model (approximately 70MB) downloads and initializes on the first query() call.
Prerequisites and Input Formats
The computer.vision methods accept four input formats interchangeably:
pil_image– APIL.Imageobject (returned bycomputer.display.screenshot())path– String path to a local image filebase_64– Base-64 encoded image stringlmc– Local Message Content dictionary with{"format": "...", "content": "..."}structure
When processing base-64 strings or LMC objects, the system writes temporary files to disk before passing them to the underlying libraries, ensuring compatibility with EasyOCR and Moondream2 file-based APIs.
Method 1: Extract Text with computer.vision.ocr()
The ocr() method in interpreter/core/computer/vision/vision.py wraps EasyOCR's Reader.readtext() function, returning concatenated text from detected regions.
Capturing and Analyzing the Active Window
To extract text from the currently focused application:
# Capture only the active window
screenshot = computer.display.screenshot(active_app_only=True, show=False)
# Perform OCR
text = computer.vision.ocr(pil_image=screenshot)
print("Extracted text:", text)
This workflow targets UI automation scenarios where you need to read dialog text, menu items, or document content without processing the entire desktop.
Processing Local Image Files
For existing screenshots or saved images:
image_path = "/path/to/screenshot.png"
extracted_text = computer.vision.ocr(path=image_path)
print(extracted_text)
The method handles file I/O internally, converting the image to the format expected by EasyOCR.
Handling Base-64 Encoded Images
When working with images from web APIs or remote sources:
import base64
import pathlib
# Encode local file to base64 (simulating API response)
raw_bytes = pathlib.Path("image.png").read_bytes()
b64_string = base64.b64encode(raw_bytes).decode()
# OCR from base64
result = computer.vision.ocr(base_64=b64_string)
print(result)
The implementation writes the decoded bytes to a temporary file in Vision.ocr() before processing, ensuring compatibility with EasyOCR's file-based interface.
Method 2: Describe Screenshots with computer.vision.query()
The query() method leverages the Moondream2 vision model (implemented in interpreter/core/computer/vision/vision.py) to generate natural language descriptions or answer specific questions about image content.
Generating Screen Descriptions
To obtain a high-level analysis of the current desktop state:
# Capture full screen
full_screen = computer.display.screenshot(active_app_only=False, show=False)
# Query the vision model
description = computer.vision.query(
pil_image=full_screen,
query="Describe what is currently displayed on the screen. List all visible applications and UI elements."
)
print("Analysis:", description)
The first execution downloads the Moondream2 model weights (approximately 70MB) and caches them locally. Subsequent calls use the loaded model without additional network requests.
Targeted Visual Questions
For specific extraction tasks:
screenshot = computer.display.screenshot()
# Extract specific information
result = computer.vision.query(
pil_image=screenshot,
query="What is the value in the 'Total' field?"
)
print(result)
This approach combines the spatial understanding of vision models with precise information extraction, useful for form reading and data entry automation.
Implementation Details and Source Code
The vision functionality is distributed across several key files in the repository:
interpreter/core/computer/vision/vision.py
Contains the main Vision class implementing ocr() and query(). The load() method handles lazy initialization of EasyOCR and Moondream2 models. Temporary file management for base64/LMC inputs occurs in lines 70-99 (OCR) and 122-131 (query).
interpreter/core/computer/display/display.py
Provides Display.screenshot() (lines 107-115), which returns PIL.Image objects compatible with the vision API. Supports active_app_only parameter for window-specific capture.
interpreter/core/computer/computer.py
Instantiates the Vision object as computer.vision during Computer.__init__(), wiring the subsystem into the main computer API.
interpreter/core/computer/utils/computer_vision.py
Contains legacy OCR utilities including pytesseract_get_text() and find_text_in_image(), though the primary modern interface resides in the main vision.py module.
The architecture emphasizes lazy loading to maintain fast interpreter startup times. Heavy dependencies (easyocr, transformers) are only imported when their respective methods are first invoked, making the vision subsystem suitable for workflows where not every session requires image analysis.
Summary
computer.vision.ocr()extracts text from screenshots using EasyOCR, supporting PIL Images, file paths, and base-64 strings.computer.vision.query()answers natural language questions about images using the lightweight Moondream2 vision model.- Both methods use lazy loading—EasyOCR and Moondream2 models load only on first use, keeping startup times minimal.
- Input flexibility includes
pil_image,path,base_64, and LMC format dictionaries, with automatic temporary file handling for non-path inputs. - The implementation resides primarily in
interpreter/core/computer/vision/vision.py, with screenshot capture provided bycomputer.display.screenshot()ininterpreter/core/computer/display/display.py.
Frequently Asked Questions
What is the difference between computer.vision.ocr() and computer.vision.query()?
computer.vision.ocr() uses EasyOCR to perform optical character recognition, returning raw text strings detected in the image. computer.vision.query() uses the Moondream2 vision-language model to answer natural language questions about image content, enabling descriptive analysis like "What applications are open?" rather than just text extraction.
Does using computer.vision require downloading large machine learning models?
The vision subsystem uses lazy loading to minimize resource usage. EasyOCR loads on the first ocr() call, and the Moondream2 model (approximately 70MB) downloads on the first query() call. If you only use ocr(), the Moondream2 model is never downloaded, and vice versa.
Can I analyze images that are already saved as files rather than live screenshots?
Yes. Both computer.vision.ocr() and computer.vision.query() accept a path parameter pointing to local image files. They also accept base_64 strings for API workflows, pil_image objects for programmatic image manipulation, and LMC format dictionaries used internally by the interpreter.
Where is the computer.vision implementation located in the source code?
The main implementation resides in interpreter/core/computer/vision/vision.py, which defines the Vision class with ocr() and query() methods. Screenshot capture functionality is in interpreter/core/computer/display/display.py. The Vision object is instantiated as computer.vision in interpreter/core/computer/computer.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →