# How the Browser-Use Framework Enables Computer Use with Open-Source Models

> Learn how the browser-use framework empowers open-source LLMs like Llama 3 to control web browsers. Discover the think act observe loop for AI agent computer use.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-17

---

**The browser-use framework turns any OpenAI-compatible open-source LLM into a computer-use agent by exposing browser actions as JSON-schema tools that the model calls via a lightweight MCP (Message Control Protocol), allowing local models like Llama 3 or Mistral to control real web browsers through a "think → act → observe" loop.**

The **browser-use** package, located in the `chapter9/browser-use-rpa/browser-use/` directory of the `bojieli/ai-agent-book` repository, provides a lightweight automation stack that abstracts complex browser interactions behind a unified LLM interface. By implementing schema-driven tool calling and a Message Control Protocol (MCP) for real-time communication, the framework enables computer use with open-source models without requiring proprietary SDKs or cloud-only APIs.

## Core Architecture of the Browser-Use Framework

The framework consists of six tightly integrated components that bridge the gap between LLM reasoning and actual browser execution.

### Browser Profile and Session Management

The `BrowserProfile` class in [`browser_use/browser/profile.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/browser/profile.py) defines how the headless browser should be launched—including viewport dimensions, proxy settings, and authentication credentials. This configuration pairs with `BrowserSession` in [`browser_use/browser/session.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/browser/session.py), which maintains per-step state including DOM snapshots, navigation history, and current page context that is serialized and sent to the LLM with every interaction.

Together, these models allow the same codebase to drive both local Playwright instances and remote cloud browsers running in Docker containers, as implemented in [`browser_use/browser/cloud.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/browser/cloud.py).

### MCP Protocol for Real-Time Communication

The **Message Control Protocol (MCP)** is a lightweight JSON-over-WebSocket layer that decouples the LLM agent process from the browser process. 

- **Server**: [`browser_use/mcp/server.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/mcp/server.py) runs inside the browser process and forwards DOM events, navigation changes, and errors.
- **Client**: [`browser_use/mcp/client.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/mcp/client.py) lives in the agent process and transmits tool commands (click, type, screenshot) initiated by the LLM.

This abstraction ensures the LLM never needs to know whether it is controlling a local or remote browser instance.

### Tool Registry and Schema Validation

Every browser capability is exposed as an OpenAI-style tool with a strict JSON schema defined in [`browser_use/tools/registry/service.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/tools/registry/service.py). When the LLM returns a tool call, the framework validates the JSON structure against the schema before execution in [`browser_use/tools/service.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/tools/service.py), preventing hallucinated or malformed commands from reaching the browser.

Supported actions include `click`, `type`, `navigate`, `extract`, `screenshot`, and file downloads, each implemented as a discrete tool with typed parameters.

### Agent Orchestration and LLM Integration

The `Agent` class in [`browser_use/agent/views.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/agent/views.py) orchestrates the main interaction loop:

1. **Observe**: Captures current page state (markdown extraction via `markdownify` plus DOM snapshot).
2. **Think**: Sends context, available tools, and short-term memory (last N steps) to the LLM.
3. **Act**: Parses the LLM's response—either a plain text answer or a JSON tool call—and forwards valid commands over MCP.
4. **Update**: Records the result, updates token usage via [`browser_use/tokens/service.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/tokens/service.py), and repeats until the task completes.

Because the agent accepts any `BaseChatModel` instance, developers can swap between GPT-4, Claude, or open-source alternatives simply by changing the LLM provider parameter.

## Why Browser-Use Works with Open-Source Models

The framework removes traditional barriers that prevented local LLMs from performing reliable computer-use tasks through four key design decisions.

**Schema-driven tool calling** eliminates ambiguity by forcing the LLM to output strictly validated JSON objects that match predefined tool signatures. The framework rejects any response that does not conform to the schema in [`browser_use/tools/registry/service.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/tools/registry/service.py), ensuring only syntactically correct browser commands execute.

**OpenAI-compatible API abstraction** allows any model server—whether Groq's Llama 4, a local vLLM endpoint, or Ollama—to act as the backbone. The agent in [`browser_use/agent/views.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/agent/views.py) treats all models identically as long as they implement the standard chat-completions format.

**Separate page-extraction LLM** support lets developers use a cheap, fast model (such as a quantized 7B parameter model) to convert raw HTML into clean markdown for the main reasoning model. This reduces context window pressure and inference costs when processing large web pages.

**Vision toggle and token accounting** in [`browser_use/tokens/service.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/tokens/service.py) automatically adjusts payloads based on model capabilities. If the open-source model lacks vision support, screenshots are omitted to reduce token usage; if it supports multimodal inputs, visual context is included. The service also pulls pricing data for cost tracking across providers, including community-maintained open-source endpoints.

## Practical Implementation Examples

### Running Llama 4 via Groq API

This minimal setup demonstrates computer use with open-source models through Groq's high-speed inference endpoint:

```python

# examples/getting_started/05_fast_agent.py

from browser_use.agent.views import Agent
from browser_use.browser.profile import BrowserProfile
from browser_use.utils import load_llm  # helper that returns a BaseChatModel

# 1️⃣  Choose a fast, open-source model (Llama 4 on Groq)

llm = load_llm(provider="groq", model="llama-4-8b", temperature=0.0, flash_mode=True)

# 2️⃣  Configure a headless Chromium instance

profile = BrowserProfile(
    headless=True,
    viewport={"width": 1280, "height": 800},
    # Optional proxy for corporate environments

    proxy={"server": "http://my-proxy:3128"},
)

# 3️⃣  Create the agent

agent = Agent(
    llm=llm,
    profile=profile,
    max_steps=15,                 # safety guard

    llm_timeout=60,               # seconds

)

# 4️⃣  Run a task – the LLM decides which browser actions to take

result = agent.run("Find the current price of the 2026 Tesla Model Y on the official website and return it as JSON.")
print(result.extracted_content)  # → {"price":"$54,990","currency":"USD"}

```

### Local Deployment with vLLM

For fully offline computer use, connect to a locally hosted Llama 3 instance via vLLM's OpenAI-compatible server:

```python

# examples/custom-functions/local_llama.py

from browser_use.agent.views import Agent
from browser_use.browser.profile import BrowserProfile
from openai import OpenAI  # vLLM exposes OpenAI-compatible endpoint

llm = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")
profile = BrowserProfile(headless=False)

agent = Agent(llm=llm, profile=profile, max_steps=20)
out = agent.run("Log in to my GitHub account (username: alice, password: x_pass), navigate to the Pull Requests tab and list the titles of open PRs.")
print(out.extracted_content)

```

### Secure Handling of Sensitive Data

The framework prevents credential leakage by substituting placeholders after the LLM call, as shown in [`examples/features/secure.py`](https://github.com/bojieli/ai-agent-book/blob/main/examples/features/secure.py):

```python

# examples/features/secure.py

from browser_use.agent.views import Agent
from browser_use.browser.profile import BrowserProfile
from browser_use.tools.service import register_sensitive_data

profile = BrowserProfile(headless=True)
agent = Agent(llm=my_open_source_llm, profile=profile, sensitive_data={"x_pass": "mySecretPwd"})

# The LLM will see only the placeholder `x_pass`; the framework injects the real value when typing.

agent.run("Log in to example.com with username 'bob' and password 'x_pass'.")

```

## Summary

- **Browser-use** enables computer use with open-source models through a **schema-driven tool registry** that validates all LLM outputs before execution.
- The **MCP protocol** ([`browser_use/mcp/server.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/mcp/server.py) and [`browser_use/mcp/client.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/mcp/client.py)) provides transparent communication between the LLM agent and browser process, supporting both local and cloud deployments.
- **OpenAI-compatible API support** allows any chat-style model—including local Llama 3, Mistral, or Groq-hosted endpoints—to drive browser automation without code changes.
- **Secure credential handling** uses placeholder substitution to ensure the LLM never receives raw passwords or secrets, as implemented in the tool service layer.
- **Automatic token accounting** in [`browser_use/tokens/service.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/tokens/service.py) tracks usage and costs across providers, optimizing for budget-conscious open-source deployments.

## Frequently Asked Questions

### What LLM providers are supported by the browser-use framework?

Any provider that implements an OpenAI-compatible chat-completions API is supported, including Groq, Together AI, Ollama, vLLM, and llama.cpp. The `Agent` class in [`browser_use/agent/views.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/agent/views.py) accepts any object conforming to the `BaseChatModel` interface, allowing you to pass a custom client pointing to local endpoints at `http://localhost:8000/v1` or similar.

### How does the framework prevent hallucinated browser commands?

All browser actions are defined as JSON-schema tools in [`browser_use/tools/registry/service.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/tools/registry/service.py). Before execution, the framework validates that the LLM's response matches the schema exactly—if the JSON is malformed or contains undefined parameters, the request is rejected and the agent prompts the model to correct its output. This validation layer ensures only syntactically valid click, type, and navigate commands reach the browser.

### Can I run browser-use entirely offline with local models?

Yes. By pointing the LLM client to a local inference server such as vLLM, Ollama, or Text Generation Inference (TGI), and setting `headless=True` in the `BrowserProfile`, you can execute the entire agent loop without external API calls. The MCP protocol functions identically whether the browser runs locally or remotely, making fully air-gapped deployments possible.

### How does credential security work in browser-use?

Sensitive data is passed via the `sensitive_data` dictionary parameter when initializing the `Agent`. The LLM receives only placeholders like `x_user` or `x_pass` in its prompt; the actual values are substituted by the tool execution layer in [`browser_use/tools/service.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/tools/service.py) immediately before the browser performs the typing action. This ensures raw credentials never appear in LLM logs, context windows, or observability traces.