How to Use Qwen-Agent for Web Browsing: A Complete Guide to BrowserQwen

To use Qwen-Agent for web browsing, deploy the workstation server with your LLM credentials, load the BrowserQwen Chrome extension in developer mode, and add web pages to Qwen’s reading list to chat with the AI about their content.

The Qwen-Agent framework (available at QwenLM/Qwen-Agent) provides BrowserQwen, a browser assistant that transforms web browsing into an interactive AI experience. This system allows you to capture web pages, PDFs, and Office documents directly from Chrome, then query them using Qwen LLMs with optional code interpretation capabilities. Understanding how to use Qwen-Agent for web browsing requires familiarity with its three-tier architecture bridging your browser and local LLM infrastructure.

Understanding the BrowserQwen Architecture

BrowserQwen consists of three tightly integrated components working together to bridge your browser and the Qwen LLM.

The Workstation Server

The core orchestration logic resides in qwen_server/workstation_server.py. This Flask-based server manages conversation history, renders the browsing list UI, and coordinates plugin execution.

Key functions include:

  • update_browser_list() (lines 153-170): Retrieves persisted browsing records using read_meta_data_by_condition and renders them as an HTML checklist in the UI.
  • bot() and pure_bot() (lines 45-108): Handle chat pipelines by assembling message history, attaching uploaded files, and invoking the LLM via Assistant or ReActChat classes.
  • choose_plugin() (lines 89-96): Toggles the Code Interpreter (CI_OPTION) and Document QA (DOC_OPTION) based on user selection.

The server maintains global state in an app_global_para dictionary that tracks messages, uploaded files, and filtering parameters.

The Database Service

The run_server.py entry point launches a FastAPI service that persists browsing metadata and conversation logs. It exposes endpoints for saving page data (save_browsing_meta_data) and retrieving records (read_meta_data_by_condition).

The Chrome Extension

Located in the browser_qwen/ directory, the unpacked extension injects UI elements into web pages and forwards content to the workstation. Key files include:

  • manifest.json: Declares permissions for activeTab, scripting, and storage.
  • src/content.js: Extracts DOM snapshots and captures screenshots, sending payloads to the workstation via HTTP POST requests.
  • src/popup.js: Manages the extension popup interface for user interactions.

Setting Up Qwen-Agent for Web Browsing

Starting the Backend Services

Launch the database and workstation services using the CLI entry point. For DashScope (Alibaba Cloud) models:

python run_server.py --llm qwen-max \
    --model_server dashscope \
    --workstation_port 7864 \
    --api_key YOUR_DASHSCOPE_API_KEY

For self-hosted models via vLLM:

python run_server.py --llm Qwen1.5-72B-Chat \
    --model_server http://localhost:8000/v1 \
    --api_key EMPTY \
    --workstation_port 7864

The workstation UI becomes available at http://127.0.0.1:7864/, presenting Editor and Chat tabs for different interaction modes.

Installing the Browser Extension

  1. Open Chrome and navigate to chrome://extensions/.
  2. Enable Developer mode using the toggle in the top-right corner.
  3. Click Load unpacked and select the browser_qwen folder from your local repository.

The BrowserQwen icon appears in your toolbar, ready to capture page content.

Capturing and Interacting with Web Content

Adding Pages to Your Reading List

When browsing, click the "Add to Qwen’s Reading List" button in the extension popup. The content.js script extracts the page metadata and sends a JSON payload to the workstation:

fetch('http://127.0.0.1:7864/add_page', {
    method: 'POST',
    headers: {'Content-Type': 'application/json'},
    body: JSON.stringify({
        url: location.href,
        title: document.title,
        html: document.documentElement.outerHTML,
        screenshot: base64ImageData  // optional
    })
});

The /add_page endpoint stores this data via the database service, making the page appear in the workstation’s browsing list panel.

Querying Content with Code Interpretation

Interact with saved content through the workstation UI or extension popup. The system constructs message payloads combining your query with page references:

content = [
    {'text': user_query},
    {'file': page_url}
]

Enable the Code Interpreter plugin from the UI dropdown to allow data visualization. The backend invokes ReActChat with function_list=['code_interpreter'], enabling Qwen to generate and execute Python code against the captured content.

Managing Browsing History

The workstation automatically refreshes your reading list using update_browser_list(). You can also query records programmatically:

curl http://127.0.0.1:7864/browsing_records?start=2024-01-01&end=2024-01-31

This returns metadata for pages saved within the specified date range, which the UI renders as interactive checkboxes alongside their URLs.

Summary

  • BrowserQwen combines a workstation server, database service, and Chrome extension to enable AI-augmented web browsing according to the QwenLM/Qwen-Agent source code.
  • Deploy the backend using run_server.py with either DashScope or self-hosted model endpoints.
  • Install the browser_qwen extension in Chrome developer mode to capture page DOM and screenshots.
  • Use the Code Interpreter plugin for data analysis tasks on saved web content.
  • Browse history persists in a local FastAPI service, accessible via REST endpoints or the workstation UI.

Frequently Asked Questions

What models does Qwen-Agent support for web browsing?

The system supports both cloud-hosted models via DashScope (such as qwen-max) and self-hosted models (such as Qwen1.5-72B-Chat) through OpenAI-compatible endpoints like vLLM. Specify your model using the --llm and --model_server flags when starting run_server.py.

How does the Chrome extension extract page content?

The extension's src/content.js script injects into web pages and captures the full DOM via document.documentElement.outerHTML, along with the page title and URL. It can optionally capture screenshots as base64-encoded images, transmitting this data to the workstation's /add_page endpoint.

Can I use Qwen-Agent with a self-hosted LLM?

Yes. Set --model_server to your local endpoint (e.g., http://localhost:8000/v1 for vLLM) and use --api_key EMPTY or your local API key. This configuration routes all inference requests to your private infrastructure rather than DashScope.

Where is the browsing history stored?

The FastAPI database service persists browsing metadata locally. The update_browser_list() function in qwen_server/workstation_server.py retrieves these records using read_meta_data_by_condition, displaying them in the workstation UI with checkboxes for selection during chat sessions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →