How to Use Qwen-Agent for Web Browsing: A Complete Guide to BrowserQwen
To use Qwen-Agent for web browsing, deploy the workstation server with your LLM credentials, load the BrowserQwen Chrome extension in developer mode, and add web pages to Qwen’s reading list to chat with the AI about their content.
The Qwen-Agent framework (available at QwenLM/Qwen-Agent) provides BrowserQwen, a browser assistant that transforms web browsing into an interactive AI experience. This system allows you to capture web pages, PDFs, and Office documents directly from Chrome, then query them using Qwen LLMs with optional code interpretation capabilities. Understanding how to use Qwen-Agent for web browsing requires familiarity with its three-tier architecture bridging your browser and local LLM infrastructure.
Understanding the BrowserQwen Architecture
BrowserQwen consists of three tightly integrated components working together to bridge your browser and the Qwen LLM.
The Workstation Server
The core orchestration logic resides in qwen_server/workstation_server.py. This Flask-based server manages conversation history, renders the browsing list UI, and coordinates plugin execution.
Key functions include:
update_browser_list()(lines 153-170): Retrieves persisted browsing records usingread_meta_data_by_conditionand renders them as an HTML checklist in the UI.bot()andpure_bot()(lines 45-108): Handle chat pipelines by assembling message history, attaching uploaded files, and invoking the LLM viaAssistantorReActChatclasses.choose_plugin()(lines 89-96): Toggles the Code Interpreter (CI_OPTION) and Document QA (DOC_OPTION) based on user selection.
The server maintains global state in an app_global_para dictionary that tracks messages, uploaded files, and filtering parameters.
The Database Service
The run_server.py entry point launches a FastAPI service that persists browsing metadata and conversation logs. It exposes endpoints for saving page data (save_browsing_meta_data) and retrieving records (read_meta_data_by_condition).
The Chrome Extension
Located in the browser_qwen/ directory, the unpacked extension injects UI elements into web pages and forwards content to the workstation. Key files include:
manifest.json: Declares permissions foractiveTab,scripting, andstorage.src/content.js: Extracts DOM snapshots and captures screenshots, sending payloads to the workstation via HTTP POST requests.src/popup.js: Manages the extension popup interface for user interactions.
Setting Up Qwen-Agent for Web Browsing
Starting the Backend Services
Launch the database and workstation services using the CLI entry point. For DashScope (Alibaba Cloud) models:
python run_server.py --llm qwen-max \
--model_server dashscope \
--workstation_port 7864 \
--api_key YOUR_DASHSCOPE_API_KEY
For self-hosted models via vLLM:
python run_server.py --llm Qwen1.5-72B-Chat \
--model_server http://localhost:8000/v1 \
--api_key EMPTY \
--workstation_port 7864
The workstation UI becomes available at http://127.0.0.1:7864/, presenting Editor and Chat tabs for different interaction modes.
Installing the Browser Extension
- Open Chrome and navigate to
chrome://extensions/. - Enable Developer mode using the toggle in the top-right corner.
- Click Load unpacked and select the
browser_qwenfolder from your local repository.
The BrowserQwen icon appears in your toolbar, ready to capture page content.
Capturing and Interacting with Web Content
Adding Pages to Your Reading List
When browsing, click the "Add to Qwen’s Reading List" button in the extension popup. The content.js script extracts the page metadata and sends a JSON payload to the workstation:
fetch('http://127.0.0.1:7864/add_page', {
method: 'POST',
headers: {'Content-Type': 'application/json'},
body: JSON.stringify({
url: location.href,
title: document.title,
html: document.documentElement.outerHTML,
screenshot: base64ImageData // optional
})
});
The /add_page endpoint stores this data via the database service, making the page appear in the workstation’s browsing list panel.
Querying Content with Code Interpretation
Interact with saved content through the workstation UI or extension popup. The system constructs message payloads combining your query with page references:
content = [
{'text': user_query},
{'file': page_url}
]
Enable the Code Interpreter plugin from the UI dropdown to allow data visualization. The backend invokes ReActChat with function_list=['code_interpreter'], enabling Qwen to generate and execute Python code against the captured content.
Managing Browsing History
The workstation automatically refreshes your reading list using update_browser_list(). You can also query records programmatically:
curl http://127.0.0.1:7864/browsing_records?start=2024-01-01&end=2024-01-31
This returns metadata for pages saved within the specified date range, which the UI renders as interactive checkboxes alongside their URLs.
Summary
- BrowserQwen combines a workstation server, database service, and Chrome extension to enable AI-augmented web browsing according to the
QwenLM/Qwen-Agentsource code. - Deploy the backend using
run_server.pywith either DashScope or self-hosted model endpoints. - Install the
browser_qwenextension in Chrome developer mode to capture page DOM and screenshots. - Use the Code Interpreter plugin for data analysis tasks on saved web content.
- Browse history persists in a local FastAPI service, accessible via REST endpoints or the workstation UI.
Frequently Asked Questions
What models does Qwen-Agent support for web browsing?
The system supports both cloud-hosted models via DashScope (such as qwen-max) and self-hosted models (such as Qwen1.5-72B-Chat) through OpenAI-compatible endpoints like vLLM. Specify your model using the --llm and --model_server flags when starting run_server.py.
How does the Chrome extension extract page content?
The extension's src/content.js script injects into web pages and captures the full DOM via document.documentElement.outerHTML, along with the page title and URL. It can optionally capture screenshots as base64-encoded images, transmitting this data to the workstation's /add_page endpoint.
Can I use Qwen-Agent with a self-hosted LLM?
Yes. Set --model_server to your local endpoint (e.g., http://localhost:8000/v1 for vLLM) and use --api_key EMPTY or your local API key. This configuration routes all inference requests to your private infrastructure rather than DashScope.
Where is the browsing history stored?
The FastAPI database service persists browsing metadata locally. The update_browser_list() function in qwen_server/workstation_server.py retrieves these records using read_meta_data_by_condition, displaying them in the workstation UI with checkboxes for selection during chat sessions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →