How Agent-Reach Handles Cookie Extraction from Browsers: Architecture and Implementation
Agent-Reach extracts authentication cookies from Chrome, Firefox, Edge, Brave, and Opera using a dual-backend system that prioritizes the Rust-based rookiepy library, falling back to browser_cookie3, then filters results against platform-specific domain patterns to populate ~/.agent-reach/config.yaml.
The Panniantong/Agent-Reach repository automates the retrieval of session cookies from local browser profiles to streamline authentication for various social platforms. This article examines how the agent-reach cookie extraction pipeline works under the hood, from SQLite database decryption to configuration file generation.
The Dual-Backend Architecture
The extractor in agent_reach/cookie_extract.py implements a resilient two-tier backend system to maximize compatibility across different environments.
Primary Backend: rookiepy
rookiepy is a Rust-based wrapper that reads browser SQLite cookie stores directly. According to the source code (lines 55-66), this is the preferred backend because it avoids sqlite3 locking issues on Windows and macOS, and returns cookies as a list of plain dictionaries with name, value, and domain keys.
Fallback Backend: browser_cookie3
If rookiepy is not installed, the system falls back to browser_cookie3, a pure-Python library that provides equivalent functionality. Both backends automatically decrypt encrypted cookie values using OS-specific keychains: DPAPI on Windows, Keychain on macOS, and GNOME-Keyring/KWallet on Linux.
The Extraction Pipeline in agent_reach/cookie_extract.py
The extract_all() function orchestrates a six-step pipeline to transform raw browser data into structured platform credentials.
Step 1: Backend Selection and Import
The function first attempts to import rookiepy (lines 55-66). If the import fails, it switches to browser_cookie3 and logs the fallback. This ensures the tool works out-of-the-box while preferring the faster Rust implementation when available.
Step 2: Browser Normalization
The supplied browser argument is lower-cased and validated against the supported list: chrome, firefox, edge, brave, and opera (lines 70-75). This normalization ensures consistent handling regardless of user input casing.
Step 3: Raw Cookie Reading
Depending on the selected backend, the code builds a dictionary of callables referencing rookiepy.chrome (or browser_cookie3.chrome, etc.) and invokes the appropriate function (lines 78-94 for rookiepy, lines 100-110 for browser_cookie3). The returned objects are normalized to a wrapper class exposing .name, .value, and .domain attributes.
Step 4: Platform Specification Mapping
A static PLATFORM_SPECS list (lines 15-41) defines extraction rules for each supported service:
- Twitter/X: Extracts specific cookies by name
- XiaoHongShu: Grabs all cookies for the domain
- Bilibili and Xueqiu: Domain-specific patterns
Each entry specifies domain patterns and required cookie names (or None to capture all cookies for that domain).
Step 5: Domain Filtering and Collection
The extractor iterates over every cookie returned by the backend (lines 18-28). It keeps only those where cookie.domain matches any pattern defined in PLATFORM_SPECS. For platforms requesting specific cookies, it filters by name; otherwise, it collects all matches (lines 31-36).
Step 6: Result Shaping
For each platform yielding data, the system creates a dictionary entry under the platform’s config_key (lines 38-47). If a platform requires a full header string, the code assembles it as name=value; .... The final output structure resembles:
{
'twitter': {'auth_token': '...', 'ct0': '...'},
'xhs': {'cookie_string': '...'},
...
}
Integration with the Configuration System
The higher-level helper configure_from_browser() (lines 25-88) bridges extraction and persistence. This function calls extract_all(), then writes discovered values into the Agent-Reach configuration file at ~/.agent-reach/config.yaml via the Config object from agent_reach/config.py.
The integration also handles platform-specific post-processing. For example, when extracting Twitter credentials, it syncs the tokens to legacy xfetch and bird credential files to maintain backward compatibility.
Practical Usage Examples
Programmatic API
To extract cookies directly in Python:
from agent_reach.cookie_extract import configure_from_browser
from agent_reach.config import Config
config = Config()
results = configure_from_browser(browser="chrome", config=config)
# Returns: [('Twitter/X', True, 'auth_token + ct0'), ('XiaoHongShu', True, '12 cookies'), ...]
print(results)
CLI Command
End-users typically invoke extraction via the command line:
# Extract from Chrome and auto-update config
agent-reach configure --from-browser chrome
# Extract from Firefox instead
agent-reach configure --from-browser firefox
The CLI parses the --from-browser flag in agent_reach/cli.py and delegates to configure_from_browser().
Manual Inspection
For debugging or custom integrations, access raw extraction results:
from agent_reach.cookie_extract import extract_all
raw = extract_all(browser="chrome")
print(raw["twitter"]["auth_token"]) # Specific token
print(raw["xhs"]["cookie_string"]) # Full header string
Summary
- Agent-reach cookie extraction supports Chrome, Firefox, Edge, Brave, and Opera through a unified interface in
agent_reach/cookie_extract.py. - The system prefers
rookiepy(Rust) for performance and reliability, falling back tobrowser_cookie3(Python) if unavailable. PLATFORM_SPECSdefines domain patterns and cookie names for Twitter/X, XiaoHongShu, Bilibili, and Xueqiu.extract_all()filters raw browser cookies against platform specifications and returns structured dictionaries.configure_from_browser()persists extracted values to~/.agent-reach/config.yamland handles platform-specific post-processing.- Both backends automatically decrypt cookies using OS-native keychains (DPAPI, Keychain, or KWallet).
Frequently Asked Questions
Which browsers does agent-reach support for cookie extraction?
Agent-Reach supports Chrome, Firefox, Edge, Brave, and Opera. The extraction logic in extract_all() normalizes browser names to lowercase and validates them against this supported list before attempting to read the SQLite cookie stores.
How does agent-reach decrypt encrypted browser cookies without manual input?
Both rookiepy and browser_cookie3 automatically handle decryption using operating-system-specific keychains. On Windows, they use DPAPI; on macOS, the Keychain; and on Linux, GNOME-Keyring or KWallet. This allows the extraction to proceed without prompting the user for passwords or decryption keys.
What happens if rookiepy is not installed?
If the rookiepy import fails, the code in extract_all() (lines 55-66) automatically falls back to browser_cookie3. While browser_cookie3 is slightly slower and may encounter SQLite locking issues on some platforms, it provides identical functionality and works out-of-the-box in pure-Python environments.
Where does agent-reach store the extracted cookies?
The configure_from_browser() function writes extracted cookies to the Agent-Reach configuration file located at ~/.agent-reach/config.yaml. The function uses the Config class from agent_reach/config.py to handle file I/O and ensures credentials are written with secure permissions (typically 0o600 as verified in tests/test_cookie_extract_perms.py).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →