# Cookie Extraction Process in Agent-Reach: How It Accesses Browser Authentication

> Discover Agent-Reach's cookie extraction process. Learn how it accesses browser authentication data for platforms like Twitter/X, Bilibili, and more using libraries like rookiepy.

- Repository: [Pnant/Agent-Reach](https://github.com/Panniantong/Agent-Reach)
- Tags: internals
- Published: 2026-06-16

---

**Agent-Reach extracts authentication cookies directly from local browsers using a dual-backend approach that prefers the Rust-based `rookiepy` library, falling back to `browser_cookie3`, then filters them against platform-specific specifications to configure API access for Twitter/X, XiaoHongShu, Bilibili, and Xueqiu.**

The cookie extraction process in the Panniantong/Agent-Reach repository automates the retrieval of session credentials from local browser storage. Located in [`agent_reach/cookie_extract.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/cookie_extract.py), this functionality eliminates manual copy-pasting of authentication tokens by reading cookie jars directly from Chrome, Firefox, Edge, Brave, and Opera.

## How the Cookie Extraction Process Works

The extraction flow follows a seven-step pipeline that isolates low-level cookie reading from configuration logic. This design allows easy addition of new platforms or swapping of extraction backends without modifying the core application.

### Browser Selection and Validation

The process begins when the caller specifies a browser name (`chrome`, `firefox`, `edge`, `brave`, or `opera`). The `normalize_browser` function validates this input against the supported list and raises a clear error for invalid selections. According to lines 70-76 in [`agent_reach/cookie_extract.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/cookie_extract.py), this normalization ensures consistent handling across different browser naming conventions.

### Backend Selection: rookiepy vs browser_cookie3

Agent-Reach implements a failover mechanism for cookie extraction backends. It first attempts to import **rookiepy**, a Rust-based library that provides more stable and faster access to browser cookie stores. If `rookiepy` is not available, the code automatically falls back to the pure-Python **browser_cookie3** package. This logic appears in lines 55-64, where the import attempt determines which backend module will be used for the remainder of the session.

### Reading the Raw Cookie Jar

Once the backend is selected, browser-specific helpers (`rookiepy.chrome`, `browser_cookie3.chrome`, etc.) return an iterable of cookie objects. Each object exposes at least `name`, `value`, and `domain` attributes. The code in lines 78-108 handles the browser-specific invocation and standardizes the output into a common format for downstream processing.

### Platform Matching and Filtering

The core filtering logic relies on **`PLATFORM_SPECS`**, a list defined in lines 15-41 that describes each supported platform (Twitter/X, XiaoHongShu, Bilibili, Xueqiu) with:
- **Domain patterns**: Suffix matches for cookie domains (e.g., `.twitter.com`)
- **Required cookies**: Specific key names required for authentication, or `null` to collect all domain cookies

The extraction process walks through every cookie in the jar (lines 22-28) and checks if its domain ends with any platform's domain patterns. If a platform defines required cookie names, only those specific keys are collected; otherwise, all cookies for the domain are joined into a single header-style string (lines 33-44).

### Building the Result Dictionary

For every platform that yields matching cookies, the code constructs a sub-dictionary stored under the platform's `config_key`. For example, Twitter cookies populate `{"twitter": {"auth_token": "...", "ct0": "..."}}` as implemented in lines 45-47. This structured dictionary separates credentials by platform while maintaining the specific key-value pairs required for API authentication.

### Legacy Credential Synchronization

When Twitter/X cookies are successfully extracted, two helper functions maintain backward compatibility with legacy tools:

- **`_sync_xfetch_session`** writes the tokens to `~/.config/xfetch/session.json` for compatibility with existing xfetch installations
- **`_sync_bird_env`** creates a source-able environment file at `~/.config/bird/credentials.env` containing `AUTH_TOKEN` and `CT0` variables for the `bird` CLI

These synchronization steps appear in lines 199-205, ensuring that credentials extracted by Agent-Reach remain accessible to other tools in the ecosystem.

## Configuration Persistence

The `configure_from_browser` function orchestrates the complete flow. It calls `extract_all` to retrieve the filtered cookies, then writes the extracted values into the central `Config` object using platform-specific keys (e.g., `config.set("twitter_auth_token", ...)`). Lines 25-31 and 44-61 handle this persistence, returning a status list that indicates which platforms succeeded—enabling user-friendly CLI feedback showing exactly which services were configured.

## Usage Examples

### Command Line Configuration

Extract cookies from Chrome and configure all supported platforms automatically:

```bash
python -m agent_reach.cli configure --from-browser chrome

```

This command invokes `configure_from_browser("chrome", config)` internally.

### Programmatic Extraction

Access the extraction functions directly in Python:

```python
from agent_reach.cookie_extract import extract_all, configure_from_browser
from agent_reach.config import Config

# Extract raw cookies from Firefox

cookies = extract_all("firefox")
print(cookies)

# → {'twitter': {'auth_token': '...', 'ct0': '...'}, 

#    'xhs': {'cookie_string': '...'}, 

#    'bilibili': {...}, 

#    'xueqiu': {...}}

# Apply to configuration

cfg = Config()
status = configure_from_browser("firefox", cfg)
print(status)

# → [('Twitter/X', True, 'auth_token + ct0'), 

#    ('XiaoHongShu', True, '12 cookies'), ...]

# Access saved values

print(cfg.get("twitter_auth_token"))

```

### Extending PLATFORM_SPECS

Add support for new platforms by extending the specification list without modifying extraction logic:

```python

# In agent_reach/cookie_extract.py

PLATFORM_SPECS.append({
    "name": "MySite",
    "domains": [".mysite.com"],
    "cookies": ["mysess", "mycsrf"],
    "config_key": "mysite",
})

```

The next `extract_all` call automatically picks up the new specification.

## Summary

- **Dual-backend extraction**: Agent-Reach prefers `rookiepy` for performance but falls back to `browser_cookie3` for compatibility
- **Platform-agnostic filtering**: The `PLATFORM_SPECS` configuration separates extraction logic from platform-specific requirements
- **Zero manual intervention**: The process reads directly from browser storage at `~/.config/` locations, eliminating manual cookie copying
- **Legacy support**: Extracted Twitter credentials automatically sync to `xfetch` and `bird` configuration files
- **Extensible architecture**: New platforms require only adding entries to `PLATFORM_SPECS` without changing extraction code

## Frequently Asked Questions

### Which browsers does Agent-Reach support for cookie extraction?

Agent-Reach supports Chrome, Firefox, Edge, Brave, and Opera. The validation logic in [`agent_reach/cookie_extract.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/cookie_extract.py) lines 70-76 normalizes browser names and validates them against this supported list before attempting extraction.

### What is the difference between rookiepy and browser_cookie3?

**rookiepy** is a Rust-based library that provides more stable and faster access to browser cookie stores, while **browser_cookie3** is a pure-Python implementation used as a fallback. The code attempts to import `rookiepy` first (lines 55-64) and only uses `browser_cookie3` if the Rust library is unavailable, ensuring maximum compatibility across different deployment environments.

### How does Agent-Reach handle platform-specific cookie requirements?

Each platform defines its requirements in `PLATFORM_SPECS`, including domain patterns and required cookie keys. The extraction process filters the raw cookie jar against these specifications (lines 22-44). If a platform lists specific cookie names, only those are extracted; otherwise, all cookies for the matching domain are collected into a single string.

### Where are extracted credentials stored after extraction?

Credentials are stored in three locations: the central Agent-Reach `Config` object (for immediate use), platform-specific configuration files (e.g., `~/.config/xfetch/session.json` for Twitter), and environment files (e.g., `~/.config/bird/credentials.env`). This multi-location storage ensures compatibility with both the Agent-Reach framework and standalone CLI tools.