How to Maintain Browser Sessions and Cookies Across Requests in crawl4ai

To maintain browser sessions and cookies across requests in crawl4ai, configure a persistent user_data_dir in BrowserConfig to save the Chromium profile to disk, and reuse the same session_id in CrawlerRunConfig to keep browser contexts alive across multiple crawl operations.

Maintaining state between HTTP requests is essential when scraping authenticated web applications or multi-page workflows. In unclecode/crawl4ai, session persistence is handled through a combination of persistent browser profiles and in-memory session caching, allowing cookies and local storage to survive across independent crawl jobs. This article explains the four core mechanisms that enable you to maintain browser sessions and cookies across requests according to the actual source implementation.

The Architecture of Session Persistence

Crawl4AI implements session continuity through four coordinated systems defined in crawl4ai/browser_manager.py and crawl4ai/managed_browser.py:

  • Persistent User Data Directories: The ManagedBrowser class launches Chromium with a --user-data-dir argument that preserves cookies, cache, and localStorage on disk.
  • Session ID Mapping: The BrowserManager maintains an internal self.sessions dictionary that maps session_id strings to active browser contexts and pages.
  • Automatic Cookie Synchronization: The setup_context method injects cookies from BrowserConfig.cookies into new contexts.
  • Runtime State Cloning: The clone_runtime_state helper copies cookies and localStorage between contexts for deep-crawl resumption.

Method 1: Persistent Browser Profiles with user_data_dir

The most reliable way to maintain cookies across separate Python runs is to specify a persistent profile directory. In crawl4ai/managed_browser.py, the ManagedBrowser.__init__ method (around lines 49-57) accepts a user_data_dir parameter that gets passed directly to the Chromium launch arguments.

When ManagedBrowser.start is called, it only creates a temporary directory if user_data_dir is None. Otherwise, it reuses the existing profile folder, preserving all browser state including cookies and localStorage.

from crawl4ai import Crawl4AI, BrowserConfig

browser_cfg = BrowserConfig(
    user_data_dir="/tmp/crawl4ai_session",  # Persistent profile location

    headless=False
)

crawler = Crawl4AI(browser_config=browser_cfg)

Method 2: Session Caching via session_id

For maintaining state across multiple run() calls within the same Python process, use the session_id parameter in CrawlerRunConfig. The BrowserManager.__init__ method (around lines 10807-10810) initializes self.sessions as a dictionary and sets self.session_ttl to 30 minutes.

When you call crawler.run() with a session_id, the BrowserManager.get_page method checks self.sessions for an existing context. If found, it returns the cached page instead of creating a new one, preserving all cookies and JavaScript state.

from crawl4ai import Crawl4AI, CrawlerRunConfig, BrowserConfig

browser_cfg = BrowserConfig(headless=True)
crawler = Crawl4AI(browser_config=browser_cfg)

# First request: creates the session

result1 = await crawler.run(CrawlerRunConfig(
    url="https://example.com/login",
    session_id="auth_session"
))

# Second request: reuses the same browser context and cookies

result2 = await crawler.run(CrawlerRunConfig(
    url="https://example.com/dashboard",
    session_id="auth_session"
))

Idle sessions are automatically purged by _cleanup_expired_sessions when their age exceeds session_ttl (30 minutes).

Method 3: Injecting Custom Cookies

The BrowserManager.setup_context method (around lines 10871-10876) automatically propagates any cookies defined in BrowserConfig.cookies to new browser contexts. This is useful for adding authentication tokens or preferences before navigation.

custom_cookies = [
    {"name": "auth_token", "value": "xyz123", "url": "https://example.com"},
    {"name": "theme", "value": "dark", "url": "https://example.com"}
]

browser_cfg = BrowserConfig(
    cookies=custom_cookies,
    headless=True
)

Method 4: Cloning Runtime State

When you need to create a fresh browser context but retain exact state from a previous session, use the clone_runtime_state function in crawl4ai/browser_manager.py (around lines 10122-10158). This utility copies cookies, localStorage, headers, and geolocation from a source context to a destination context.

This is particularly useful for resuming deep crawls or duplicating authenticated sessions across parallel workers.

from crawl4ai.browser_manager import clone_runtime_state

# Assuming src_ctx and dst_ctx are BrowserContext objects

await clone_runtime_state(
    src_ctx, 
    dst_ctx, 
    crawlerRunConfig=run_cfg, 
    browserConfig=browser_cfg
)

Summary

  • Use user_data_dir in BrowserConfig to persist cookies to disk across Python processes via ManagedBrowser in crawl4ai/managed_browser.py.
  • Use session_id in CrawlerRunConfig to cache live browser contexts in memory via BrowserManager.sessions, with automatic TTL cleanup after 30 minutes.
  • Set BrowserConfig.cookies to inject specific cookies on context creation in BrowserManager.setup_context (lines 10871-10876).
  • Call clone_runtime_state to copy cookies and localStorage between contexts when resuming complex crawl workflows (lines 10122-10158).

Frequently Asked Questions

How long do sessions stay cached in BrowserManager?

Sessions remain cached for 30 minutes by default, controlled by self.session_ttl in BrowserManager.__init__. The _cleanup_expired_sessions method runs automatically before each new page request and removes entries older than this threshold to prevent memory leaks.

Can I use multiple different sessions simultaneously?

Yes. The self.sessions dictionary in BrowserManager supports multiple concurrent session IDs. Each unique session_id maintains its own isolated browser context, cookies, and localStorage, allowing you to crawl as different users concurrently within the same process.

Where are cookies physically stored when using user_data_dir?

When BrowserConfig.user_data_dir is set, Chromium writes its SQLite cookie database and localStorage files to that directory path. According to the ManagedBrowser.start implementation in crawl4ai/managed_browser.py, this directory persists between program runs unless explicitly deleted.

How do I clear cookies between requests?

To clear cookies, either omit the session_id to create a fresh anonymous context, or use a new session_id to start a separate session. If using user_data_dir, delete the profile directory or point to a different path to ensure a clean state.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →