How to Maintain Browser Sessions and Cookies Across Requests in crawl4ai
To maintain browser sessions and cookies across requests in crawl4ai, configure a persistent user_data_dir in BrowserConfig to save the Chromium profile to disk, and reuse the same session_id in CrawlerRunConfig to keep browser contexts alive across multiple crawl operations.
Maintaining state between HTTP requests is essential when scraping authenticated web applications or multi-page workflows. In unclecode/crawl4ai, session persistence is handled through a combination of persistent browser profiles and in-memory session caching, allowing cookies and local storage to survive across independent crawl jobs. This article explains the four core mechanisms that enable you to maintain browser sessions and cookies across requests according to the actual source implementation.
The Architecture of Session Persistence
Crawl4AI implements session continuity through four coordinated systems defined in crawl4ai/browser_manager.py and crawl4ai/managed_browser.py:
- Persistent User Data Directories: The
ManagedBrowserclass launches Chromium with a--user-data-dirargument that preserves cookies, cache, and localStorage on disk. - Session ID Mapping: The
BrowserManagermaintains an internalself.sessionsdictionary that mapssession_idstrings to active browser contexts and pages. - Automatic Cookie Synchronization: The
setup_contextmethod injects cookies fromBrowserConfig.cookiesinto new contexts. - Runtime State Cloning: The
clone_runtime_statehelper copies cookies and localStorage between contexts for deep-crawl resumption.
Method 1: Persistent Browser Profiles with user_data_dir
The most reliable way to maintain cookies across separate Python runs is to specify a persistent profile directory. In crawl4ai/managed_browser.py, the ManagedBrowser.__init__ method (around lines 49-57) accepts a user_data_dir parameter that gets passed directly to the Chromium launch arguments.
When ManagedBrowser.start is called, it only creates a temporary directory if user_data_dir is None. Otherwise, it reuses the existing profile folder, preserving all browser state including cookies and localStorage.
from crawl4ai import Crawl4AI, BrowserConfig
browser_cfg = BrowserConfig(
user_data_dir="/tmp/crawl4ai_session", # Persistent profile location
headless=False
)
crawler = Crawl4AI(browser_config=browser_cfg)
Method 2: Session Caching via session_id
For maintaining state across multiple run() calls within the same Python process, use the session_id parameter in CrawlerRunConfig. The BrowserManager.__init__ method (around lines 10807-10810) initializes self.sessions as a dictionary and sets self.session_ttl to 30 minutes.
When you call crawler.run() with a session_id, the BrowserManager.get_page method checks self.sessions for an existing context. If found, it returns the cached page instead of creating a new one, preserving all cookies and JavaScript state.
from crawl4ai import Crawl4AI, CrawlerRunConfig, BrowserConfig
browser_cfg = BrowserConfig(headless=True)
crawler = Crawl4AI(browser_config=browser_cfg)
# First request: creates the session
result1 = await crawler.run(CrawlerRunConfig(
url="https://example.com/login",
session_id="auth_session"
))
# Second request: reuses the same browser context and cookies
result2 = await crawler.run(CrawlerRunConfig(
url="https://example.com/dashboard",
session_id="auth_session"
))
Idle sessions are automatically purged by _cleanup_expired_sessions when their age exceeds session_ttl (30 minutes).
Method 3: Injecting Custom Cookies
The BrowserManager.setup_context method (around lines 10871-10876) automatically propagates any cookies defined in BrowserConfig.cookies to new browser contexts. This is useful for adding authentication tokens or preferences before navigation.
custom_cookies = [
{"name": "auth_token", "value": "xyz123", "url": "https://example.com"},
{"name": "theme", "value": "dark", "url": "https://example.com"}
]
browser_cfg = BrowserConfig(
cookies=custom_cookies,
headless=True
)
Method 4: Cloning Runtime State
When you need to create a fresh browser context but retain exact state from a previous session, use the clone_runtime_state function in crawl4ai/browser_manager.py (around lines 10122-10158). This utility copies cookies, localStorage, headers, and geolocation from a source context to a destination context.
This is particularly useful for resuming deep crawls or duplicating authenticated sessions across parallel workers.
from crawl4ai.browser_manager import clone_runtime_state
# Assuming src_ctx and dst_ctx are BrowserContext objects
await clone_runtime_state(
src_ctx,
dst_ctx,
crawlerRunConfig=run_cfg,
browserConfig=browser_cfg
)
Summary
- Use
user_data_dirinBrowserConfigto persist cookies to disk across Python processes viaManagedBrowserincrawl4ai/managed_browser.py. - Use
session_idinCrawlerRunConfigto cache live browser contexts in memory viaBrowserManager.sessions, with automatic TTL cleanup after 30 minutes. - Set
BrowserConfig.cookiesto inject specific cookies on context creation inBrowserManager.setup_context(lines 10871-10876). - Call
clone_runtime_stateto copy cookies and localStorage between contexts when resuming complex crawl workflows (lines 10122-10158).
Frequently Asked Questions
How long do sessions stay cached in BrowserManager?
Sessions remain cached for 30 minutes by default, controlled by self.session_ttl in BrowserManager.__init__. The _cleanup_expired_sessions method runs automatically before each new page request and removes entries older than this threshold to prevent memory leaks.
Can I use multiple different sessions simultaneously?
Yes. The self.sessions dictionary in BrowserManager supports multiple concurrent session IDs. Each unique session_id maintains its own isolated browser context, cookies, and localStorage, allowing you to crawl as different users concurrently within the same process.
Where are cookies physically stored when using user_data_dir?
When BrowserConfig.user_data_dir is set, Chromium writes its SQLite cookie database and localStorage files to that directory path. According to the ManagedBrowser.start implementation in crawl4ai/managed_browser.py, this directory persists between program runs unless explicitly deleted.
How do I clear cookies between requests?
To clear cookies, either omit the session_id to create a fresh anonymous context, or use a new session_id to start a separate session. If using user_data_dir, delete the profile directory or point to a different path to ensure a clean state.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →