How to Handle Cookies and Simulate Authenticated Sessions in Crawl4AI
Crawl4AI provides three primary mechanisms for handling cookies and simulating authenticated sessions: static cookie injection via BrowserConfig.cookies, dynamic cookie management through the on_page_context_created hook using context.add_cookies(), and session reuse across crawls via clone_runtime_state().
When scraping protected content or maintaining state across multiple pages, managing authentication is critical. Crawl4AI, which leverages Playwright's browser automation capabilities, offers flexible approaches to handle cookies and simulate authenticated sessions without requiring external tools. The library stores cookies in the browser's native storage state, ensuring they persist through navigation, redirects, and subsequent HTTP requests.
Static Cookie Injection with BrowserConfig
The simplest approach to handle cookies in Crawl4AI involves passing pre-defined cookies directly to the browser configuration. In crawl4ai/async_configs.py, the BrowserConfig class accepts a cookies parameter—a list of cookie dictionaries that Playwright automatically applies when creating a new browser context.
This method works best when you already possess session tokens or authentication cookies from a previous login, or when testing with known static values.
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
# Define cookies obtained from previous authentication
my_cookies = [
{
"name": "session_id",
"value": "abcd1234",
"domain": ".example.com",
"path": "/",
},
{
"name": "auth_token",
"value": "jwt-token-goes-here",
"domain": ".example.com",
"path": "/",
},
]
# Inject cookies via BrowserConfig
browser_cfg = BrowserConfig(
headless=True,
cookies=my_cookies,
)
crawler = AsyncWebCrawler(config=browser_cfg)
await crawler.start()
run_cfg = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS,
)
result = await crawler.arun("https://example.com/dashboard", config=run_cfg)
print(result.html) # Content returned as authenticated user
await crawler.close()
Dynamic Authentication Using Hooks
For scenarios requiring interactive login flows—such as submitting credentials through a form—Crawl4AI exposes the on_page_context_created hook. This hook receives the Playwright BrowserContext and Page instances immediately after creation, allowing you to execute login sequences and then capture or inject cookies using context.add_cookies().
The implementation in crawl4ai/browser_manager.py handles the underlying add_cookies calls, while practical examples appear in docs/examples/hooks_example.py.
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig
from playwright.async_api import BrowserContext, Page
async def on_page_context_created(page: Page, context: BrowserContext, **kwargs):
# Navigate to login page
await page.goto("https://example.com/login")
# Fill credentials (adjust selectors to match target site)
await page.fill('input[name="user"]', "my_user")
await page.fill('input[name="pass"]', "my_pass")
await page.click('button[type="submit"]')
# Wait for post-login indicator
await page.wait_for_selector('.profile')
# Capture cookies set by server after authentication
cookies = await context.cookies()
print("Authenticated cookies:", cookies)
return page
browser_cfg = BrowserConfig(headless=False) # Set visible for debugging
crawler = AsyncWebCrawler(config=browser_cfg)
# Register the authentication hook
crawler.crawler_strategy.set_hook("on_page_context_created", on_page_context_created)
await crawler.start()
result = await crawler.arun("https://example.com/protected")
print("Protected content length:", len(result.html))
await crawler.close()
Reusing Authenticated Sessions Across Crawls
When running multiple crawler instances or sequential jobs that require the same authenticated state, re-logging in for each session wastes resources and risks rate-limiting. Crawl4AI solves this through clone_runtime_state(), implemented in crawl4ai/browser_manager.py, which copies cookies and optionally localStorage from a source BrowserContext to a destination context.
This utility enables you to authenticate once, then propagate that session to subsequent crawlers without repeating the login flow.
from crawl4ai import AsyncWebCrawler, BrowserConfig
from crawl4ai.browser_manager import clone_runtime_state
# First crawler: establish authenticated session
auth_browser_cfg = BrowserConfig(headless=True)
auth_crawler = AsyncWebCrawler(config=auth_browser_cfg)
await auth_crawler.start()
# Assume authentication happened via hook or pre-set cookies
src_context = await auth_crawler.browser_manager.default_context
# Second crawler: new instance that needs same session
new_browser_cfg = BrowserConfig(headless=True)
new_crawler = AsyncWebCrawler(config=new_browser_cfg)
await new_crawler.start()
dst_context = await new_crawler.browser_manager.default_context
# Transfer cookies and storage state
await clone_runtime_state(src_context, dst_context)
# Now new_crawler can access protected resources
result = await new_crawler.arun("https://example.com/secret")
print("Secret page content:", result.html[:200])
await new_crawler.close()
await auth_crawler.close()
Key Implementation Files
Understanding where these mechanisms reside helps when debugging or extending functionality:
| File | Purpose |
|---|---|
crawl4ai/async_configs.py |
Defines BrowserConfig including the cookies parameter for static injection. |
crawl4ai/browser_manager.py |
Implements clone_runtime_state() and manages add_cookies calls for context manipulation. |
docs/examples/hooks_example.py |
Demonstrates practical usage of on_page_context_created for dynamic authentication flows. |
tests/docker/test_hooks_utility.py |
Contains unit tests verifying cookie handling within the hook system. |
Summary
- Static authentication works by passing cookie dictionaries to
BrowserConfig.cookiesbefore starting the crawler, ideal for pre-existing session tokens. - Dynamic authentication leverages the
on_page_context_createdhook to execute login flows and inject cookies viacontext.add_cookies()after page creation. - Session persistence across multiple crawler instances uses
clone_runtime_state()fromcrawl4ai/browser_manager.pyto copy cookies and storage between browser contexts. - All methods rely on Playwright's native cookie storage, ensuring compatibility with modern web authentication mechanisms including JWT tokens and session IDs.
Frequently Asked Questions
How do I handle cookies in Crawl4AI?
You can handle cookies in Crawl4AI either statically by passing a list of cookie dictionaries to BrowserConfig.cookies before starting the crawler, or dynamically by using the on_page_context_created hook to call context.add_cookies() after performing a login flow. Both methods store cookies in Playwright's native storage state, making them available for all subsequent requests.
Can I simulate a login flow before crawling protected content?
Yes, simulate authentication by registering an on_page_context_created hook that receives the Page and BrowserContext objects. Inside the hook, navigate to the login page, fill credentials using Playwright's API (e.g., page.fill() and page.click()), wait for the post-login state, and optionally capture the cookies for reuse. This approach is demonstrated in docs/examples/hooks_example.py.
How do I reuse an authenticated session across multiple Crawl4AI instances?
To reuse sessions, use the clone_runtime_state() function from crawl4ai/browser_manager.py after authenticating with your first crawler instance. This utility copies cookies and optionally localStorage from the source BrowserContext to a destination context created by a new AsyncWebCrawler instance, eliminating the need to re-authenticate for subsequent crawls.
Where are cookies stored in Crawl4AI?
Cookies are stored in Playwright's native browser storage state within each BrowserContext. When you provide cookies via BrowserConfig or inject them through hooks, Crawl4AI applies them to the Playwright context using context.add_cookies(). This storage persists across page navigations and HTTP requests within the same context, and can be transferred between contexts using clone_runtime_state() as implemented in crawl4ai/browser_manager.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →