How to Set Up Persistent Browser Profiles with Authentication in Crawl4AI
Use BrowserProfiler.create_profile() to generate a persistent browser profile in ~/.crawl4ai/profiles/, authenticate manually or programmatically, then pass the profile path to BrowserConfig.user_data_dir to maintain authenticated sessions across multiple crawl jobs.
Crawl4AI is an open-source asynchronous web crawling framework that supports persistent browser profiles to preserve cookies, local storage, and session data between runs. Setting up persistent browser profiles with authentication in Crawl4AI allows you to perform a one-time login and reuse those credentials across unlimited subsequent scraping operations without re-authenticating.
Creating a Persistent Browser Profile Interactively
The BrowserProfiler class in crawl4ai/browser_profiler.py provides a high-level API for managing browser profiles. The create_profile() coroutine launches a visible browser window, waits for you to complete authentication or any other UI interactions, and then persists the profile directory to disk.
When you invoke create_profile(), the method:
- Creates a profile folder under
~/.crawl4ai/profiles/(computed viaget_home_folder()incrawl4ai/utils.py) - Starts a managed browser instance attached to that profile
- Prints the CDP (Chrome DevTools Protocol) URL for external connections
- Blocks until you press
qin the terminal, then automatically saves the profile state
Manual Authentication Workflow
Run the following script to generate a profile and log in manually:
import asyncio
from crawl4ai.browser_profiler import BrowserProfiler
async def make_profile():
profiler = BrowserProfiler()
# profile_name is optional; a timestamped name is generated if omitted
profile_path = await profiler.create_profile(profile_name="my-login-profile")
print(f"Profile stored at: {profile_path}")
if __name__ == "__main__":
asyncio.run(make_profile())
Execute the script, complete the login flow in the opened browser window, then press q in your terminal. The profile—including cookies, local storage, and session tokens—is saved to ~/.crawl4ai/profiles/my-login-profile.
Reusing Profiles for Authenticated Crawling
To leverage the saved authentication state, pass the profile directory to BrowserConfig.user_data_dir. The BrowserManager (implemented in crawl4ai/browser_manager.py) injects this path into the browser launch flags via ManagedBrowser._get_browser_args(), appending --user-data-dir= to the Chromium command line.
import asyncio
from crawl4ai.async_webcrawler import AsyncWebCrawler
from crawl4ai.async_configs import BrowserConfig, CrawlerRunConfig
async def crawl_with_profile():
browser_cfg = BrowserConfig(
browser_type="chromium",
headless=False, # Set True only after verifying authentication works
user_data_dir="~/.crawl4ai/profiles/my-login-profile",
verbose=True,
)
crawler_cfg = CrawlerRunConfig()
async with AsyncWebCrawler(config=browser_cfg) as crawler:
result = await crawler.arun(
url="https://example.com/protected-page",
config=crawler_cfg,
)
print(result.markdown.raw_markdown)
if __name__ == "__main__":
asyncio.run(crawl_with_profile())
Because the crawler attaches to the existing profile, the server recognizes the authenticated session from previous cookies, granting access to protected resources without additional login steps.
Automating Login Programmatically
For scenarios requiring automated authentication without manual interaction, connect to the profile using Playwright over the CDP URL printed during profile creation. This allows you to script form filling while still persisting the resulting session to disk.
import asyncio
from crawl4ai.browser_profiler import BrowserProfiler
from crawl4ai.async_configs import BrowserConfig
from crawl4ai.async_webcrawler import AsyncWebCrawler
async def login_and_crawl():
profiler = BrowserProfiler()
profile_path = await profiler.create_profile(profile_name="auto-login")
# Programmatic login via Playwright
from playwright.async_api import async_playwright
async with async_playwright() as p:
browser = await p.chromium.connect_over_cdp("http://localhost:9222")
context = browser.contexts[0]
page = await context.new_page()
await page.goto("https://example.com/login")
await page.fill('input[name="username"]', "my_user")
await page.fill('input[name="password"]', "my_secret")
await page.click('button[type="submit"]')
await page.wait_for_load_state("networkidle")
await page.close()
await browser.close()
# Reuse authenticated profile
browser_cfg = BrowserConfig(
browser_type="chromium",
headless=False,
user_data_dir=profile_path,
verbose=True,
)
async with AsyncWebCrawler(config=browser_cfg) as crawler:
result = await crawler.arun("https://example.com/secure-area")
print(result.markdown.raw_markdown)
if __name__ == "__main__":
asyncio.run(login_and_crawl())
This approach combines the durability of persistent profiles with the efficiency of automated authentication workflows.
Running a Standalone Browser for Long-Lived Sessions
For high-throughput scenarios requiring multiple crawls over an extended period, use launch_standalone_browser() (also in crawl4ai/browser_profiler.py). This starts a persistent browser instance that remains active until manually terminated, returning a CDP URL that multiple crawler instances can reuse.
import asyncio
from crawl4ai.browser_profiler import BrowserProfiler
from crawl4ai.async_configs import BrowserConfig
from crawl4ai.async_webcrawler import AsyncWebCrawler
async def long_running():
profiler = BrowserProfiler()
cdp_url = await profiler.launch_standalone_browser(
user_data_dir="~/.crawl4ai/profiles/long-run",
headless=False,
)
# Browser remains alive until 'q' is pressed in terminal
browser_cfg = BrowserConfig(
browser_type="chromium",
headless=False,
cdp_url=cdp_url,
verbose=True,
)
async with AsyncWebCrawler(config=browser_cfg) as crawler:
for url in ["https://site1.com", "https://site2.com"]:
res = await crawler.arun(url)
print(res.url, "→", len(res.markdown.raw_markdown), "chars")
if __name__ == "__main__":
asyncio.run(long_running())
Sharing a single browser instance across crawls eliminates the overhead of repeated browser launches while maintaining the authenticated state stored in the profile directory.
Summary
- Persistent profiles store cookies, local storage, and session data in
~/.crawl4ai/profiles/via theBrowserProfilerclass incrawl4ai/browser_profiler.py. - Interactive creation uses
create_profile()to launch a visible browser, capture manual login actions, and persist state when you pressq. - Profile reuse requires passing the profile path to
BrowserConfig.user_data_dir, whichManagedBrowser(incrawl4ai/browser_manager.py) injects into the--user-data-dir=launch flag. - Programmatic authentication connects to the profile via Playwright over the CDP URL returned during creation, allowing scripted logins while preserving the session.
- Long-running sessions use
launch_standalone_browser()to start a persistent CDP-enabled browser that multiple crawler instances can share, minimizing startup overhead.
Frequently Asked Questions
How do I update an existing persistent profile with new cookies?
Run the create_profile() method again with the same profile_name argument. The BrowserProfiler loads the existing profile directory, launches the browser with that state intact, and overwrites the profile with any new cookies or storage changes when you press q.
Can I use persistent profiles with headless mode?
Yes, but only after the initial authentication is complete. Create the profile with headless=False to perform the login visually, then set headless=True in BrowserConfig for subsequent crawls. The authenticated session remains valid because the profile directory contains the saved cookies.
What is the difference between user_data_dir and cdp_url?
user_data_dir specifies a filesystem path to a Chrome profile directory (containing Cookies, Local Storage, etc.), which ManagedBrowser uses to launch a new browser instance. cdp_url connects to an already-running browser instance via the Chrome DevTools Protocol, useful for sharing a single browser across multiple crawler processes or long-running sessions.
Where does Crawl4AI store persistent profiles by default?
By default, profiles are stored in ~/.crawl4ai/profiles/. This path is computed by the get_home_folder() helper in crawl4ai/utils.py. You can override this location by passing an absolute path to the user_data_dir parameter in BrowserConfig.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →