How to Manage Browser Contexts and Prevent Memory Leaks in MediaCrawler
MediaCrawler prevents memory leaks by using isolated incognito browser contexts with explicit lifecycle management, ensuring each crawl job creates, uses, and properly disposes of its CDP session and context cleanup.
MediaCrawler (NanmiCoder/MediaCrawler) is a high-throughput media extraction framework that automates Chromium browsers via the Chrome DevTools Protocol. Managing browser contexts correctly is critical—each leaked context accumulates orphaned Chrome processes, open WebSockets, and retained JavaScript objects that can exhaust system resources within hours of continuous operation.
Understanding Browser Context Architecture in MediaCrawler
The codebase implements a three-layer browser management stack:
BrowserLauncher(tools/browser_launcher.py) – singleton Chromium process managementCDPBrowser(tools/cdp_browser.py) – CDP session wrapper with context factoryBaseCrawler(base/base_crawler.py) – per-job lifecycle orchestration
This separation ensures that expensive Chromium launches are reused across thousands of crawl jobs while keeping each job's browsing data strictly isolated.
Creating Isolated Contexts Per Crawl Job
MediaCrawler uses incognito browser contexts to guarantee complete state isolation. Unlike persistent contexts, incognito contexts discard all cookies, localStorage, IndexedDB, and service worker state upon closure.
In tools/cdp_browser.py, the context creation follows this pattern:
class CDPBrowser:
async def new_context(self):
# Create an isolated incognito context for a single crawl
self.context = await self.browser.create_incognito_browser_context()
self.page = await self.context.new_page()
return self.page
Each call to create_incognito_browser_context() spins up a lightweight browsing partition without the overhead of launching a new Chromium instance.
Enforcing Cleanup with Finally-Block Guarantees
The BaseCrawler class implements defensive cleanup patterns to prevent leaks during exception scenarios. The run() method in base/base_crawler.py wraps all crawl logic in a try/finally block:
async def run(self, *args, **kwargs):
try:
page = await self.browser.new_context()
await self.perform_crawl(page, *args, **kwargs)
finally:
# Guarantees cleanup even on error
await self.browser.close_context()
This pattern ensures that close_context() executes regardless of whether the crawl succeeds, times out, or raises an unhandled exception.
Implementing Proper Context and Session Disposal
Complete resource liberation requires two discrete cleanup operations. In tools/cdp_browser.py, the close_context() method handles both:
async def close_context(self):
# Close pages and release the context
await self.context.close()
Additionally, the CDP wrapper calls session.detach() and session.close() to release the underlying WebSocket connection back to the browser's connection pool.
Graceful Launcher Shutdown with Process Termination
The BrowserLauncher class in tools/browser_launcher.py manages the singleton Chromium process:
from playwright.async_api import async_playwright
class BrowserLauncher:
async def launch(self):
self.playwright = await async_playwright().start()
self.browser = await self.playwright.chromium.launch(headless=True)
async def close(self):
await self.browser.close()
await self.playwright.stop()
The launcher registers shutdown() with Python's atexit handler, guaranteeing that the Chromium process is terminated even if the interpreter exits abnormally due to SIGTERM or unhandled exceptions.
Monitoring Resource Limits and Global Cleanup
The caching layer provides context-aware resource monitoring. cache/redis_cache.py and cache/local_cache.py track active context counts and trigger threshold-based cleanup operations when memory pressure exceeds configured limits.
The cache_factory.py exposes a unified interface for this monitoring, while api/services/crawler_manager.py orchestrates multiple concurrent crawlers—each receiving its own isolated context from the pool without risk of cross-contamination.
Key Implementation Files
| File | Purpose |
|---|---|
tools/browser_launcher.py |
Singleton Chromium launch and graceful process termination |
tools/cdp_browser.py |
CDP session management, incognito context creation, and disposal |
base/base_crawler.py |
Base crawler enforcing finally-block context cleanup |
cache/cache_factory.py |
Unified cache layer with context monitoring |
cache/redis_cache.py / cache/local_cache.py |
Resource threshold tracking and global cleanup triggers |
api/services/crawler_manager.py |
Multi-crawler orchestration with per-job context allocation |
Common Anti-Patterns to Avoid
Several leak-prone practices are explicitly not used in MediaCrawler's implementation:
- Reusing persistent contexts between unrelated crawl jobs (retains cookies/cache)
- Omitting
awaiton close operations (creates dangling coroutines) - Creating pages without closing their parent context (leaves partition allocated)
- Launching Chromium per job (prohibitive startup cost, harder process management)
Summary
- Incognito contexts provide complete isolation without persistent state accumulation
- Try/finally blocks in
BaseCrawler.run()guarantee cleanup on all exit paths - Dual disposal (context close + session detach) releases all held resources
- Singleton launcher with atexit registration prevents orphaned Chrome processes
- Cache-layer monitoring enables proactive cleanup under resource pressure
Frequently Asked Questions
What causes browser memory leaks in web crawling?
Memory leaks typically stem from unclosed pages, persistent contexts retaining cookies and cache, or dangling WebSocket connections to the CDP. Each leaked context maintains its own JavaScript heap, network socket pool, and storage partitions—multiplied across thousands of jobs, this exhausts system memory and file descriptors.
How does Playwright's incognito context differ from a new browser instance?
create_incognito_browser_context() creates a lightweight browsing partition within an existing Chromium process, sharing the browser's core engine while isolating storage, cookies, and permissions. This avoids the ~100MB+ startup cost and process overhead of launching separate Chromium instances per job.
Should I call browser.close() or context.close() in MediaCrawler?
Call context.close() after each crawl job to release that partition, and reserve browser.close() for application shutdown handled by BrowserLauncher.close(). Closing the browser terminates all contexts and should only occur when the crawler pool is being destroyed.
How does MediaCrawler handle crashes where cleanup code doesn't run?
The atexit-registered BrowserLauncher.shutdown() provides a final safety net. While it cannot execute if the process receives SIGKILL, it handles SIGTERM and normal interpreter exits. For maximum resilience, containerized deployments should combine this with process supervisors that reclaim zombie Chrome processes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →