How to Manage Browser Contexts and Prevent Memory Leaks in MediaCrawler

MediaCrawler prevents memory leaks by using isolated incognito browser contexts with explicit lifecycle management, ensuring each crawl job creates, uses, and properly disposes of its CDP session and context cleanup.

MediaCrawler (NanmiCoder/MediaCrawler) is a high-throughput media extraction framework that automates Chromium browsers via the Chrome DevTools Protocol. Managing browser contexts correctly is critical—each leaked context accumulates orphaned Chrome processes, open WebSockets, and retained JavaScript objects that can exhaust system resources within hours of continuous operation.

Understanding Browser Context Architecture in MediaCrawler

The codebase implements a three-layer browser management stack:

  1. BrowserLauncher (tools/browser_launcher.py) – singleton Chromium process management
  2. CDPBrowser (tools/cdp_browser.py) – CDP session wrapper with context factory
  3. BaseCrawler (base/base_crawler.py) – per-job lifecycle orchestration

This separation ensures that expensive Chromium launches are reused across thousands of crawl jobs while keeping each job's browsing data strictly isolated.

Creating Isolated Contexts Per Crawl Job

MediaCrawler uses incognito browser contexts to guarantee complete state isolation. Unlike persistent contexts, incognito contexts discard all cookies, localStorage, IndexedDB, and service worker state upon closure.

In tools/cdp_browser.py, the context creation follows this pattern:

class CDPBrowser:
    async def new_context(self):
        # Create an isolated incognito context for a single crawl

        self.context = await self.browser.create_incognito_browser_context()
        self.page = await self.context.new_page()
        return self.page

Each call to create_incognito_browser_context() spins up a lightweight browsing partition without the overhead of launching a new Chromium instance.

Enforcing Cleanup with Finally-Block Guarantees

The BaseCrawler class implements defensive cleanup patterns to prevent leaks during exception scenarios. The run() method in base/base_crawler.py wraps all crawl logic in a try/finally block:

async def run(self, *args, **kwargs):
    try:
        page = await self.browser.new_context()
        await self.perform_crawl(page, *args, **kwargs)
    finally:
        # Guarantees cleanup even on error

        await self.browser.close_context()

This pattern ensures that close_context() executes regardless of whether the crawl succeeds, times out, or raises an unhandled exception.

Implementing Proper Context and Session Disposal

Complete resource liberation requires two discrete cleanup operations. In tools/cdp_browser.py, the close_context() method handles both:

async def close_context(self):
    # Close pages and release the context

    await self.context.close()

Additionally, the CDP wrapper calls session.detach() and session.close() to release the underlying WebSocket connection back to the browser's connection pool.

Graceful Launcher Shutdown with Process Termination

The BrowserLauncher class in tools/browser_launcher.py manages the singleton Chromium process:

from playwright.async_api import async_playwright

class BrowserLauncher:
    async def launch(self):
        self.playwright = await async_playwright().start()
        self.browser = await self.playwright.chromium.launch(headless=True)

    async def close(self):
        await self.browser.close()
        await self.playwright.stop()

The launcher registers shutdown() with Python's atexit handler, guaranteeing that the Chromium process is terminated even if the interpreter exits abnormally due to SIGTERM or unhandled exceptions.

Monitoring Resource Limits and Global Cleanup

The caching layer provides context-aware resource monitoring. cache/redis_cache.py and cache/local_cache.py track active context counts and trigger threshold-based cleanup operations when memory pressure exceeds configured limits.

The cache_factory.py exposes a unified interface for this monitoring, while api/services/crawler_manager.py orchestrates multiple concurrent crawlers—each receiving its own isolated context from the pool without risk of cross-contamination.

Key Implementation Files

File Purpose
tools/browser_launcher.py Singleton Chromium launch and graceful process termination
tools/cdp_browser.py CDP session management, incognito context creation, and disposal
base/base_crawler.py Base crawler enforcing finally-block context cleanup
cache/cache_factory.py Unified cache layer with context monitoring
cache/redis_cache.py / cache/local_cache.py Resource threshold tracking and global cleanup triggers
api/services/crawler_manager.py Multi-crawler orchestration with per-job context allocation

Common Anti-Patterns to Avoid

Several leak-prone practices are explicitly not used in MediaCrawler's implementation:

  • Reusing persistent contexts between unrelated crawl jobs (retains cookies/cache)
  • Omitting await on close operations (creates dangling coroutines)
  • Creating pages without closing their parent context (leaves partition allocated)
  • Launching Chromium per job (prohibitive startup cost, harder process management)

Summary

  • Incognito contexts provide complete isolation without persistent state accumulation
  • Try/finally blocks in BaseCrawler.run() guarantee cleanup on all exit paths
  • Dual disposal (context close + session detach) releases all held resources
  • Singleton launcher with atexit registration prevents orphaned Chrome processes
  • Cache-layer monitoring enables proactive cleanup under resource pressure

Frequently Asked Questions

What causes browser memory leaks in web crawling?

Memory leaks typically stem from unclosed pages, persistent contexts retaining cookies and cache, or dangling WebSocket connections to the CDP. Each leaked context maintains its own JavaScript heap, network socket pool, and storage partitions—multiplied across thousands of jobs, this exhausts system memory and file descriptors.

How does Playwright's incognito context differ from a new browser instance?

create_incognito_browser_context() creates a lightweight browsing partition within an existing Chromium process, sharing the browser's core engine while isolating storage, cookies, and permissions. This avoids the ~100MB+ startup cost and process overhead of launching separate Chromium instances per job.

Should I call browser.close() or context.close() in MediaCrawler?

Call context.close() after each crawl job to release that partition, and reserve browser.close() for application shutdown handled by BrowserLauncher.close(). Closing the browser terminates all contexts and should only occur when the crawler pool is being destroyed.

How does MediaCrawler handle crashes where cleanup code doesn't run?

The atexit-registered BrowserLauncher.shutdown() provides a final safety net. While it cannot execute if the process receives SIGKILL, it handles SIGTERM and normal interpreter exits. For maximum resilience, containerized deployments should combine this with process supervisors that reclaim zombie Chrome processes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →