How to Implement Custom Hooks at Different Crawling Stages in crawl4ai

TLDR: You can inject custom logic into crawl4ai's crawling pipeline at four distinct stages—pre_crawl, pre_fetch, post_fetch, and post_crawl—either by passing a hooks dictionary to the run() method or by using the @register_hook decorator to register functions globally in the hooks_registry.

crawl4ai is an open-source asynchronous web crawler that exposes a flexible hook system for manipulating the crawling lifecycle. By implementing custom hooks at different crawling stages in crawl4ai, you can modify crawler configuration, inspect HTTP responses, or transform final results without subclassing the core Crawl4AI class. The system routes all hook calls through the _invoke_hooks method in crawl4ai/crawler.py, which executes callables in registration order.

Architecture of the Hook System

The hook implementation consists of two primary components: the global registry defined in crawl4ai/hooks/__init__.py and the invocation logic inside crawl4ai/crawler.py.

The registry is a simple dictionary mapping stage names to lists of callables:


# crawl4ai/hooks/__init__.py (lines 6-12)

hooks_registry: dict[str, list[callable]] = {}

def register_hook(stage: str):
    def decorator(func):
        hooks_registry.setdefault(stage, []).append(func)
        return func
    return decorator

The crawler invokes hooks via the _invoke_hooks method (lines 43-46 in crawl4ai/crawler.py), which iterates over the callables stored for a specific stage:

def _invoke_hooks(self, stage: str, context: Dict[str, Any]):
    for hook in self.hooks.get(stage, []):
        hook(self, context)

This method is called by four private lifecycle methods—_pre_crawl, _pre_fetch, _post_fetch, and _post_crawl—defined in crawl4ai/crawler.py (lines 47-62).

The Four Crawling Stages

Each hook receives the Crawl4AI instance and a context dictionary specific to the current stage:

  • pre_crawl: Executes once before any network activity. Context: {}
  • pre_fetch: Executes before each HTTP request. Context: {"url": url}
  • post_fetch: Executes after receiving the Playwright Response. Context: {"response": response}
  • post_crawl: Executes after the final result is constructed. Context: {"result": crawler.result}

Method 1: Global Registration with Decorators

For reusable hooks that apply across multiple crawler instances, use the register_hook decorator from crawl4ai/hooks/__init__.py. This appends your function to the global hooks_registry (line 10).

from crawl4ai import Crawl4AI
from crawl4ai.hooks import register_hook, clear_hooks

# Ensure a clean registry

clear_hooks()

@register_hook("pre_crawl")
def enforce_minimum_depth(crawler, context):
    """Force max_depth to be at least 2."""
    if crawler.max_depth < 2:
        crawler.max_depth = 2

@register_hook("post_fetch")
def log_status_code(crawler, context):
    """Capture HTTP status from the response."""
    response = context["response"]
    print(f"Status: {response.status} for {context.get('url')}")

crawler = Crawl4AI()
crawler.run(
    url="https://example.com",
    max_depth=1,  # Hook will override this to 2

    hooks={}      # Global hooks execute even with empty dict

)

The decorator preserves registration order by appending to a list, ensuring hooks execute sequentially as demonstrated in tests/docker/test_hooks_utility.py (lines 16-36).

Method 2: Per-Run Hook Dictionary

For single-use or context-specific logic, pass a hooks dictionary directly to the run() method. This dictionary maps stage names to lists of callables, as shown in crawl4ai/crawler.py (lines 29-39).

from crawl4ai import Crawl4AI

def add_auth_header(crawler, context):
    """Inject custom header before the request."""
    crawler.config.headers["X-Auth-Token"] = "secret123"

def extract_metadata(crawler, context):
    """Process the final crawl result."""
    result = context["result"]
    print(f"Title: {result.get('title')}")

crawler = Crawl4AI()
crawler.run(
    url="https://httpbin.org/headers",
    max_depth=1,
    hooks={
        "pre_fetch": [add_auth_header],
        "post_crawl": [extract_metadata],
    },
)

The run() method stores this dictionary in self.hooks (line 39), and _invoke_hooks retrieves the appropriate list via self.hooks.get(stage, []) (line 44).

Practical Implementation Examples

Modifying Crawler State Dynamically

Hooks can mutate the crawler instance to alter behavior mid-flight. This example rotates the user agent based on the target URL:

def rotate_user_agent(crawler, context):
    url = context["url"]
    if "mobile" in url:
        crawler.config.user_agent = "MobileBot/1.0"
    else:
        crawler.config.user_agent = "DesktopBot/1.0"

crawler = Crawl4AI()
crawler.run(
    url="https://example.com",
    hooks={"pre_fetch": [rotate_user_agent]}
)

Inspecting HTTP Responses

The post_fetch hook receives the raw Playwright Response object, allowing you to inspect headers or status codes before processing continues:

def check_cache_policy(crawler, context):
    response = context["response"]
    cache_control = response.headers.get("cache-control", "")
    if "no-store" in cache_control:
        print("Warning: Page prohibits storage")

crawler.run(
    url="https://example.com",
    hooks={"post_fetch": [check_cache_policy]}
)

Summary

  • crawl4ai exposes four hook stages—pre_crawl, pre_fetch, post_fetch, and post_crawl—implemented in crawl4ai/crawler.py via the _invoke_hooks method.
  • Register hooks globally using the @register_hook decorator from crawl4ai/hooks/__init__.py (lines 8-12), which stores callables in the hooks_registry dictionary.
  • Alternatively, pass a hooks dictionary to Crawl4AI.run() for per-run customization, as stored in self.hooks (crawler.py line 39).
  • Each hook receives the crawler instance and a context dict containing stage-specific data (URL, Response object, or final result).
  • Hooks execute in the order they are registered, allowing predictable chaining of modifications.

Frequently Asked Questions

What is the execution order when multiple hooks are registered for the same stage?

Hooks execute in the order they are appended to the registry or list. In crawl4ai/hooks/__init__.py (line 10), hooks_registry.setdefault(stage, []).append(func) preserves insertion order. When invoked via _invoke_hooks in crawl4ai/crawler.py (lines 44-45), the code iterates through the list sequentially, ensuring hook_one runs before hook_two if registered first.

Can I combine global decorator hooks with per-run dictionary hooks?

Yes. The crawler's run() method accepts a hooks parameter that merges with the global registry. If you register a hook via @register_hook and also pass a hook in the dictionary for the same stage, both will execute. The global hooks typically run first if the implementation checks the registry before the instance dictionary, though you should verify the exact merge logic in your version of crawl4ai/crawler.py.

How do I modify the HTTP request before it is sent?

Use the pre_fetch stage, which fires immediately before the network request. The context dictionary contains the "url" key, and you can modify crawler.config.headers or other attributes to influence the upcoming fetch. Note that you cannot directly modify the Playwright request object at this stage; for advanced request interception, you would need to use Playwright's route handling features within the hook.

What data is available in the context dictionary for each hook stage?

  • pre_crawl: Empty dictionary {}
  • pre_fetch: {"url": str} containing the target URL
  • post_fetch: {"response": Response} containing the Playwright Response object
  • post_crawl: {"result": dict} containing the final crawl result payload

This data is passed via the context parameter in _invoke_hooks (crawl4ai/crawler.py, lines 43-46) and allows hooks to access relevant state without reaching into private attributes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →