How to Set Up Webhooks for Crawling Event Notifications in Crawl4AI

Configure a webhook_config object when you submit a crawl job; Crawl4AI will automatically POST event notifications to your endpoint with exponential‑backoff retries.

Crawl4AI is an open‑source async crawling framework that can notify external systems the moment a crawl finishes (or fails). By leveraging the webhook subsystem implemented in the Docker deployment layer, you can receive JSON payloads containing task status, URLs, and even the full crawl result without polling the API.


Understanding the Webhook Architecture in Crawl4AI

The webhook delivery pipeline is split into three logical pieces:

  1. Configuration schema – WebhookConfig (lines 91‑96) defines the fields a client may send: webhook_url, webhook_data_in_payload, and webhook_headers.
  2. Delivery service – WebhookDeliveryService (lines 18‑29) handles the actual HTTP POST, retry logic with exponential back‑off, and merging of global defaults with per‑job overrides.
  3. Payload contract – WebhookPayload (lines 98‑106) describes the JSON structure you will receive: task_id, task_type, status, urls, optional data, and timestamps.

When a crawl job finishes, the server invokes WebhookDeliveryService.notify_job_completion (lines 99‑159 in webhook.py), which assembles the payload and calls send_webhook (lines 53‑92) to perform the HTTP request.


Configuring Webhooks for a Crawl Job

Per‑Job Webhook Configuration

Pass a webhook_config dictionary when you create a job via the Python client:

from crawl4ai import Crawl4AI

crawler = Crawl4AI()

job = crawler.crawl(
    urls=["https://example.com", "https://example.org"],
    webhook_config={
        "webhook_url": "https://myserver.com/webhooks/crawl4ai",
        "webhook_data_in_payload": True,   # embed full crawl result

        "webhook_headers": {
            "Authorization": "Bearer sk_live_12345",
            "X-Custom-Id": "abc"
        }
    }
)

The webhook_config object is validated against the WebhookConfig schema. If you omit webhook_data_in_payload, the payload will contain only metadata (task ID, status, URLs) and omit the potentially large crawl result.

Raw HTTP API Request

If you prefer to call the REST endpoint directly (e.g., from a CI pipeline), include the same webhook_config object in the JSON body:

curl -X POST https://crawl4ai.example.com/crawl/job \
  -H "Content-Type: application/json" \
  -d '{
        "urls": ["https://example.com"],
        "webhook_config": {
          "webhook_url": "https://myserver.com/webhook",
          "webhook_data_in_payload": false,
          "webhook_headers": {"X-API-Key": "secret"}
        }
      }'

The server stores the configuration alongside the job record and triggers the webhook once the crawl reaches a terminal state (completed, failed, or cancelled).


Handling Webhook Delivery and Retries

The WebhookDeliveryService implements a resilient delivery mechanism:

  • Exponential back‑off – After a failed attempt, the service waits initial_delay_ms (default 500 ms) and doubles the delay on each subsequent failure up to max_delay_ms (default 8000 ms).
  • Maximum attempts – The service stops after max_attempts (default 5) and logs the final failure.
  • Timeout – Each HTTP request respects timeout_ms (default 15000 ms).

These defaults can be tuned globally via the Docker configuration file (see the next section) or overridden on a per‑job basis using the webhook_config fields retry_policy, timeout_ms, etc., if the schema is extended in your fork.


Receiving Webhook Notifications

Example FastAPI Receiver

Below is a minimal Python server that validates the incoming Crawl4AI payload:

from fastapi import FastAPI, Request, HTTPException
from pydantic import BaseModel, Field
from typing import List, Optional, Dict, Any
import datetime

app = FastAPI()

class WebhookPayload(BaseModel):
    task_id: str
    task_type: str
    status: str
    urls: List[str]
    data: Optional[Dict[str, Any]] = None
    created_at: datetime.datetime
    updated_at: datetime.datetime

@app.post("/webhook")
async def receive_webhook(payload: WebhookPayload):
    print(f"Received notification for task {payload.task_id}")
    print(f"Status: {payload.status}, URLs: {payload.urls}")
    if payload.data:
        print("Crawl result included!")
    return {"status": "received"}

Run the server with uvicorn main:app --port 8000 and point webhook_url to http://<your-host>:8000/webhook.

Payload Structure

The JSON you receive conforms to the WebhookPayload schema:

Field Type Description
task_id string UUID of the crawl job
task_type string "crawl" or "llm"
status string completed, failed, cancelled
urls list Original URLs submitted
data object Only present if webhook_data_in_payload was true
created_at ISO‑8601 Job creation timestamp
updated_at ISO‑8601 Last status change timestamp

Global Webhook Defaults (Optional)

When deploying Crawl4AI via Docker, you can supply a config.yml that defines default webhook behavior so you don’t have to repeat the URL in every request:

webhooks:
  enabled: true
  default_url: "https://myapp.com/default-webhook"
  data_in_payload: false
  headers:
    Authorization: "Bearer global-token"
  retry:
    max_attempts: 4
    initial_delay_ms: 500
    max_delay_ms: 8000
    timeout_ms: 15000

The WebhookDeliveryService loads these values on startup (see webhook.py lines 25‑29) and merges them with any per‑job overrides. If enabled is false, the service skips delivery entirely.


Summary


Frequently Asked Questions

What URL should I use for the webhook endpoint?

Use any HTTPS (or HTTP for local testing) endpoint that you control. Common choices are:

  • A FastAPI/Flask route on your own server (e.g., https://api.myapp.com/webhooks/crawl4ai).
  • A serverless function (AWS Lambda, Google Cloud Function) with a public URL.
  • A Zapier or Make (Integromat) webhook URL to forward events to other SaaS tools.

Can I include custom headers in the webhook request?

Yes. The webhook_config object accepts a webhook_headers dictionary (see WebhookConfig in [schemas.py](https://github.com/unclecode/crawl4ai/blob/main/deploy/docker/schemas.py)). These headers are merged with any global headers defined in config.yml and sent with every POST request to your endpoint. This is useful for passing API keys or request signatures.

How does Crawl4AI handle webhook delivery failures?

The WebhookDeliveryService implements an exponential back‑off retry strategy (defined in [webhook.py](https://github.com/unclecode/crawl4ai/blob/main/deploy/docker/webhook.py) lines 53‑92). If the initial POST fails (network error or non‑2xx response), the service waits initial_delay_ms (default 500 ms), then retries, doubling the delay each time up to max_delay_ms (default 8000 ms). After max_attempts (default 5) failures, the delivery is abandoned and the failure is logged.

Is it possible to disable webhooks globally?

Yes. When running the Docker deployment, set webhooks.enabled: false in your config.yml (see the Global Webhook Defaults section). When disabled, the WebhookDeliveryService skips all delivery attempts regardless of per‑job settings, which is useful for local development or air‑gapped environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →