How to Set Up Webhooks for Crawling Event Notifications in Crawl4AI
Configure a webhook_config object when you submit a crawl job; Crawl4AI will automatically POST event notifications to your endpoint with exponential‑backoff retries.
Crawl4AI is an open‑source async crawling framework that can notify external systems the moment a crawl finishes (or fails). By leveraging the webhook subsystem implemented in the Docker deployment layer, you can receive JSON payloads containing task status, URLs, and even the full crawl result without polling the API.
Understanding the Webhook Architecture in Crawl4AI
The webhook delivery pipeline is split into three logical pieces:
- Configuration schema –
WebhookConfig(lines 91‑96) defines the fields a client may send:webhook_url,webhook_data_in_payload, andwebhook_headers. - Delivery service –
WebhookDeliveryService(lines 18‑29) handles the actual HTTP POST, retry logic with exponential back‑off, and merging of global defaults with per‑job overrides. - Payload contract –
WebhookPayload(lines 98‑106) describes the JSON structure you will receive:task_id,task_type,status,urls, optionaldata, and timestamps.
When a crawl job finishes, the server invokes WebhookDeliveryService.notify_job_completion (lines 99‑159 in webhook.py), which assembles the payload and calls send_webhook (lines 53‑92) to perform the HTTP request.
Configuring Webhooks for a Crawl Job
Per‑Job Webhook Configuration
Pass a webhook_config dictionary when you create a job via the Python client:
from crawl4ai import Crawl4AI
crawler = Crawl4AI()
job = crawler.crawl(
urls=["https://example.com", "https://example.org"],
webhook_config={
"webhook_url": "https://myserver.com/webhooks/crawl4ai",
"webhook_data_in_payload": True, # embed full crawl result
"webhook_headers": {
"Authorization": "Bearer sk_live_12345",
"X-Custom-Id": "abc"
}
}
)
The webhook_config object is validated against the WebhookConfig schema. If you omit webhook_data_in_payload, the payload will contain only metadata (task ID, status, URLs) and omit the potentially large crawl result.
Raw HTTP API Request
If you prefer to call the REST endpoint directly (e.g., from a CI pipeline), include the same webhook_config object in the JSON body:
curl -X POST https://crawl4ai.example.com/crawl/job \
-H "Content-Type: application/json" \
-d '{
"urls": ["https://example.com"],
"webhook_config": {
"webhook_url": "https://myserver.com/webhook",
"webhook_data_in_payload": false,
"webhook_headers": {"X-API-Key": "secret"}
}
}'
The server stores the configuration alongside the job record and triggers the webhook once the crawl reaches a terminal state (completed, failed, or cancelled).
Handling Webhook Delivery and Retries
The WebhookDeliveryService implements a resilient delivery mechanism:
- Exponential back‑off – After a failed attempt, the service waits
initial_delay_ms(default 500 ms) and doubles the delay on each subsequent failure up tomax_delay_ms(default 8000 ms). - Maximum attempts – The service stops after
max_attempts(default 5) and logs the final failure. - Timeout – Each HTTP request respects
timeout_ms(default 15000 ms).
These defaults can be tuned globally via the Docker configuration file (see the next section) or overridden on a per‑job basis using the webhook_config fields retry_policy, timeout_ms, etc., if the schema is extended in your fork.
Receiving Webhook Notifications
Example FastAPI Receiver
Below is a minimal Python server that validates the incoming Crawl4AI payload:
from fastapi import FastAPI, Request, HTTPException
from pydantic import BaseModel, Field
from typing import List, Optional, Dict, Any
import datetime
app = FastAPI()
class WebhookPayload(BaseModel):
task_id: str
task_type: str
status: str
urls: List[str]
data: Optional[Dict[str, Any]] = None
created_at: datetime.datetime
updated_at: datetime.datetime
@app.post("/webhook")
async def receive_webhook(payload: WebhookPayload):
print(f"Received notification for task {payload.task_id}")
print(f"Status: {payload.status}, URLs: {payload.urls}")
if payload.data:
print("Crawl result included!")
return {"status": "received"}
Run the server with uvicorn main:app --port 8000 and point webhook_url to http://<your-host>:8000/webhook.
Payload Structure
The JSON you receive conforms to the WebhookPayload schema:
| Field | Type | Description |
|---|---|---|
task_id |
string | UUID of the crawl job |
task_type |
string | "crawl" or "llm" |
status |
string | completed, failed, cancelled |
urls |
list | Original URLs submitted |
data |
object | Only present if webhook_data_in_payload was true |
created_at |
ISO‑8601 | Job creation timestamp |
updated_at |
ISO‑8601 | Last status change timestamp |
Global Webhook Defaults (Optional)
When deploying Crawl4AI via Docker, you can supply a config.yml that defines default webhook behavior so you don’t have to repeat the URL in every request:
webhooks:
enabled: true
default_url: "https://myapp.com/default-webhook"
data_in_payload: false
headers:
Authorization: "Bearer global-token"
retry:
max_attempts: 4
initial_delay_ms: 500
max_delay_ms: 8000
timeout_ms: 15000
The WebhookDeliveryService loads these values on startup (see webhook.py lines 25‑29) and merges them with any per‑job overrides. If enabled is false, the service skips delivery entirely.
Summary
- Webhook configuration is supplied at job creation time via the
webhook_configobject (or globally viaconfig.yml). - Schemas in [
deploy/docker/schemas.py](https://github.com/unclecode/crawl4ai/blob/main/deploy/docker/schemas.py) defineWebhookConfig(input) andWebhookPayload(output). - Delivery logic lives in [
deploy/docker/webhook.py](https://github.com/unclecode/crawl4ai/blob/main/deploy/docker/webhook.py) inside theWebhookDeliveryServiceclass, which handles retries, back‑off, and payload assembly. - Per‑job overrides take precedence over global defaults, letting you tailor URLs, headers, and whether the full crawl result is included.
- Receiver implementation can be as simple as a FastAPI endpoint that accepts the
WebhookPayloadJSON and processes the status or result data.
Frequently Asked Questions
What URL should I use for the webhook endpoint?
Use any HTTPS (or HTTP for local testing) endpoint that you control. Common choices are:
- A FastAPI/Flask route on your own server (e.g.,
https://api.myapp.com/webhooks/crawl4ai). - A serverless function (AWS Lambda, Google Cloud Function) with a public URL.
- A Zapier or Make (Integromat) webhook URL to forward events to other SaaS tools.
Can I include custom headers in the webhook request?
Yes. The webhook_config object accepts a webhook_headers dictionary (see WebhookConfig in [schemas.py](https://github.com/unclecode/crawl4ai/blob/main/deploy/docker/schemas.py)). These headers are merged with any global headers defined in config.yml and sent with every POST request to your endpoint. This is useful for passing API keys or request signatures.
How does Crawl4AI handle webhook delivery failures?
The WebhookDeliveryService implements an exponential back‑off retry strategy (defined in [webhook.py](https://github.com/unclecode/crawl4ai/blob/main/deploy/docker/webhook.py) lines 53‑92). If the initial POST fails (network error or non‑2xx response), the service waits initial_delay_ms (default 500 ms), then retries, doubling the delay each time up to max_delay_ms (default 8000 ms). After max_attempts (default 5) failures, the delivery is abandoned and the failure is logged.
Is it possible to disable webhooks globally?
Yes. When running the Docker deployment, set webhooks.enabled: false in your config.yml (see the Global Webhook Defaults section). When disabled, the WebhookDeliveryService skips all delivery attempts regardless of per‑job settings, which is useful for local development or air‑gapped environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →