How to Deploy Crawl4AI in Docker with the Official Image

Deploy Crawl4AI by pulling the official unclecode/crawl4ai image from Docker Hub and running it with port mapping and shared memory allocation, or orchestrate production-ready stacks using the repository's docker-compose.yml file.

Crawl4AI provides an official Docker image that packages the FastAPI service, Redis cache, and Playwright browser engine into a hardened, single-container deployment. Whether you need to deploy crawl4ai in docker for local development or production scraping workflows, this guide covers the official image architecture, runtime configuration, and API access patterns based on the source code in the unclecode/crawl4ai repository.

Understanding the Official Image Architecture

The official image (unclecode/crawl4ai:latest) is built from the repository's Dockerfile, which uses python:3.12-slim-bookworm as its secure, minimal base. The build process installs system dependencies including build-essential, git, redis-server, supervisor, and a full suite of Chromium libraries required by Playwright (lines 34-66). For AMD64 systems requiring hardware acceleration, passing ENABLE_GPU=true as a build argument installs the CUDA toolkit (lines 77-84).

Security Hardening and Process Management

The container enforces security through a dedicated non-root appuser created in lines 102-107 of the Dockerfile, ensuring the process runs without elevated privileges. A health check defined in lines 87-95 periodically verifies that the FastAPI service responds at http://localhost:11235/health. Inside the container, Supervisor (configured via deploy/docker/supervisord.conf at line 205) manages both the Redis server and the FastAPI application, automatically restarting either service if they crash.

Quick Start with Docker Run

For immediate testing without custom builds, pull the pre-built image and expose port 11235:

docker pull unclecode/crawl4ai:latest

docker run -d \
  -p 11235:11235 \
  --shm-size=1g \
  --name crawl4ai \
  unclecode/crawl4ai:latest

The --shm-size=1g flag is critical for Chromium stability, allocating sufficient shared memory for browser process communication. Upon startup, Supervisor launches Redis and the FastAPI server automatically, making the API available at http://localhost:11235.

Production Deployment with Docker Compose

For persistent deployments with health monitoring and resource constraints, use the provided [docker-compose.yml](https://github.com/unclecode/crawl4ai/blob/main/docker-compose.yml):

version: '3.8'

x-base-config: &base-config
  ports:
    - "11235:11235"
  volumes:
    - /dev/shm:/dev/shm
  env_file:
    - .llm.env
  deploy:
    resources:
      limits:
        memory: 4G
  restart: unless-stopped
  healthcheck:
    test: ["CMD", "curl", "-f", "http://localhost:11235/health"]
    interval: 30s
    timeout: 10s
    retries: 3
    start_period: 40s
  user: "appuser"

services:
  crawl4ai:
    image: ${IMAGE:-unclecode/crawl4ai:${TAG:-latest}}
    build:
      context: .
      dockerfile: Dockerfile
      args:
        INSTALL_TYPE: ${INSTALL_TYPE:-default}
        ENABLE_GPU: ${ENABLE_GPU:-false}
    <<: *base-config

Deploy the stack with:

docker compose up -d

This configuration mounts the host's shared memory into the container, enforces the non-root appuser at runtime (line 34), and applies a 4GB memory limit to prevent resource exhaustion.

Build Arguments for Customization

When building locally rather than pulling the official image, you can customize the deployment using these Dockerfile arguments:

  • INSTALL_TYPE: Controls dependency installation profiles (default or full)
  • ENABLE_GPU: Enables CUDA support when set to true on AMD64 architectures
  • USE_LOCAL: Determines whether to install the package from local source (/tmp/project) or directly from GitHub

Accessing the API and Dashboard

Once the container is healthy, interact with the crawler through the REST API or web interface. Submit crawl jobs programmatically:

import requests

response = requests.post(
    "http://localhost:11235/crawl",
    json={"urls": ["https://example.com"], "priority": 10}
)

if response.ok:
    task_id = response.json()["task_id"]
    result = requests.get(f"http://localhost:11235/task/{task_id}").json()
    print(result["results"])

Access the real-time monitoring dashboard at http://localhost:11235/dashboard to view browser pool utilization, memory consumption, and active crawl tasks.

Summary

  • The official unclecode/crawl4ai image combines FastAPI, Redis, and Playwright into a single container orchestrated by Supervisor according to deploy/docker/supervisord.conf
  • Use docker run -p 11235:11235 --shm-size=1g for quick deployments, ensuring adequate shared memory for Chromium rendering processes
  • Production environments should use the provided docker-compose.yml with explicit health checks, 4GB memory limits, and the non-root appuser for security isolation
  • The container exposes port 11235 for the REST API and serves a monitoring dashboard at the /dashboard endpoint
  • GPU acceleration is available when building the image with ENABLE_GPU=true on AMD64 architectures

Frequently Asked Questions

What port does Crawl4AI use in Docker?

Crawl4AI exposes the FastAPI service on port 11235 inside the container. You must map this to a host port using -p 11235:11235 in Docker run commands or the ports section of Docker Compose to access the API and web dashboard from outside the container.

How do I enable GPU support in the Crawl4AI container?

GPU support requires building the image locally with the ENABLE_GPU=true build argument on AMD64 architecture. The Dockerfile installs CUDA toolkit dependencies in lines 77-84 when this flag is enabled, allowing Playwright to leverage GPU acceleration for browser rendering tasks.

Is the Crawl4AI Docker image secure for production?

Yes. The Dockerfile implements security hardening by creating a dedicated non-root appuser in lines 102-107, disabling file URL access and external hooks by default as documented in v0.8.0 security fixes. The health check defined in lines 87-95 provides automated failure detection for orchestration platforms.

Why does the container need --shm-size=1g?

Chromium and Playwright require shared memory (/dev/shm) for inter-process communication and rendering complex JavaScript-heavy pages. Without adequate shared memory (Docker defaults to 64MB), Chromium crashes during execution. The docker-compose.yml persists this by mounting the host's /dev/shm volume directly into the container.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →