How to Use the crwl CLI for Quick Crawling Tasks in Crawl4AI

The crwl command is Crawl4AI's Click-based CLI entry point, defined in crawl4ai/cli.py and registered via pyproject.toml, that defaults to the crawl subcommand to let you fetch pages, extract structured data, and query LLMs from a single terminal line without writing Python code.

The crwl CLI wraps the full Crawl4AI SDK into a lightweight terminal interface. Whether you need a quick HTML dump, a markdown conversion, or a deep crawl across multiple pages, the CLI exposes every major feature through command-line flags and YAML-based configuration files.

CLI Architecture and Entry Point

The crwl binary is registered as a console script in pyproject.toml (lines 78‑84):

[project.scripts]
crwl = "crawl4ai.cli:main"

When you type crwl, Python invokes the main() function in crawl4ai/cli.py. This function creates a Click command group and automatically forwards the call to the crawl subcommand if no other subcommand (such as profiles, browser, or config) is specified.

Installation and First Run

Install Crawl4AI via pip to make the crwl command available globally:

pip install crawl4ai

On first use of LLM-powered features, the CLI may prompt you to download required models. This is handled by the helper logic in crawl4ai/model_loader.py, which ensures local models are cached in ~/.crawl4ai/.

Core Commands for Quick Crawling

Basic Single-Page Crawling

The simplest invocation fetches raw HTML and prints it to STDOUT:

crwl https://example.com

To convert the page to clean markdown, specify the output format:

crwl https://example.com -o markdown

Available formats include json, markdown, html, and raw.

Structured Data Extraction

Pass an extraction configuration file and a JSON schema to pull specific fields without writing Python:

crwl "https://www.infoq.com/ai-ml-data-eng/" \
    -e docs/examples/cli/extract_css.yml \
    -s docs/examples/cli/css_schema.json \
    -o json

The -e flag loads a YAML extraction config (e.g., json-css type), while -s points to the schema that defines which CSS selectors map to which output fields.

Deep Crawling with BFS or DFS

Enable multi-page crawling with the --deep-crawl flag, specifying either bfs or dfs strategy:

crwl https://news.ycombinator.com \
    --deep-crawl bfs \
    --max-pages 5 \
    -o json

This command performs a breadth-first crawl starting from the seed URL, halting after five pages.

LLM-Powered Q&A

Ask a question about the page content directly from the terminal:

crwl https://example.com -q "What are the main take-aways?"

If you have not configured an LLM provider, the CLI interactively prompts for API keys or local model paths and persists the configuration to ~/.crawl4ai/global.yml.

Profile-Based Browsing

Reuse authenticated sessions by creating a browser profile first:


# Interactive profile creation

crwl profiles

Then reference the profile name in subsequent crawls:

crwl https://linkedin.com/in/someone -p linkedin-profile -o markdown

The profile stores cookies, local storage, and authentication state, allowing you to crawl sites that require login without re-authenticating each time.

Advanced Configuration with YAML and Inline Overrides

For complex tasks, combine configuration files with inline overrides:

crwl https://example.com \
    -B my_browser.yml \
    -b "headless=false,viewport_width=1600"
  • -B loads a full browser configuration YAML file.
  • -b accepts comma-separated key-value pairs that override specific keys in the loaded config or defaults.

Similarly, use -C for crawler configuration files and -c for inline crawler overrides, or -e and -s for extraction pipelines as shown earlier.

Summary

  • The crwl command is the entry point defined in pyproject.toml and implemented in crawl4ai/cli.py using the Click framework.
  • It defaults to the crawl subcommand, enabling one-liner usage for fetching pages.
  • Use -o to select output formats (json, markdown, html, raw).
  • Use -e and -s for structured data extraction via YAML configs and JSON schemas.
  • Enable multi-page crawling with --deep-crawl bfs|dfs and --max-pages.
  • Leverage -q for LLM-powered questions and -p for persistent browser profiles.
  • Combine -B/-b, -C/-c to mix YAML configuration files with inline parameter overrides.

Frequently Asked Questions

How do I install the crwl CLI?

Install Crawl4AI via pip, which registers the crwl console script globally. Run pip install crawl4ai and then verify installation with crwl --help. The entry point is defined in pyproject.toml at lines 78‑84, mapping crwl to crawl4ai.cli:main.

What is the difference between the crwl CLI and the Python SDK?

The crwl CLI is a thin Click-based wrapper around the Crawl4AI Python SDK. While the SDK requires writing Python code to instantiate Crawl4AI and call arun(), the CLI exposes the same functionality through command-line flags and YAML configuration files, making it ideal for shell scripts, cron jobs, and one-off data extraction tasks.

Can I use crwl for deep crawling multiple pages?

Yes. Pass the --deep-crawl flag with either bfs (breadth-first) or dfs (depth-first) strategy, and limit the scope with --max-pages. For example: crwl https://example.com --deep-crawl bfs --max-pages 10 -o json. This invokes the crawler's internal BFS/DFS logic defined in the core SDK without requiring Python code.

How do I extract specific data fields using crwl?

Use the -e flag to point to a YAML extraction configuration file and -s to provide a JSON schema that maps CSS selectors or LLM prompts to output fields. For instance: crwl https://site.com -e extract.yml -s schema.json -o json. The CLI loads these configurations and passes them to the extraction pipeline in crawl4ai/cli.py, serializing the results to your chosen format.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →