How to Use LLM Extraction Strategies for Structured JSON Data with Crawl4AI

Crawl4AI converts unstructured web pages into typed JSON objects by sending HTML content to an LLM with a predefined schema, returning structured data through the extracted_data attribute of the crawl result.

Crawl4AI provides a dedicated LLM-based extraction strategy that eliminates the need for brittle CSS selectors when dealing with complex or varying page layouts. By leveraging the LLMExtractionStrategy class in crawl4ai/extraction_strategy.py, you can define JSON schemas that guide LLMs to extract specific fields from crawled content. This approach integrates seamlessly with both synchronous and asynchronous crawling workflows.

Architecture of LLM-Based Extraction

The extraction system isolates all LLM-specific concerns within dedicated classes, allowing you to declare a schema and let the framework handle prompt engineering, backoff logic, and response parsing.

Core Components

Three primary components work in concert within the crawl4ai codebase:

  • LLMExtractionStrategy (located in crawl4ai/extraction_strategy.py): This class orchestrates the extraction process, supporting both "block" mode (free-form text) and "schema" mode (strict JSON). When initialized with a schema parameter, it automatically switches extraction_type to "schema" and interpolates the schema into specialized prompt templates like PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION.

  • LLMConfig (defined in crawl4ai/prompts.py): A configuration container that specifies the provider string (e.g., "openai/gpt-4o-mini"), API token, endpoint URL, and retry backoff settings. The strategy creates a default configuration if none is supplied.

  • perform_completion_with_backoff (implemented in crawl4ai/utils.py): A utility function that handles the actual API communication with exponential backoff, token usage tracking, and error resilience.

The Schema-Based Execution Flow

When using structured extraction, the following sequence occurs inside the extract method:

  1. Schema Injection: The provided Python dictionary or Pydantic model is JSON-pretty-printed and embedded into PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION along with the target HTML.

  2. LLM Invocation: The strategy calls perform_completion_with_backoff to send the composed prompt to the configured provider.

  3. Response Parsing: When force_json_response=True (recommended for schema mode), the raw LLM output is sanitized and parsed using json.loads into Python dictionaries.

  4. Result Population: The extracted list of dictionaries populates the extracted_data attribute of the CrawlerResult object returned by the crawler.

Implementing LLM Extraction Strategies

Basic Synchronous Setup

For single-page extractions, use the synchronous Crawler class with an explicit schema definition. Configure authentication through environment variables and initialize LLMConfig with your provider details:

import os
from crawl4ai import Crawler, CrawlerRunConfig, LLMExtractionStrategy, LLMConfig

product_schema = {
    "title": "string",
    "price": "float",
    "currency": "string",
    "available": "bool",
    "rating": "float",
    "review_count": "int"
}

strategy = LLMExtractionStrategy(
    llm_config=LLMConfig(
        provider="openai/gpt-4o-mini",
        api_token=os.getenv("OPENAI_API_KEY")
    ),
    schema=product_schema,
    extraction_type="schema",
    force_json_response=True,
    verbose=True
)

run_cfg = CrawlerRunConfig(extraction_strategy=strategy)
result = Crawler(url="https://example.com/product/12345").run(run_cfg)

print(result.extracted_data)  # Returns list of dicts matching product_schema

Setting force_json_response=True is critical for schema mode, ensuring the LLM returns parseable JSON rather than markdown-formatted code blocks or explanatory text.

Asynchronous Bulk Processing

For high-throughput scenarios, use AsyncCrawler with AsyncCrawlerRunConfig (both defined in crawl4ai/config.py) to process multiple URLs concurrently with asyncio.gather:

import os
import asyncio
from crawl4ai import AsyncCrawler, AsyncCrawlerRunConfig, LLMExtractionStrategy, LLMConfig

article_schema = {
    "headline": "string",
    "author": "string",
    "date": "string",
    "content": "string"
}

strategy = LLMExtractionStrategy(
    llm_config=LLMConfig(
        provider="anthropic/claude-3-opus-20240229",
        api_token=os.getenv("ANTHROPIC_API_KEY")
    ),
    schema=article_schema,
    extraction_type="schema",
    force_json_response=True
)

async def extract_from_url(url):
    cfg = AsyncCrawlerRunConfig(
        extraction_strategy=strategy,
        bypass_cache=True
    )
    async with AsyncCrawler(url=url) as crawler:
        result = await crawler.run(cfg)
        return result.extracted_data

async def main():
    urls = [
        "https://news.ycombinator.com/item?id=39582325",
        "https://example.com/blog/ai-trends"
    ]
    results = await asyncio.gather(*(extract_from_url(u) for u in urls))
    for url, data in zip(urls, results):
        print(f"{url} -> {data}")

asyncio.run(main())

Type-Safe Extraction with Pydantic

For production environments, define schemas using Pydantic models to enable validation, autocomplete, and type checking. Convert the model to a JSON schema dictionary using model_json_schema():

from pydantic import BaseModel, Field
from crawl4ai import LLMExtractionStrategy, LLMConfig

class ProductSchema(BaseModel):
    title: str = Field(..., description="Product title")
    price: float = Field(..., description="Price in USD")
    in_stock: bool = Field(..., description="Availability flag")
    rating: float = Field(..., description="Average rating out of 5")
    review_count: int = Field(..., description="Number of reviews")

strategy = LLMExtractionStrategy(
    llm_config=LLMConfig(
        provider="openai/gpt-4o-mini", 
        api_token=os.getenv("OPENAI_API_KEY")
    ),
    schema=ProductSchema.model_json_schema(),
    extraction_type="schema",
    force_json_response=True,
    verbose=True
)

Key Configuration Files and Utilities

Understanding the underlying source files helps with debugging and advanced customization:

  • crawl4ai/extraction_strategy.py: Contains the LLMExtractionStrategy class and prompt templates including PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION used when extraction_type="schema".

  • crawl4ai/prompts.py: Defines the LLMConfig dataclass for provider credentials, model identifiers, and retry configuration.

  • crawl4ai/utils.py: Implements perform_completion_with_backoff for resilient API calls with exponential backoff, plus TokenUsage tracking for monitoring costs.

  • crawl4ai/config.py: Houses CrawlerRunConfig and AsyncCrawlerRunConfig, the containers that bind the extraction strategy to the crawler runtime.

Summary

  • Schema Definition: Define extraction targets using Python dictionaries or Pydantic models to enable strict typing and validation.
  • Strategy Configuration: Instantiate LLMExtractionStrategy with LLMConfig containing provider details; always set force_json_response=True for reliable JSON parsing.
  • Crawler Integration: Pass the strategy to CrawlerRunConfig (synchronous) or AsyncCrawlerRunConfig (asynchronous); extracted data appears in result.extracted_data.
  • Resilience: The framework automatically handles rate limiting and retries through perform_completion_with_backoff in crawl4ai/utils.py.

Frequently Asked Questions

What is the difference between block mode and schema mode in LLMExtractionStrategy?

Block mode extracts free-form text segments without structural constraints, suitable for content summarization or open-ended analysis. Schema mode (activated by providing a schema parameter) forces the LLM to return valid JSON matching your specified structure, enabling direct database insertion or API responses. The constructor automatically sets extraction_type="schema" when a schema is detected.

How does Crawl4AI handle LLM API failures or rate limits?

The perform_completion_with_backoff function in crawl4ai/utils.py implements exponential backoff and retry logic for all LLM API calls. Configure connection timeouts and maximum retry attempts through the LLMConfig class parameters when initializing your extraction strategy.

Can I use local LLM models or Azure OpenAI instead of commercial APIs?

Yes. The provider field in LLMConfig accepts various string identifiers including local model endpoints, "openai/gpt-4o-mini", "anthropic/claude-3-opus-20240229", or custom base URLs. Ensure your API token and endpoint configuration match your provider's requirements.

Why is my extracted_data empty or containing markdown instead of JSON?

Verify that force_json_response=True is set in your LLMExtractionStrategy initialization. Without this flag, the LLM may return markdown-formatted JSON or explanatory text that fails parsing. Additionally, check that your schema JSON is valid and that the crawled HTML actually contains the information requested in your schema definition.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →