# How to Use LLM Extraction Strategies for Structured JSON Data with Crawl4AI

> Unlock structured JSON data from web pages using LLM extraction strategies with Crawl4AI. Convert unstructured HTML into typed JSON objects effortlessly.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Crawl4AI converts unstructured web pages into typed JSON objects by sending HTML content to an LLM with a predefined schema, returning structured data through the `extracted_data` attribute of the crawl result.**

Crawl4AI provides a dedicated **LLM-based extraction strategy** that eliminates the need for brittle CSS selectors when dealing with complex or varying page layouts. By leveraging the `LLMExtractionStrategy` class in [`crawl4ai/extraction_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/extraction_strategy.py), you can define JSON schemas that guide LLMs to extract specific fields from crawled content. This approach integrates seamlessly with both synchronous and asynchronous crawling workflows.

## Architecture of LLM-Based Extraction

The extraction system isolates all LLM-specific concerns within dedicated classes, allowing you to declare a schema and let the framework handle prompt engineering, backoff logic, and response parsing.

### Core Components

Three primary components work in concert within the crawl4ai codebase:

- **`LLMExtractionStrategy`** (located in [`crawl4ai/extraction_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/extraction_strategy.py)): This class orchestrates the extraction process, supporting both "block" mode (free-form text) and "schema" mode (strict JSON). When initialized with a schema parameter, it automatically switches `extraction_type` to `"schema"` and interpolates the schema into specialized prompt templates like `PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION`.

- **`LLMConfig`** (defined in [`crawl4ai/prompts.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/prompts.py)): A configuration container that specifies the provider string (e.g., `"openai/gpt-4o-mini"`), API token, endpoint URL, and retry backoff settings. The strategy creates a default configuration if none is supplied.

- **`perform_completion_with_backoff`** (implemented in [`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py)): A utility function that handles the actual API communication with exponential backoff, token usage tracking, and error resilience.

### The Schema-Based Execution Flow

When using structured extraction, the following sequence occurs inside the `extract` method:

1. **Schema Injection**: The provided Python dictionary or Pydantic model is JSON-pretty-printed and embedded into `PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION` along with the target HTML.

2. **LLM Invocation**: The strategy calls `perform_completion_with_backoff` to send the composed prompt to the configured provider.

3. **Response Parsing**: When `force_json_response=True` (recommended for schema mode), the raw LLM output is sanitized and parsed using `json.loads` into Python dictionaries.

4. **Result Population**: The extracted list of dictionaries populates the `extracted_data` attribute of the `CrawlerResult` object returned by the crawler.

## Implementing LLM Extraction Strategies

### Basic Synchronous Setup

For single-page extractions, use the synchronous `Crawler` class with an explicit schema definition. Configure authentication through environment variables and initialize `LLMConfig` with your provider details:

```python
import os
from crawl4ai import Crawler, CrawlerRunConfig, LLMExtractionStrategy, LLMConfig

product_schema = {
    "title": "string",
    "price": "float",
    "currency": "string",
    "available": "bool",
    "rating": "float",
    "review_count": "int"
}

strategy = LLMExtractionStrategy(
    llm_config=LLMConfig(
        provider="openai/gpt-4o-mini",
        api_token=os.getenv("OPENAI_API_KEY")
    ),
    schema=product_schema,
    extraction_type="schema",
    force_json_response=True,
    verbose=True
)

run_cfg = CrawlerRunConfig(extraction_strategy=strategy)
result = Crawler(url="https://example.com/product/12345").run(run_cfg)

print(result.extracted_data)  # Returns list of dicts matching product_schema

```

Setting `force_json_response=True` is critical for schema mode, ensuring the LLM returns parseable JSON rather than markdown-formatted code blocks or explanatory text.

### Asynchronous Bulk Processing

For high-throughput scenarios, use `AsyncCrawler` with `AsyncCrawlerRunConfig` (both defined in [`crawl4ai/config.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/config.py)) to process multiple URLs concurrently with `asyncio.gather`:

```python
import os
import asyncio
from crawl4ai import AsyncCrawler, AsyncCrawlerRunConfig, LLMExtractionStrategy, LLMConfig

article_schema = {
    "headline": "string",
    "author": "string",
    "date": "string",
    "content": "string"
}

strategy = LLMExtractionStrategy(
    llm_config=LLMConfig(
        provider="anthropic/claude-3-opus-20240229",
        api_token=os.getenv("ANTHROPIC_API_KEY")
    ),
    schema=article_schema,
    extraction_type="schema",
    force_json_response=True
)

async def extract_from_url(url):
    cfg = AsyncCrawlerRunConfig(
        extraction_strategy=strategy,
        bypass_cache=True
    )
    async with AsyncCrawler(url=url) as crawler:
        result = await crawler.run(cfg)
        return result.extracted_data

async def main():
    urls = [
        "https://news.ycombinator.com/item?id=39582325",
        "https://example.com/blog/ai-trends"
    ]
    results = await asyncio.gather(*(extract_from_url(u) for u in urls))
    for url, data in zip(urls, results):
        print(f"{url} -> {data}")

asyncio.run(main())

```

### Type-Safe Extraction with Pydantic

For production environments, define schemas using Pydantic models to enable validation, autocomplete, and type checking. Convert the model to a JSON schema dictionary using `model_json_schema()`:

```python
from pydantic import BaseModel, Field
from crawl4ai import LLMExtractionStrategy, LLMConfig

class ProductSchema(BaseModel):
    title: str = Field(..., description="Product title")
    price: float = Field(..., description="Price in USD")
    in_stock: bool = Field(..., description="Availability flag")
    rating: float = Field(..., description="Average rating out of 5")
    review_count: int = Field(..., description="Number of reviews")

strategy = LLMExtractionStrategy(
    llm_config=LLMConfig(
        provider="openai/gpt-4o-mini", 
        api_token=os.getenv("OPENAI_API_KEY")
    ),
    schema=ProductSchema.model_json_schema(),
    extraction_type="schema",
    force_json_response=True,
    verbose=True
)

```

## Key Configuration Files and Utilities

Understanding the underlying source files helps with debugging and advanced customization:

- **[`crawl4ai/extraction_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/extraction_strategy.py)**: Contains the `LLMExtractionStrategy` class and prompt templates including `PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION` used when `extraction_type="schema"`.

- **[`crawl4ai/prompts.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/prompts.py)**: Defines the `LLMConfig` dataclass for provider credentials, model identifiers, and retry configuration.

- **[`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py)**: Implements `perform_completion_with_backoff` for resilient API calls with exponential backoff, plus `TokenUsage` tracking for monitoring costs.

- **[`crawl4ai/config.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/config.py)**: Houses `CrawlerRunConfig` and `AsyncCrawlerRunConfig`, the containers that bind the extraction strategy to the crawler runtime.

## Summary

- **Schema Definition**: Define extraction targets using Python dictionaries or Pydantic models to enable strict typing and validation.
- **Strategy Configuration**: Instantiate `LLMExtractionStrategy` with `LLMConfig` containing provider details; always set `force_json_response=True` for reliable JSON parsing.
- **Crawler Integration**: Pass the strategy to `CrawlerRunConfig` (synchronous) or `AsyncCrawlerRunConfig` (asynchronous); extracted data appears in `result.extracted_data`.
- **Resilience**: The framework automatically handles rate limiting and retries through `perform_completion_with_backoff` in [`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py).

## Frequently Asked Questions

### What is the difference between block mode and schema mode in LLMExtractionStrategy?

Block mode extracts free-form text segments without structural constraints, suitable for content summarization or open-ended analysis. Schema mode (activated by providing a `schema` parameter) forces the LLM to return valid JSON matching your specified structure, enabling direct database insertion or API responses. The constructor automatically sets `extraction_type="schema"` when a schema is detected.

### How does Crawl4AI handle LLM API failures or rate limits?

The `perform_completion_with_backoff` function in [`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py) implements exponential backoff and retry logic for all LLM API calls. Configure connection timeouts and maximum retry attempts through the `LLMConfig` class parameters when initializing your extraction strategy.

### Can I use local LLM models or Azure OpenAI instead of commercial APIs?

Yes. The `provider` field in `LLMConfig` accepts various string identifiers including local model endpoints, `"openai/gpt-4o-mini"`, `"anthropic/claude-3-opus-20240229"`, or custom base URLs. Ensure your API token and endpoint configuration match your provider's requirements.

### Why is my extracted_data empty or containing markdown instead of JSON?

Verify that `force_json_response=True` is set in your `LLMExtractionStrategy` initialization. Without this flag, the LLM may return markdown-formatted JSON or explanatory text that fails parsing. Additionally, check that your schema JSON is valid and that the crawled HTML actually contains the information requested in your schema definition.