How to Use LLM Extraction Strategies for Structured JSON Data with Crawl4AI
Crawl4AI converts unstructured web pages into typed JSON objects by sending HTML content to an LLM with a predefined schema, returning structured data through the extracted_data attribute of the crawl result.
Crawl4AI provides a dedicated LLM-based extraction strategy that eliminates the need for brittle CSS selectors when dealing with complex or varying page layouts. By leveraging the LLMExtractionStrategy class in crawl4ai/extraction_strategy.py, you can define JSON schemas that guide LLMs to extract specific fields from crawled content. This approach integrates seamlessly with both synchronous and asynchronous crawling workflows.
Architecture of LLM-Based Extraction
The extraction system isolates all LLM-specific concerns within dedicated classes, allowing you to declare a schema and let the framework handle prompt engineering, backoff logic, and response parsing.
Core Components
Three primary components work in concert within the crawl4ai codebase:
-
LLMExtractionStrategy(located incrawl4ai/extraction_strategy.py): This class orchestrates the extraction process, supporting both "block" mode (free-form text) and "schema" mode (strict JSON). When initialized with a schema parameter, it automatically switchesextraction_typeto"schema"and interpolates the schema into specialized prompt templates likePROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION. -
LLMConfig(defined incrawl4ai/prompts.py): A configuration container that specifies the provider string (e.g.,"openai/gpt-4o-mini"), API token, endpoint URL, and retry backoff settings. The strategy creates a default configuration if none is supplied. -
perform_completion_with_backoff(implemented incrawl4ai/utils.py): A utility function that handles the actual API communication with exponential backoff, token usage tracking, and error resilience.
The Schema-Based Execution Flow
When using structured extraction, the following sequence occurs inside the extract method:
-
Schema Injection: The provided Python dictionary or Pydantic model is JSON-pretty-printed and embedded into
PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTIONalong with the target HTML. -
LLM Invocation: The strategy calls
perform_completion_with_backoffto send the composed prompt to the configured provider. -
Response Parsing: When
force_json_response=True(recommended for schema mode), the raw LLM output is sanitized and parsed usingjson.loadsinto Python dictionaries. -
Result Population: The extracted list of dictionaries populates the
extracted_dataattribute of theCrawlerResultobject returned by the crawler.
Implementing LLM Extraction Strategies
Basic Synchronous Setup
For single-page extractions, use the synchronous Crawler class with an explicit schema definition. Configure authentication through environment variables and initialize LLMConfig with your provider details:
import os
from crawl4ai import Crawler, CrawlerRunConfig, LLMExtractionStrategy, LLMConfig
product_schema = {
"title": "string",
"price": "float",
"currency": "string",
"available": "bool",
"rating": "float",
"review_count": "int"
}
strategy = LLMExtractionStrategy(
llm_config=LLMConfig(
provider="openai/gpt-4o-mini",
api_token=os.getenv("OPENAI_API_KEY")
),
schema=product_schema,
extraction_type="schema",
force_json_response=True,
verbose=True
)
run_cfg = CrawlerRunConfig(extraction_strategy=strategy)
result = Crawler(url="https://example.com/product/12345").run(run_cfg)
print(result.extracted_data) # Returns list of dicts matching product_schema
Setting force_json_response=True is critical for schema mode, ensuring the LLM returns parseable JSON rather than markdown-formatted code blocks or explanatory text.
Asynchronous Bulk Processing
For high-throughput scenarios, use AsyncCrawler with AsyncCrawlerRunConfig (both defined in crawl4ai/config.py) to process multiple URLs concurrently with asyncio.gather:
import os
import asyncio
from crawl4ai import AsyncCrawler, AsyncCrawlerRunConfig, LLMExtractionStrategy, LLMConfig
article_schema = {
"headline": "string",
"author": "string",
"date": "string",
"content": "string"
}
strategy = LLMExtractionStrategy(
llm_config=LLMConfig(
provider="anthropic/claude-3-opus-20240229",
api_token=os.getenv("ANTHROPIC_API_KEY")
),
schema=article_schema,
extraction_type="schema",
force_json_response=True
)
async def extract_from_url(url):
cfg = AsyncCrawlerRunConfig(
extraction_strategy=strategy,
bypass_cache=True
)
async with AsyncCrawler(url=url) as crawler:
result = await crawler.run(cfg)
return result.extracted_data
async def main():
urls = [
"https://news.ycombinator.com/item?id=39582325",
"https://example.com/blog/ai-trends"
]
results = await asyncio.gather(*(extract_from_url(u) for u in urls))
for url, data in zip(urls, results):
print(f"{url} -> {data}")
asyncio.run(main())
Type-Safe Extraction with Pydantic
For production environments, define schemas using Pydantic models to enable validation, autocomplete, and type checking. Convert the model to a JSON schema dictionary using model_json_schema():
from pydantic import BaseModel, Field
from crawl4ai import LLMExtractionStrategy, LLMConfig
class ProductSchema(BaseModel):
title: str = Field(..., description="Product title")
price: float = Field(..., description="Price in USD")
in_stock: bool = Field(..., description="Availability flag")
rating: float = Field(..., description="Average rating out of 5")
review_count: int = Field(..., description="Number of reviews")
strategy = LLMExtractionStrategy(
llm_config=LLMConfig(
provider="openai/gpt-4o-mini",
api_token=os.getenv("OPENAI_API_KEY")
),
schema=ProductSchema.model_json_schema(),
extraction_type="schema",
force_json_response=True,
verbose=True
)
Key Configuration Files and Utilities
Understanding the underlying source files helps with debugging and advanced customization:
-
crawl4ai/extraction_strategy.py: Contains theLLMExtractionStrategyclass and prompt templates includingPROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTIONused whenextraction_type="schema". -
crawl4ai/prompts.py: Defines theLLMConfigdataclass for provider credentials, model identifiers, and retry configuration. -
crawl4ai/utils.py: Implementsperform_completion_with_backofffor resilient API calls with exponential backoff, plusTokenUsagetracking for monitoring costs. -
crawl4ai/config.py: HousesCrawlerRunConfigandAsyncCrawlerRunConfig, the containers that bind the extraction strategy to the crawler runtime.
Summary
- Schema Definition: Define extraction targets using Python dictionaries or Pydantic models to enable strict typing and validation.
- Strategy Configuration: Instantiate
LLMExtractionStrategywithLLMConfigcontaining provider details; always setforce_json_response=Truefor reliable JSON parsing. - Crawler Integration: Pass the strategy to
CrawlerRunConfig(synchronous) orAsyncCrawlerRunConfig(asynchronous); extracted data appears inresult.extracted_data. - Resilience: The framework automatically handles rate limiting and retries through
perform_completion_with_backoffincrawl4ai/utils.py.
Frequently Asked Questions
What is the difference between block mode and schema mode in LLMExtractionStrategy?
Block mode extracts free-form text segments without structural constraints, suitable for content summarization or open-ended analysis. Schema mode (activated by providing a schema parameter) forces the LLM to return valid JSON matching your specified structure, enabling direct database insertion or API responses. The constructor automatically sets extraction_type="schema" when a schema is detected.
How does Crawl4AI handle LLM API failures or rate limits?
The perform_completion_with_backoff function in crawl4ai/utils.py implements exponential backoff and retry logic for all LLM API calls. Configure connection timeouts and maximum retry attempts through the LLMConfig class parameters when initializing your extraction strategy.
Can I use local LLM models or Azure OpenAI instead of commercial APIs?
Yes. The provider field in LLMConfig accepts various string identifiers including local model endpoints, "openai/gpt-4o-mini", "anthropic/claude-3-opus-20240229", or custom base URLs. Ensure your API token and endpoint configuration match your provider's requirements.
Why is my extracted_data empty or containing markdown instead of JSON?
Verify that force_json_response=True is set in your LLMExtractionStrategy initialization. Without this flag, the LLM may return markdown-formatted JSON or explanatory text that fails parsing. Additionally, check that your schema JSON is valid and that the crawled HTML actually contains the information requested in your schema definition.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →