How to Extract Data Using CSS Selectors with JsonCssExtractionStrategy in crawl4ai
JsonCssExtractionStrategy enables structured data extraction from HTML by mapping CSS selectors to JSON schema fields, utilizing BeautifulSoup's select() method to parse and extract text, attributes, or HTML content from specific DOM elements.
The crawl4ai library provides a robust framework for transforming unstructured web content into structured datasets. When you extract data using CSS selectors with JsonCssExtractionStrategy in crawl4ai, you leverage a concrete implementation of the library's JSON-based extraction engine that parses HTML with BeautifulSoup and traverses the DOM using familiar CSS selector syntax.
Architecture of JsonCssExtractionStrategy
JsonCssExtractionStrategy inherits from JsonElementExtractionStrategy, the abstract base class defined in crawl4ai/extraction_strategy.py. This separation of concerns allows the same extraction logic to support both CSS and XPath selector strategies while maintaining a consistent interface.
The concrete CSS implementation resides at line 1436 of crawl4ai/extraction_strategy.py. It specializes the base class by implementing CSS-specific element selection using BeautifulSoup's select() method. For automatic schema creation, the library provides an LLM-driven helper at line 1344 that constructs valid CSS-based schemas from sample HTML content.
The CSS Extraction Pipeline
The extraction process follows a systematic pipeline to transform HTML into structured JSON:
-
HTML Parsing: The
_parse_htmlmethod creates a BeautifulSoup object using thelxmlparser for fast DOM traversal. -
Base Element Selection:
_get_base_elementsexecutessoup.select(baseSelector)to identify repeating containers such as product cards or article elements. -
Relative Field Extraction: For each base element,
_extract_itemprocesses thefieldsarray, where_get_elementsrunselement.select(selector)relative to the current base element. -
Data Transformation: Raw values undergo optional processing via
_apply_transform(supporting operations likelowercase,strip, orcapitalize), while computed fields leverage_compute_fieldwithevalor callable functions.
Each schema field supports multiple extraction types: text (default), attribute (requires an attribute key), html, regex, or computed.
Practical Implementation Examples
Manual Schema Definition
Define your extraction schema explicitly to target specific DOM structures. The baseSelector identifies container elements, while individual fields use selectors relative to those containers.
from crawl4ai import Crawl, JsonCssExtractionStrategy
schema = {
"name": "product_schema",
"baseSelector": ".product-card",
"baseFields": [],
"fields": [
{"name": "title", "selector": ".title", "type": "text"},
{"name": "price", "selector": ".price", "type": "text", "transform": "strip"},
{"name": "link", "selector": "a.buy", "type": "attribute", "attribute": "href"},
{"name": "image", "selector": "img", "type": "attribute", "attribute": "src"}
]
}
strategy = JsonCssExtractionStrategy(schema)
crawler = Crawl()
result = crawler.run(url="https://example.com/shop", extraction_strategy=strategy)
print(result) # Returns list of dictionaries, one per product card
Auto-Generating Schemas with LLMs
For dynamic or unfamiliar page structures, use the generate_schema class method to automatically build extraction schemas using Large Language Models.
from crawl4ai import Crawl, JsonCssExtractionStrategy, create_llm_config
html = Crawl().fetch_html("https://example.com/blog")
schema = JsonCssExtractionStrategy.generate_schema(
html=html,
schema_type="CSS",
query="Extract the article title, author, and publication date",
llm_config=create_llm_config()
)
strategy = JsonCssExtractionStrategy(schema)
result = Crawl().run(url="https://example.com/blog", extraction_strategy=strategy)
The generate_schema method processes the HTML through perform_completion_with_backoff using prompts defined in crawl4ai/prompts.py, returning a valid schema that you can inspect or modify before extraction.
Async Integration with FastAPI
The Crawl class provides arun for asynchronous operations, making it compatible with modern web frameworks while keeping the CSS extraction logic thread-safe.
from fastapi import FastAPI
from crawl4ai import Crawl, JsonCssExtractionStrategy
app = FastAPI()
@app.get("/scrape")
async def scrape(url: str):
strategy = JsonCssExtractionStrategy(schema=my_schema)
crawl = Crawl()
data = await crawl.arun(url=url, extraction_strategy=strategy)
return {"extracted": data}
The arun method internally executes strategy.run within an asyncio-friendly thread pool, ensuring non-blocking I/O while maintaining the computational efficiency of BeautifulSoup's CSS selection.
Key Source Files and Implementation Details
Several core files comprise the CSS extraction pipeline in the unclecode/crawl4ai repository:
-
crawl4ai/extraction_strategy.py: Contains both the abstractJsonElementExtractionStrategyand concreteJsonCssExtractionStrategyimplementations. This file defines the parsing logic, selector handling via_get_elements, field extraction methods (_get_element_text,_get_element_attribute), and thegenerate_schemahelper at line 1344. -
crawl4ai/__init__.py: Provides the public API entry points including theCrawlclass andcreate_llm_configutility used in the examples above. -
crawl4ai/prompts.py: Houses theJSON_SCHEMA_BUILDERtemplates and system prompts that power the LLM-driven schema generation feature. -
crawl4ai/utils.py: Supplies helper functions for HTML sanitization, safeevalexecution for computed fields, and LLM interaction utilities used throughout the extraction process.
Summary
- JsonCssExtractionStrategy extends
JsonElementExtractionStrategyto provide CSS-based DOM targeting using BeautifulSoup'sselect()method. - The extraction process relies on a
baseSelectorto identify container elements, with field selectors evaluated relative to each container. - Schema definitions support multiple extraction types including text, attributes, HTML content, and computed values with optional transformations.
- The
generate_schemamethod at line 1344 ofextraction_strategy.pyenables automatic schema creation via LLM analysis of sample HTML. - Both synchronous (
run) and asynchronous (arun) execution paths are available through theCrawlclass incrawl4ai/__init__.py.
Frequently Asked Questions
What is the difference between baseSelector and field selectors?
The baseSelector defines the root container for each data record (such as a product card or article), while field selectors target specific elements within that container. According to the implementation in crawl4ai/extraction_strategy.py, the strategy first collects all elements matching baseSelector, then evaluates each field's selector relative to those base elements using element.select().
Can JsonCssExtractionStrategy extract multiple values from a single selector?
Yes. When a CSS selector matches multiple elements within a base container, the _get_elements method returns all matches. By default, the strategy handles these according to the field definition, and you can process arrays by configuring the schema appropriately or using computed fields to aggregate results.
How does the transform parameter work in field definitions?
The transform parameter accepts string values like lowercase, uppercase, strip, or capitalize that are applied via _apply_transform after extraction. This occurs in the _extract_item method before the final value is assigned to the output dictionary, allowing immediate data normalization during the scraping process.
Is it possible to extract data from JavaScript-rendered pages?
While JsonCssExtractionStrategy itself operates on static HTML parsed by BeautifulSoup, the crawl4ai framework handles JavaScript rendering upstream in the Crawl class. The strategy receives already-rendered HTML, meaning you can target elements generated by client-side JavaScript as long as the crawler executes the page with JavaScript support enabled before extraction begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →