How to Extract Data Using CSS Selectors with JsonCssExtractionStrategy in crawl4ai

JsonCssExtractionStrategy enables structured data extraction from HTML by mapping CSS selectors to JSON schema fields, utilizing BeautifulSoup's select() method to parse and extract text, attributes, or HTML content from specific DOM elements.

The crawl4ai library provides a robust framework for transforming unstructured web content into structured datasets. When you extract data using CSS selectors with JsonCssExtractionStrategy in crawl4ai, you leverage a concrete implementation of the library's JSON-based extraction engine that parses HTML with BeautifulSoup and traverses the DOM using familiar CSS selector syntax.

Architecture of JsonCssExtractionStrategy

JsonCssExtractionStrategy inherits from JsonElementExtractionStrategy, the abstract base class defined in crawl4ai/extraction_strategy.py. This separation of concerns allows the same extraction logic to support both CSS and XPath selector strategies while maintaining a consistent interface.

The concrete CSS implementation resides at line 1436 of crawl4ai/extraction_strategy.py. It specializes the base class by implementing CSS-specific element selection using BeautifulSoup's select() method. For automatic schema creation, the library provides an LLM-driven helper at line 1344 that constructs valid CSS-based schemas from sample HTML content.

The CSS Extraction Pipeline

The extraction process follows a systematic pipeline to transform HTML into structured JSON:

  1. HTML Parsing: The _parse_html method creates a BeautifulSoup object using the lxml parser for fast DOM traversal.

  2. Base Element Selection: _get_base_elements executes soup.select(baseSelector) to identify repeating containers such as product cards or article elements.

  3. Relative Field Extraction: For each base element, _extract_item processes the fields array, where _get_elements runs element.select(selector) relative to the current base element.

  4. Data Transformation: Raw values undergo optional processing via _apply_transform (supporting operations like lowercase, strip, or capitalize), while computed fields leverage _compute_field with eval or callable functions.

Each schema field supports multiple extraction types: text (default), attribute (requires an attribute key), html, regex, or computed.

Practical Implementation Examples

Manual Schema Definition

Define your extraction schema explicitly to target specific DOM structures. The baseSelector identifies container elements, while individual fields use selectors relative to those containers.

from crawl4ai import Crawl, JsonCssExtractionStrategy

schema = {
    "name": "product_schema",
    "baseSelector": ".product-card",
    "baseFields": [],
    "fields": [
        {"name": "title", "selector": ".title", "type": "text"},
        {"name": "price", "selector": ".price", "type": "text", "transform": "strip"},
        {"name": "link", "selector": "a.buy", "type": "attribute", "attribute": "href"},
        {"name": "image", "selector": "img", "type": "attribute", "attribute": "src"}
    ]
}

strategy = JsonCssExtractionStrategy(schema)
crawler = Crawl()
result = crawler.run(url="https://example.com/shop", extraction_strategy=strategy)

print(result)  # Returns list of dictionaries, one per product card

Auto-Generating Schemas with LLMs

For dynamic or unfamiliar page structures, use the generate_schema class method to automatically build extraction schemas using Large Language Models.

from crawl4ai import Crawl, JsonCssExtractionStrategy, create_llm_config

html = Crawl().fetch_html("https://example.com/blog")

schema = JsonCssExtractionStrategy.generate_schema(
    html=html,
    schema_type="CSS",
    query="Extract the article title, author, and publication date",
    llm_config=create_llm_config()
)

strategy = JsonCssExtractionStrategy(schema)
result = Crawl().run(url="https://example.com/blog", extraction_strategy=strategy)

The generate_schema method processes the HTML through perform_completion_with_backoff using prompts defined in crawl4ai/prompts.py, returning a valid schema that you can inspect or modify before extraction.

Async Integration with FastAPI

The Crawl class provides arun for asynchronous operations, making it compatible with modern web frameworks while keeping the CSS extraction logic thread-safe.

from fastapi import FastAPI
from crawl4ai import Crawl, JsonCssExtractionStrategy

app = FastAPI()

@app.get("/scrape")
async def scrape(url: str):
    strategy = JsonCssExtractionStrategy(schema=my_schema)
    crawl = Crawl()
    data = await crawl.arun(url=url, extraction_strategy=strategy)
    return {"extracted": data}

The arun method internally executes strategy.run within an asyncio-friendly thread pool, ensuring non-blocking I/O while maintaining the computational efficiency of BeautifulSoup's CSS selection.

Key Source Files and Implementation Details

Several core files comprise the CSS extraction pipeline in the unclecode/crawl4ai repository:

  • crawl4ai/extraction_strategy.py: Contains both the abstract JsonElementExtractionStrategy and concrete JsonCssExtractionStrategy implementations. This file defines the parsing logic, selector handling via _get_elements, field extraction methods (_get_element_text, _get_element_attribute), and the generate_schema helper at line 1344.

  • crawl4ai/__init__.py: Provides the public API entry points including the Crawl class and create_llm_config utility used in the examples above.

  • crawl4ai/prompts.py: Houses the JSON_SCHEMA_BUILDER templates and system prompts that power the LLM-driven schema generation feature.

  • crawl4ai/utils.py: Supplies helper functions for HTML sanitization, safe eval execution for computed fields, and LLM interaction utilities used throughout the extraction process.

Summary

  • JsonCssExtractionStrategy extends JsonElementExtractionStrategy to provide CSS-based DOM targeting using BeautifulSoup's select() method.
  • The extraction process relies on a baseSelector to identify container elements, with field selectors evaluated relative to each container.
  • Schema definitions support multiple extraction types including text, attributes, HTML content, and computed values with optional transformations.
  • The generate_schema method at line 1344 of extraction_strategy.py enables automatic schema creation via LLM analysis of sample HTML.
  • Both synchronous (run) and asynchronous (arun) execution paths are available through the Crawl class in crawl4ai/__init__.py.

Frequently Asked Questions

What is the difference between baseSelector and field selectors?

The baseSelector defines the root container for each data record (such as a product card or article), while field selectors target specific elements within that container. According to the implementation in crawl4ai/extraction_strategy.py, the strategy first collects all elements matching baseSelector, then evaluates each field's selector relative to those base elements using element.select().

Can JsonCssExtractionStrategy extract multiple values from a single selector?

Yes. When a CSS selector matches multiple elements within a base container, the _get_elements method returns all matches. By default, the strategy handles these according to the field definition, and you can process arrays by configuring the schema appropriately or using computed fields to aggregate results.

How does the transform parameter work in field definitions?

The transform parameter accepts string values like lowercase, uppercase, strip, or capitalize that are applied via _apply_transform after extraction. This occurs in the _extract_item method before the final value is assigned to the output dictionary, allowing immediate data normalization during the scraping process.

Is it possible to extract data from JavaScript-rendered pages?

While JsonCssExtractionStrategy itself operates on static HTML parsed by BeautifulSoup, the crawl4ai framework handles JavaScript rendering upstream in the Crawl class. The strategy receives already-rendered HTML, meaning you can target elements generated by client-side JavaScript as long as the crawler executes the page with JavaScript support enabled before extraction begins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →