# How to Extract Data Using CSS Selectors with JsonCssExtractionStrategy in crawl4ai

> Learn to extract structured data from HTML using CSS selectors and JsonCssExtractionStrategy in crawl4ai. Map CSS selectors to JSON schema fields with BeautifulSoup for efficient parsing.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: how-to-guide
- Published: 2026-03-05

---

**JsonCssExtractionStrategy enables structured data extraction from HTML by mapping CSS selectors to JSON schema fields, utilizing BeautifulSoup's `select()` method to parse and extract text, attributes, or HTML content from specific DOM elements.**

The crawl4ai library provides a robust framework for transforming unstructured web content into structured datasets. When you extract data using CSS selectors with JsonCssExtractionStrategy in crawl4ai, you leverage a concrete implementation of the library's JSON-based extraction engine that parses HTML with BeautifulSoup and traverses the DOM using familiar CSS selector syntax.

## Architecture of JsonCssExtractionStrategy

`JsonCssExtractionStrategy` inherits from `JsonElementExtractionStrategy`, the abstract base class defined in [`crawl4ai/extraction_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/extraction_strategy.py). This separation of concerns allows the same extraction logic to support both CSS and XPath selector strategies while maintaining a consistent interface.

The concrete CSS implementation resides at line 1436 of [`crawl4ai/extraction_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/extraction_strategy.py). It specializes the base class by implementing CSS-specific element selection using BeautifulSoup's `select()` method. For automatic schema creation, the library provides an LLM-driven helper at line 1344 that constructs valid CSS-based schemas from sample HTML content.

## The CSS Extraction Pipeline

The extraction process follows a systematic pipeline to transform HTML into structured JSON:

1. **HTML Parsing**: The `_parse_html` method creates a BeautifulSoup object using the `lxml` parser for fast DOM traversal.

2. **Base Element Selection**: `_get_base_elements` executes `soup.select(baseSelector)` to identify repeating containers such as product cards or article elements.

3. **Relative Field Extraction**: For each base element, `_extract_item` processes the `fields` array, where `_get_elements` runs `element.select(selector)` relative to the current base element.

4. **Data Transformation**: Raw values undergo optional processing via `_apply_transform` (supporting operations like `lowercase`, `strip`, or `capitalize`), while computed fields leverage `_compute_field` with `eval` or callable functions.

Each schema field supports multiple extraction types: `text` (default), `attribute` (requires an `attribute` key), `html`, `regex`, or `computed`.

## Practical Implementation Examples

### Manual Schema Definition

Define your extraction schema explicitly to target specific DOM structures. The `baseSelector` identifies container elements, while individual `fields` use selectors relative to those containers.

```python
from crawl4ai import Crawl, JsonCssExtractionStrategy

schema = {
    "name": "product_schema",
    "baseSelector": ".product-card",
    "baseFields": [],
    "fields": [
        {"name": "title", "selector": ".title", "type": "text"},
        {"name": "price", "selector": ".price", "type": "text", "transform": "strip"},
        {"name": "link", "selector": "a.buy", "type": "attribute", "attribute": "href"},
        {"name": "image", "selector": "img", "type": "attribute", "attribute": "src"}
    ]
}

strategy = JsonCssExtractionStrategy(schema)
crawler = Crawl()
result = crawler.run(url="https://example.com/shop", extraction_strategy=strategy)

print(result)  # Returns list of dictionaries, one per product card

```

### Auto-Generating Schemas with LLMs

For dynamic or unfamiliar page structures, use the `generate_schema` class method to automatically build extraction schemas using Large Language Models.

```python
from crawl4ai import Crawl, JsonCssExtractionStrategy, create_llm_config

html = Crawl().fetch_html("https://example.com/blog")

schema = JsonCssExtractionStrategy.generate_schema(
    html=html,
    schema_type="CSS",
    query="Extract the article title, author, and publication date",
    llm_config=create_llm_config()
)

strategy = JsonCssExtractionStrategy(schema)
result = Crawl().run(url="https://example.com/blog", extraction_strategy=strategy)

```

The `generate_schema` method processes the HTML through `perform_completion_with_backoff` using prompts defined in [`crawl4ai/prompts.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/prompts.py), returning a valid schema that you can inspect or modify before extraction.

### Async Integration with FastAPI

The `Crawl` class provides `arun` for asynchronous operations, making it compatible with modern web frameworks while keeping the CSS extraction logic thread-safe.

```python
from fastapi import FastAPI
from crawl4ai import Crawl, JsonCssExtractionStrategy

app = FastAPI()

@app.get("/scrape")
async def scrape(url: str):
    strategy = JsonCssExtractionStrategy(schema=my_schema)
    crawl = Crawl()
    data = await crawl.arun(url=url, extraction_strategy=strategy)
    return {"extracted": data}

```

The `arun` method internally executes `strategy.run` within an `asyncio`-friendly thread pool, ensuring non-blocking I/O while maintaining the computational efficiency of BeautifulSoup's CSS selection.

## Key Source Files and Implementation Details

Several core files comprise the CSS extraction pipeline in the unclecode/crawl4ai repository:

- **[`crawl4ai/extraction_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/extraction_strategy.py)**: Contains both the abstract `JsonElementExtractionStrategy` and concrete `JsonCssExtractionStrategy` implementations. This file defines the parsing logic, selector handling via `_get_elements`, field extraction methods (`_get_element_text`, `_get_element_attribute`), and the `generate_schema` helper at line 1344.

- **[`crawl4ai/__init__.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/__init__.py)**: Provides the public API entry points including the `Crawl` class and `create_llm_config` utility used in the examples above.

- **[`crawl4ai/prompts.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/prompts.py)**: Houses the `JSON_SCHEMA_BUILDER` templates and system prompts that power the LLM-driven schema generation feature.

- **[`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py)**: Supplies helper functions for HTML sanitization, safe `eval` execution for computed fields, and LLM interaction utilities used throughout the extraction process.

## Summary

- **JsonCssExtractionStrategy** extends `JsonElementExtractionStrategy` to provide CSS-based DOM targeting using BeautifulSoup's `select()` method.
- The extraction process relies on a `baseSelector` to identify container elements, with field selectors evaluated relative to each container.
- Schema definitions support multiple extraction types including text, attributes, HTML content, and computed values with optional transformations.
- The `generate_schema` method at line 1344 of [`extraction_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/extraction_strategy.py) enables automatic schema creation via LLM analysis of sample HTML.
- Both synchronous (`run`) and asynchronous (`arun`) execution paths are available through the `Crawl` class in [`crawl4ai/__init__.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/__init__.py).

## Frequently Asked Questions

### What is the difference between baseSelector and field selectors?

The `baseSelector` defines the root container for each data record (such as a product card or article), while field selectors target specific elements within that container. According to the implementation in [`crawl4ai/extraction_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/extraction_strategy.py), the strategy first collects all elements matching `baseSelector`, then evaluates each field's `selector` relative to those base elements using `element.select()`.

### Can JsonCssExtractionStrategy extract multiple values from a single selector?

Yes. When a CSS selector matches multiple elements within a base container, the `_get_elements` method returns all matches. By default, the strategy handles these according to the field definition, and you can process arrays by configuring the schema appropriately or using computed fields to aggregate results.

### How does the transform parameter work in field definitions?

The `transform` parameter accepts string values like `lowercase`, `uppercase`, `strip`, or `capitalize` that are applied via `_apply_transform` after extraction. This occurs in the `_extract_item` method before the final value is assigned to the output dictionary, allowing immediate data normalization during the scraping process.

### Is it possible to extract data from JavaScript-rendered pages?

While `JsonCssExtractionStrategy` itself operates on static HTML parsed by BeautifulSoup, the crawl4ai framework handles JavaScript rendering upstream in the `Crawl` class. The strategy receives already-rendered HTML, meaning you can target elements generated by client-side JavaScript as long as the crawler executes the page with JavaScript support enabled before extraction begins.