# How to Create a Custom Markdown Generation Strategy in Crawl4AI

> Learn how to create a custom Markdown generation strategy in Crawl4AI. Subclass MarkdownGenerationStrategy, implement generate_markdown, and inject your custom generator for tailored output.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: how-to-guide
- Published: 2026-03-05

---

**To create a custom markdown generation strategy in Crawl4AI, subclass `MarkdownGenerationStrategy` from [`crawl4ai/markdown_generation_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/markdown_generation_strategy.py), implement the `generate_markdown` method to return a `MarkdownGenerationResult`, and inject your instance into `CrawlerRunConfig(markdown_generator=...)` to override the default HTML-to-Markdown conversion.**

Crawl4AI separates content extraction from formatting logic through a pluggable strategy pattern. If you need to create a custom markdown generation strategy in Crawl4AI—whether to inject YAML front matter, change citation styles, or preserve specific HTML tags—you can extend the abstract base class and plug your implementation into the crawl configuration without modifying the core crawler code.

## Understanding the Markdown Generation Architecture

Crawl4AI decouples HTML processing from crawling by defining a strict contract in [`crawl4ai/markdown_generation_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/markdown_generation_strategy.py). The architecture centers on three primary components that handle conversion, configuration, and execution.

### The Abstract Base Class

`MarkdownGenerationStrategy` (defined at line 26 in [`crawl4ai/markdown_generation_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/markdown_generation_strategy.py)) establishes the interface for all markdown generators. It declares a single abstract method, `generate_markdown`, which receives raw HTML and metadata and must return a `MarkdownGenerationResult` from [`crawl4ai/models.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/models.py).

### The Default Implementation

`DefaultMarkdownGenerator` (line 55 in [`crawl4ai/markdown_generation_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/markdown_generation_strategy.py)) provides the built-in conversion logic. It utilizes `CustomHTML2Text` from [`crawl4ai/html2text/__init__.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/html2text/__init__.py) for HTML parsing and implements `convert_links_to_citations` to transform URLs into reference-style links. This generator produces raw markdown, citation-enhanced versions, and filtered "fit" markdown when content filters are supplied.

### Configuration Integration

The `CrawlerRunConfig` class in [`crawl4ai/async_configs.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_configs.py) exposes the `markdown_generator` field (line 22). When `AsyncWebCrawler` executes in [`crawl4ai/async_webcrawler.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_webcrawler.py), it instantiates your custom strategy or falls back to `DefaultMarkdownGenerator`, then invokes `generate_markdown` with the scraped HTML, base URL, and processing options.

## Implementing a Custom Markdown Generation Strategy

Follow these three steps to create a custom markdown generation strategy in Crawl4AI that integrates seamlessly with the existing crawling pipeline.

### Step 1 - Subclass MarkdownGenerationStrategy

Create a new Python class that inherits from `MarkdownGenerationStrategy`. Import the base class and the result model to ensure type safety and adherence to the interface contract.

```python
from crawl4ai.markdown_generation_strategy import (
    MarkdownGenerationStrategy,
    MarkdownGenerationResult,
)
from crawl4ai.html2text import CustomHTML2Text

```

### Step 2 - Implement the generate_markdown Method

Your implementation must accept the following parameters:
- `input_html`: The raw or cleaned HTML string extracted from the page
- `base_url`: The source URL for resolving relative links
- `html2text_options`: Dictionary of configuration options for `CustomHTML2Text`
- `content_filter`: Optional filter object for generating "fit" markdown
- `citations`: Boolean flag indicating whether to process citations
- `**kwargs`: Additional arguments for future extensibility

Return a `MarkdownGenerationResult` containing at minimum the `raw_markdown` field. You may also populate `markdown_with_citations`, `references_markdown`, `fit_markdown`, and `fit_html` depending on your processing logic.

### Step 3 - Configure the Crawler

Instantiate your custom strategy and pass it to the crawler configuration. The async crawler will automatically delegate HTML-to-Markdown conversion to your implementation.

```python
from crawl4ai.async_configs import CrawlerRunConfig
from crawl4ai.async_webcrawler import AsyncWebCrawler
from my_module import MyCustomGenerator

config = CrawlerRunConfig(
    markdown_generator=MyCustomGenerator(),
)

async with AsyncWebCrawler() as crawler:
    result = await crawler.arun("https://example.com", config=config)

```

## Practical Examples

These runnable examples demonstrate common customization patterns when you create a custom markdown generation strategy in Crawl4AI.

### Example 1 - Adding a Static Header to Crawl Output

This generator prepends a branded H1 header to every crawled page, useful for organizing reports or adding timestamps:

```python

# my_markdown_strategy.py

from crawl4ai.markdown_generation_strategy import (
    MarkdownGenerationStrategy,
    MarkdownGenerationResult,
    DefaultMarkdownGenerator,
)
from crawl4ai.html2text import CustomHTML2Text


class SimpleHeaderGenerator(MarkdownStrategy):
    """
    Prepends a static H1 header before the converted markdown.
    """

    def generate_markdown(
        self,
        input_html: str,
        base_url: str = "",
        html2text_options: dict | None = None,
        content_filter=None,
        citations: bool = True,
        **kwargs,
    ) -> MarkdownGenerationResult:
        # Initialize the HTML-to-markdown converter

        h = CustomHTML2Text(baseurl=base_url)
        if html2text_options:
            h.update_params(**html2text_options)

        # Convert HTML to markdown

        raw_md = h.handle(input_html)

        # Prepend custom header

        header = "# My Custom Crawl Report\n\n"

        raw_md = header + raw_md

        # Reuse default citation handling

        if citations:
            md, refs = DefaultMarkdownGenerator().convert_links_to_citations(
                raw_md, base_url
            )
        else:
            md, refs = raw_md, ""

        return MarkdownGenerationResult(
            raw_markdown=raw_md,
            markdown_with_citations=md,
            references_markdown=refs,
            fit_markdown=None,
            fit_html=None,
        )

```

**Usage:**

```python
from crawl4ai.async_webcrawler import AsyncWebCrawler
from crawl4ai.async_configs import CrawlerRunConfig
from my_markdown_strategy import SimpleHeaderGenerator

config = CrawlerRunConfig(
    markdown_generator=SimpleHeaderGenerator(),
)

async def run():
    async with AsyncWebCrawler() as crawler:
        result = await crawler.arun("https://example.com", config=config)
        print(result.markdown.raw_markdown)  # Contains the custom H1 header

```

### Example 2 - Generating Markdown with YAML Front Matter

For static site generators or documentation pipelines, you may need YAML front matter containing crawl metadata:

```python

# advanced_markdown.py

import yaml
from datetime import datetime
from crawl4ai.markdown_generation_strategy import (
    MarkdownGenerationStrategy,
    MarkdownGenerationResult,
    LINK_PATTERN,
    fast_urljoin,
)
from crawl4ai.html2text import CustomHTML2Text


class FrontMatterMarkdownGenerator(MarkdownGenerationStrategy):
    """
    Generates markdown with YAML front matter and numbered footnote citations.
    """

    def generate_markdown(
        self,
        input_html: str,
        base_url: str = "",
        html2text_options: dict | None = None,
        content_filter=None,
        citations: bool = True,
        **kwargs,
    ) -> MarkdownGenerationResult:
        # Convert HTML to markdown

        h = CustomHTML2Text(baseurl=base_url)
        if html2text_options:
            h.update_params(**html2text_options)
        raw_md = h.handle(input_html)

        # Build YAML front matter

        fm = {
            "source_url": base_url,
            "crawled_at": datetime.utcnow().isoformat() + "Z",
            "title": kwargs.get("title", "Untitled"),
        }
        front_matter = "---\n" + yaml.safe_dump(fm, sort_keys=False) + "---\n\n"
        raw_md = front_matter + raw_md

        # Custom citation handling with numbered footnotes

        if citations:
            link_map = {}
            parts = []
            last_end = 0
            counter = 1

            for match in LINK_PATTERN.finditer(raw_md):
                parts.append(raw_md[last_end:match.start()])
                text, url, title = match.groups()
                if base_url and not url.startswith(("http://", "https://", "mailto:")):
                    url = fast_urljoin(base_url, url)
                if url not in link_map:
                    link_map[url] = counter
                    counter += 1
                footnote_num = link_map[url]
                parts.append(f"{text}[{footnote_num}]")
                last_end = match.end()
            parts.append(raw_md[last_end:])
            markdown_with_citations = "".join(parts)

            # Build footnote reference list

            references = "\n\n---\n\n"
            for url, num in sorted(link_map.items(), key=lambda kv: kv[1]):
                references += f"[{num}]: {url}\n"
        else:
            markdown_with_citations = raw_md
            references = ""

        return MarkdownGenerationResult(
            raw_markdown=raw_md,
            markdown_with_citations=markdown_with_citations,
            references_markdown=references,
            fit_markdown=None,
            fit_html=None,
        )

```

**Usage:**

```python
config = CrawlerRunConfig(
    markdown_generator=FrontMatterMarkdownGenerator(),
)

# Run the crawler as shown in the previous example.

```

These examples demonstrate how to inject static content, customize citation formats, and leverage existing utilities while maintaining full control over the markdown output structure.

## Summary

To create a custom markdown generation strategy in Crawl4AI, follow these essential steps:

- **Subclass `MarkdownGenerationStrategy`** from [`crawl4ai/markdown_generation_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/markdown_generation_strategy.py) to define your custom logic and adhere to the interface contract.
- **Implement `generate_markdown`** with the standard signature to receive HTML content and metadata, then return a populated `MarkdownGenerationResult` from [`crawl4ai/models.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/models.py).
- **Leverage existing utilities** such as `CustomHTML2Text` for HTML parsing and `fast_urljoin` for URL resolution to avoid reimplementing core functionality.
- **Inject via configuration** by passing your strategy instance to `CrawlerRunConfig(markdown_generator=...)` defined in [`crawl4ai/async_configs.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_configs.py), allowing the async crawler in [`crawl4ai/async_webcrawler.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_webcrawler.py) to invoke your custom logic automatically.

## Frequently Asked Questions

### What is the difference between raw markdown and markdown with citations in Crawl4AI?

Raw markdown represents the direct HTML-to-text conversion performed by `CustomHTML2Text` without any link processing. Markdown with citations replaces inline URLs with numbered reference markers (such as `[1]`) and appends a reference list at the end of the document. The `DefaultMarkdownGenerator` produces both variants using the `convert_links_to_citations` method, and custom strategies can implement alternative citation styles by overriding this logic or implementing their own URL handling.

### Can I reuse the default HTML-to-markdown conversion in my custom strategy?

Yes. Import `CustomHTML2Text` from `crawl4ai/html2text` and instantiate it with `CustomHTML2Text(baseurl=base_url)`. Call `handle(input_html)` to obtain the raw markdown string, then apply your custom post-processing such as prepending headers or modifying citation formats. You can also pass `html2text_options` to the `update_params()` method to control tag preservation and code block handling without modifying the underlying HTML parser.

### How do I handle relative URLs when generating markdown with custom citations?

Use the `fast_urljoin` utility imported from `crawl4ai/markdown_generation_strategy`. This helper resolves relative URLs against the `base_url` parameter passed to `generate_markdown`. When iterating over links using `LINK_PATTERN`, check if the URL starts with `http://`, `https://`, or `mailto:`; if not, pass it through `fast_urljoin(base_url, url)` before storing it in your citation mapping to ensure all links are absolute in the final output.

### Where do I configure the markdown generator when running the crawler?

Set the `markdown_generator` parameter in `CrawlerRunConfig` (defined in [`crawl4ai/async_configs.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_configs.py)). When you instantiate `CrawlerRunConfig(markdown_generator=YourCustomStrategy())` and pass it to `AsyncWebCrawler.arun()`, the crawler automatically delegates HTML-to-Markdown conversion to your strategy's `generate_markdown` method instead of the default implementation. This configuration-based injection allows you to switch strategies per-crawl without modifying the crawler's core logic in [`crawl4ai/async_webcrawler.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_webcrawler.py).