How to Create a Custom Markdown Generation Strategy in Crawl4AI

To create a custom markdown generation strategy in Crawl4AI, subclass MarkdownGenerationStrategy from crawl4ai/markdown_generation_strategy.py, implement the generate_markdown method to return a MarkdownGenerationResult, and inject your instance into CrawlerRunConfig(markdown_generator=...) to override the default HTML-to-Markdown conversion.

Crawl4AI separates content extraction from formatting logic through a pluggable strategy pattern. If you need to create a custom markdown generation strategy in Crawl4AI—whether to inject YAML front matter, change citation styles, or preserve specific HTML tags—you can extend the abstract base class and plug your implementation into the crawl configuration without modifying the core crawler code.

Understanding the Markdown Generation Architecture

Crawl4AI decouples HTML processing from crawling by defining a strict contract in crawl4ai/markdown_generation_strategy.py. The architecture centers on three primary components that handle conversion, configuration, and execution.

The Abstract Base Class

MarkdownGenerationStrategy (defined at line 26 in crawl4ai/markdown_generation_strategy.py) establishes the interface for all markdown generators. It declares a single abstract method, generate_markdown, which receives raw HTML and metadata and must return a MarkdownGenerationResult from crawl4ai/models.py.

The Default Implementation

DefaultMarkdownGenerator (line 55 in crawl4ai/markdown_generation_strategy.py) provides the built-in conversion logic. It utilizes CustomHTML2Text from crawl4ai/html2text/__init__.py for HTML parsing and implements convert_links_to_citations to transform URLs into reference-style links. This generator produces raw markdown, citation-enhanced versions, and filtered "fit" markdown when content filters are supplied.

Configuration Integration

The CrawlerRunConfig class in crawl4ai/async_configs.py exposes the markdown_generator field (line 22). When AsyncWebCrawler executes in crawl4ai/async_webcrawler.py, it instantiates your custom strategy or falls back to DefaultMarkdownGenerator, then invokes generate_markdown with the scraped HTML, base URL, and processing options.

Implementing a Custom Markdown Generation Strategy

Follow these three steps to create a custom markdown generation strategy in Crawl4AI that integrates seamlessly with the existing crawling pipeline.

Step 1 - Subclass MarkdownGenerationStrategy

Create a new Python class that inherits from MarkdownGenerationStrategy. Import the base class and the result model to ensure type safety and adherence to the interface contract.

from crawl4ai.markdown_generation_strategy import (
    MarkdownGenerationStrategy,
    MarkdownGenerationResult,
)
from crawl4ai.html2text import CustomHTML2Text

Step 2 - Implement the generate_markdown Method

Your implementation must accept the following parameters:

  • input_html: The raw or cleaned HTML string extracted from the page
  • base_url: The source URL for resolving relative links
  • html2text_options: Dictionary of configuration options for CustomHTML2Text
  • content_filter: Optional filter object for generating "fit" markdown
  • citations: Boolean flag indicating whether to process citations
  • **kwargs: Additional arguments for future extensibility

Return a MarkdownGenerationResult containing at minimum the raw_markdown field. You may also populate markdown_with_citations, references_markdown, fit_markdown, and fit_html depending on your processing logic.

Step 3 - Configure the Crawler

Instantiate your custom strategy and pass it to the crawler configuration. The async crawler will automatically delegate HTML-to-Markdown conversion to your implementation.

from crawl4ai.async_configs import CrawlerRunConfig
from crawl4ai.async_webcrawler import AsyncWebCrawler
from my_module import MyCustomGenerator

config = CrawlerRunConfig(
    markdown_generator=MyCustomGenerator(),
)

async with AsyncWebCrawler() as crawler:
    result = await crawler.arun("https://example.com", config=config)

Practical Examples

These runnable examples demonstrate common customization patterns when you create a custom markdown generation strategy in Crawl4AI.

Example 1 - Adding a Static Header to Crawl Output

This generator prepends a branded H1 header to every crawled page, useful for organizing reports or adding timestamps:


# my_markdown_strategy.py

from crawl4ai.markdown_generation_strategy import (
    MarkdownGenerationStrategy,
    MarkdownGenerationResult,
    DefaultMarkdownGenerator,
)
from crawl4ai.html2text import CustomHTML2Text


class SimpleHeaderGenerator(MarkdownStrategy):
    """
    Prepends a static H1 header before the converted markdown.
    """

    def generate_markdown(
        self,
        input_html: str,
        base_url: str = "",
        html2text_options: dict | None = None,
        content_filter=None,
        citations: bool = True,
        **kwargs,
    ) -> MarkdownGenerationResult:
        # Initialize the HTML-to-markdown converter

        h = CustomHTML2Text(baseurl=base_url)
        if html2text_options:
            h.update_params(**html2text_options)

        # Convert HTML to markdown

        raw_md = h.handle(input_html)

        # Prepend custom header

        header = "# My Custom Crawl Report\n\n"

        raw_md = header + raw_md

        # Reuse default citation handling

        if citations:
            md, refs = DefaultMarkdownGenerator().convert_links_to_citations(
                raw_md, base_url
            )
        else:
            md, refs = raw_md, ""

        return MarkdownGenerationResult(
            raw_markdown=raw_md,
            markdown_with_citations=md,
            references_markdown=refs,
            fit_markdown=None,
            fit_html=None,
        )

Usage:

from crawl4ai.async_webcrawler import AsyncWebCrawler
from crawl4ai.async_configs import CrawlerRunConfig
from my_markdown_strategy import SimpleHeaderGenerator

config = CrawlerRunConfig(
    markdown_generator=SimpleHeaderGenerator(),
)

async def run():
    async with AsyncWebCrawler() as crawler:
        result = await crawler.arun("https://example.com", config=config)
        print(result.markdown.raw_markdown)  # Contains the custom H1 header

Example 2 - Generating Markdown with YAML Front Matter

For static site generators or documentation pipelines, you may need YAML front matter containing crawl metadata:


# advanced_markdown.py

import yaml
from datetime import datetime
from crawl4ai.markdown_generation_strategy import (
    MarkdownGenerationStrategy,
    MarkdownGenerationResult,
    LINK_PATTERN,
    fast_urljoin,
)
from crawl4ai.html2text import CustomHTML2Text


class FrontMatterMarkdownGenerator(MarkdownGenerationStrategy):
    """
    Generates markdown with YAML front matter and numbered footnote citations.
    """

    def generate_markdown(
        self,
        input_html: str,
        base_url: str = "",
        html2text_options: dict | None = None,
        content_filter=None,
        citations: bool = True,
        **kwargs,
    ) -> MarkdownGenerationResult:
        # Convert HTML to markdown

        h = CustomHTML2Text(baseurl=base_url)
        if html2text_options:
            h.update_params(**html2text_options)
        raw_md = h.handle(input_html)

        # Build YAML front matter

        fm = {
            "source_url": base_url,
            "crawled_at": datetime.utcnow().isoformat() + "Z",
            "title": kwargs.get("title", "Untitled"),
        }
        front_matter = "---\n" + yaml.safe_dump(fm, sort_keys=False) + "---\n\n"
        raw_md = front_matter + raw_md

        # Custom citation handling with numbered footnotes

        if citations:
            link_map = {}
            parts = []
            last_end = 0
            counter = 1

            for match in LINK_PATTERN.finditer(raw_md):
                parts.append(raw_md[last_end:match.start()])
                text, url, title = match.groups()
                if base_url and not url.startswith(("http://", "https://", "mailto:")):
                    url = fast_urljoin(base_url, url)
                if url not in link_map:
                    link_map[url] = counter
                    counter += 1
                footnote_num = link_map[url]
                parts.append(f"{text}[{footnote_num}]")
                last_end = match.end()
            parts.append(raw_md[last_end:])
            markdown_with_citations = "".join(parts)

            # Build footnote reference list

            references = "\n\n---\n\n"
            for url, num in sorted(link_map.items(), key=lambda kv: kv[1]):
                references += f"[{num}]: {url}\n"
        else:
            markdown_with_citations = raw_md
            references = ""

        return MarkdownGenerationResult(
            raw_markdown=raw_md,
            markdown_with_citations=markdown_with_citations,
            references_markdown=references,
            fit_markdown=None,
            fit_html=None,
        )

Usage:

config = CrawlerRunConfig(
    markdown_generator=FrontMatterMarkdownGenerator(),
)

# Run the crawler as shown in the previous example.

These examples demonstrate how to inject static content, customize citation formats, and leverage existing utilities while maintaining full control over the markdown output structure.

Summary

To create a custom markdown generation strategy in Crawl4AI, follow these essential steps:

  • Subclass MarkdownGenerationStrategy from crawl4ai/markdown_generation_strategy.py to define your custom logic and adhere to the interface contract.
  • Implement generate_markdown with the standard signature to receive HTML content and metadata, then return a populated MarkdownGenerationResult from crawl4ai/models.py.
  • Leverage existing utilities such as CustomHTML2Text for HTML parsing and fast_urljoin for URL resolution to avoid reimplementing core functionality.
  • Inject via configuration by passing your strategy instance to CrawlerRunConfig(markdown_generator=...) defined in crawl4ai/async_configs.py, allowing the async crawler in crawl4ai/async_webcrawler.py to invoke your custom logic automatically.

Frequently Asked Questions

What is the difference between raw markdown and markdown with citations in Crawl4AI?

Raw markdown represents the direct HTML-to-text conversion performed by CustomHTML2Text without any link processing. Markdown with citations replaces inline URLs with numbered reference markers (such as [1]) and appends a reference list at the end of the document. The DefaultMarkdownGenerator produces both variants using the convert_links_to_citations method, and custom strategies can implement alternative citation styles by overriding this logic or implementing their own URL handling.

Can I reuse the default HTML-to-markdown conversion in my custom strategy?

Yes. Import CustomHTML2Text from crawl4ai/html2text and instantiate it with CustomHTML2Text(baseurl=base_url). Call handle(input_html) to obtain the raw markdown string, then apply your custom post-processing such as prepending headers or modifying citation formats. You can also pass html2text_options to the update_params() method to control tag preservation and code block handling without modifying the underlying HTML parser.

How do I handle relative URLs when generating markdown with custom citations?

Use the fast_urljoin utility imported from crawl4ai/markdown_generation_strategy. This helper resolves relative URLs against the base_url parameter passed to generate_markdown. When iterating over links using LINK_PATTERN, check if the URL starts with http://, https://, or mailto:; if not, pass it through fast_urljoin(base_url, url) before storing it in your citation mapping to ensure all links are absolute in the final output.

Where do I configure the markdown generator when running the crawler?

Set the markdown_generator parameter in CrawlerRunConfig (defined in crawl4ai/async_configs.py). When you instantiate CrawlerRunConfig(markdown_generator=YourCustomStrategy()) and pass it to AsyncWebCrawler.arun(), the crawler automatically delegates HTML-to-Markdown conversion to your strategy's generate_markdown method instead of the default implementation. This configuration-based injection allows you to switch strategies per-crawl without modifying the crawler's core logic in crawl4ai/async_webcrawler.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →