How to Create a Custom Markdown Generation Strategy in Crawl4AI
To create a custom markdown generation strategy in Crawl4AI, subclass MarkdownGenerationStrategy from crawl4ai/markdown_generation_strategy.py, implement the generate_markdown method to return a MarkdownGenerationResult, and inject your instance into CrawlerRunConfig(markdown_generator=...) to override the default HTML-to-Markdown conversion.
Crawl4AI separates content extraction from formatting logic through a pluggable strategy pattern. If you need to create a custom markdown generation strategy in Crawl4AI—whether to inject YAML front matter, change citation styles, or preserve specific HTML tags—you can extend the abstract base class and plug your implementation into the crawl configuration without modifying the core crawler code.
Understanding the Markdown Generation Architecture
Crawl4AI decouples HTML processing from crawling by defining a strict contract in crawl4ai/markdown_generation_strategy.py. The architecture centers on three primary components that handle conversion, configuration, and execution.
The Abstract Base Class
MarkdownGenerationStrategy (defined at line 26 in crawl4ai/markdown_generation_strategy.py) establishes the interface for all markdown generators. It declares a single abstract method, generate_markdown, which receives raw HTML and metadata and must return a MarkdownGenerationResult from crawl4ai/models.py.
The Default Implementation
DefaultMarkdownGenerator (line 55 in crawl4ai/markdown_generation_strategy.py) provides the built-in conversion logic. It utilizes CustomHTML2Text from crawl4ai/html2text/__init__.py for HTML parsing and implements convert_links_to_citations to transform URLs into reference-style links. This generator produces raw markdown, citation-enhanced versions, and filtered "fit" markdown when content filters are supplied.
Configuration Integration
The CrawlerRunConfig class in crawl4ai/async_configs.py exposes the markdown_generator field (line 22). When AsyncWebCrawler executes in crawl4ai/async_webcrawler.py, it instantiates your custom strategy or falls back to DefaultMarkdownGenerator, then invokes generate_markdown with the scraped HTML, base URL, and processing options.
Implementing a Custom Markdown Generation Strategy
Follow these three steps to create a custom markdown generation strategy in Crawl4AI that integrates seamlessly with the existing crawling pipeline.
Step 1 - Subclass MarkdownGenerationStrategy
Create a new Python class that inherits from MarkdownGenerationStrategy. Import the base class and the result model to ensure type safety and adherence to the interface contract.
from crawl4ai.markdown_generation_strategy import (
MarkdownGenerationStrategy,
MarkdownGenerationResult,
)
from crawl4ai.html2text import CustomHTML2Text
Step 2 - Implement the generate_markdown Method
Your implementation must accept the following parameters:
input_html: The raw or cleaned HTML string extracted from the pagebase_url: The source URL for resolving relative linkshtml2text_options: Dictionary of configuration options forCustomHTML2Textcontent_filter: Optional filter object for generating "fit" markdowncitations: Boolean flag indicating whether to process citations**kwargs: Additional arguments for future extensibility
Return a MarkdownGenerationResult containing at minimum the raw_markdown field. You may also populate markdown_with_citations, references_markdown, fit_markdown, and fit_html depending on your processing logic.
Step 3 - Configure the Crawler
Instantiate your custom strategy and pass it to the crawler configuration. The async crawler will automatically delegate HTML-to-Markdown conversion to your implementation.
from crawl4ai.async_configs import CrawlerRunConfig
from crawl4ai.async_webcrawler import AsyncWebCrawler
from my_module import MyCustomGenerator
config = CrawlerRunConfig(
markdown_generator=MyCustomGenerator(),
)
async with AsyncWebCrawler() as crawler:
result = await crawler.arun("https://example.com", config=config)
Practical Examples
These runnable examples demonstrate common customization patterns when you create a custom markdown generation strategy in Crawl4AI.
Example 1 - Adding a Static Header to Crawl Output
This generator prepends a branded H1 header to every crawled page, useful for organizing reports or adding timestamps:
# my_markdown_strategy.py
from crawl4ai.markdown_generation_strategy import (
MarkdownGenerationStrategy,
MarkdownGenerationResult,
DefaultMarkdownGenerator,
)
from crawl4ai.html2text import CustomHTML2Text
class SimpleHeaderGenerator(MarkdownStrategy):
"""
Prepends a static H1 header before the converted markdown.
"""
def generate_markdown(
self,
input_html: str,
base_url: str = "",
html2text_options: dict | None = None,
content_filter=None,
citations: bool = True,
**kwargs,
) -> MarkdownGenerationResult:
# Initialize the HTML-to-markdown converter
h = CustomHTML2Text(baseurl=base_url)
if html2text_options:
h.update_params(**html2text_options)
# Convert HTML to markdown
raw_md = h.handle(input_html)
# Prepend custom header
header = "# My Custom Crawl Report\n\n"
raw_md = header + raw_md
# Reuse default citation handling
if citations:
md, refs = DefaultMarkdownGenerator().convert_links_to_citations(
raw_md, base_url
)
else:
md, refs = raw_md, ""
return MarkdownGenerationResult(
raw_markdown=raw_md,
markdown_with_citations=md,
references_markdown=refs,
fit_markdown=None,
fit_html=None,
)
Usage:
from crawl4ai.async_webcrawler import AsyncWebCrawler
from crawl4ai.async_configs import CrawlerRunConfig
from my_markdown_strategy import SimpleHeaderGenerator
config = CrawlerRunConfig(
markdown_generator=SimpleHeaderGenerator(),
)
async def run():
async with AsyncWebCrawler() as crawler:
result = await crawler.arun("https://example.com", config=config)
print(result.markdown.raw_markdown) # Contains the custom H1 header
Example 2 - Generating Markdown with YAML Front Matter
For static site generators or documentation pipelines, you may need YAML front matter containing crawl metadata:
# advanced_markdown.py
import yaml
from datetime import datetime
from crawl4ai.markdown_generation_strategy import (
MarkdownGenerationStrategy,
MarkdownGenerationResult,
LINK_PATTERN,
fast_urljoin,
)
from crawl4ai.html2text import CustomHTML2Text
class FrontMatterMarkdownGenerator(MarkdownGenerationStrategy):
"""
Generates markdown with YAML front matter and numbered footnote citations.
"""
def generate_markdown(
self,
input_html: str,
base_url: str = "",
html2text_options: dict | None = None,
content_filter=None,
citations: bool = True,
**kwargs,
) -> MarkdownGenerationResult:
# Convert HTML to markdown
h = CustomHTML2Text(baseurl=base_url)
if html2text_options:
h.update_params(**html2text_options)
raw_md = h.handle(input_html)
# Build YAML front matter
fm = {
"source_url": base_url,
"crawled_at": datetime.utcnow().isoformat() + "Z",
"title": kwargs.get("title", "Untitled"),
}
front_matter = "---\n" + yaml.safe_dump(fm, sort_keys=False) + "---\n\n"
raw_md = front_matter + raw_md
# Custom citation handling with numbered footnotes
if citations:
link_map = {}
parts = []
last_end = 0
counter = 1
for match in LINK_PATTERN.finditer(raw_md):
parts.append(raw_md[last_end:match.start()])
text, url, title = match.groups()
if base_url and not url.startswith(("http://", "https://", "mailto:")):
url = fast_urljoin(base_url, url)
if url not in link_map:
link_map[url] = counter
counter += 1
footnote_num = link_map[url]
parts.append(f"{text}[{footnote_num}]")
last_end = match.end()
parts.append(raw_md[last_end:])
markdown_with_citations = "".join(parts)
# Build footnote reference list
references = "\n\n---\n\n"
for url, num in sorted(link_map.items(), key=lambda kv: kv[1]):
references += f"[{num}]: {url}\n"
else:
markdown_with_citations = raw_md
references = ""
return MarkdownGenerationResult(
raw_markdown=raw_md,
markdown_with_citations=markdown_with_citations,
references_markdown=references,
fit_markdown=None,
fit_html=None,
)
Usage:
config = CrawlerRunConfig(
markdown_generator=FrontMatterMarkdownGenerator(),
)
# Run the crawler as shown in the previous example.
These examples demonstrate how to inject static content, customize citation formats, and leverage existing utilities while maintaining full control over the markdown output structure.
Summary
To create a custom markdown generation strategy in Crawl4AI, follow these essential steps:
- Subclass
MarkdownGenerationStrategyfromcrawl4ai/markdown_generation_strategy.pyto define your custom logic and adhere to the interface contract. - Implement
generate_markdownwith the standard signature to receive HTML content and metadata, then return a populatedMarkdownGenerationResultfromcrawl4ai/models.py. - Leverage existing utilities such as
CustomHTML2Textfor HTML parsing andfast_urljoinfor URL resolution to avoid reimplementing core functionality. - Inject via configuration by passing your strategy instance to
CrawlerRunConfig(markdown_generator=...)defined incrawl4ai/async_configs.py, allowing the async crawler incrawl4ai/async_webcrawler.pyto invoke your custom logic automatically.
Frequently Asked Questions
What is the difference between raw markdown and markdown with citations in Crawl4AI?
Raw markdown represents the direct HTML-to-text conversion performed by CustomHTML2Text without any link processing. Markdown with citations replaces inline URLs with numbered reference markers (such as [1]) and appends a reference list at the end of the document. The DefaultMarkdownGenerator produces both variants using the convert_links_to_citations method, and custom strategies can implement alternative citation styles by overriding this logic or implementing their own URL handling.
Can I reuse the default HTML-to-markdown conversion in my custom strategy?
Yes. Import CustomHTML2Text from crawl4ai/html2text and instantiate it with CustomHTML2Text(baseurl=base_url). Call handle(input_html) to obtain the raw markdown string, then apply your custom post-processing such as prepending headers or modifying citation formats. You can also pass html2text_options to the update_params() method to control tag preservation and code block handling without modifying the underlying HTML parser.
How do I handle relative URLs when generating markdown with custom citations?
Use the fast_urljoin utility imported from crawl4ai/markdown_generation_strategy. This helper resolves relative URLs against the base_url parameter passed to generate_markdown. When iterating over links using LINK_PATTERN, check if the URL starts with http://, https://, or mailto:; if not, pass it through fast_urljoin(base_url, url) before storing it in your citation mapping to ensure all links are absolute in the final output.
Where do I configure the markdown generator when running the crawler?
Set the markdown_generator parameter in CrawlerRunConfig (defined in crawl4ai/async_configs.py). When you instantiate CrawlerRunConfig(markdown_generator=YourCustomStrategy()) and pass it to AsyncWebCrawler.arun(), the crawler automatically delegates HTML-to-Markdown conversion to your strategy's generate_markdown method instead of the default implementation. This configuration-based injection allows you to switch strategies per-crawl without modifying the crawler's core logic in crawl4ai/async_webcrawler.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →