How Distilly Parses Markdown Files: Regex-Based Section Extraction Explained

Distilly parses Markdown files using compiled regular expressions to detect level-2 headings, enabling idempotent section merging and structured research extraction without relying on external parsing libraries.

The titanwings/distilly repository treats all skill-related content as Markdown documents, employing lightweight regex utilities to locate, extract, and replace structural elements. Understanding how Distilly parses Markdown files reveals a streamlined approach that prioritizes speed and determinism over heavy dependencies, making it ideal for AI skill generation workflows.

Core Markdown Parsing Pipeline

Distilly’s parsing strategy centers on three distinct stages: detecting section boundaries, merging incremental updates, and extracting structured research data.

Section Heading Detection with SECTION_HEADING_RE

In tools/skill_writer.py (lines 14-15), Distilly defines a compiled multiline regex to identify logical content chunks:

SECTION_HEADING_RE = re.compile(r'^##\s+.+$', re.MULTILINE)

This pattern matches level-2 headings (##) that serve as delimiters for skill sections. When processing a document, Distilly applies this regex using re.finditer to capture each heading’s text, start index, and the content slice ending just before the next heading (or EOF).

Idempotent Patch Merging via merge_markdown_patch

The merge_markdown_patch function (lines 317-334 in tools/skill_writer.py) handles incremental updates by treating patches as collections of heading-delimited sections. When a user supplies a Markdown fragment, the utility walks through each heading in the patch and either replaces the matching section in the existing document or appends the new section if no match exists.

This logic guarantees that updates are idempotent—running the same patch twice produces identical results without creating duplicate sections.

Structured Research Extraction

For research notes, tools/research/merge_research.py (lines 24-25) implements a capture group pattern:

SECTION_PATTERN = re.compile(r'^##\s+(.+)$', re.MULTILINE)

This extracts section names (e.g., "Evidence", "Contradictions", "Key Findings") to determine how subsequent bullet lines (- ) should be classified. The parser then walks line-by-line, counting bullets into categories such as contradictions, inferences, patterns, or gaps, while additional regexes—URL_PATTERN, SOURCE_WEIGHT_PATTERN, and TIMESTAMP_PATTERN—gather metadata for the final Research Summary.

Technical Implementation Details

The Section Replacement Algorithm

For every heading detected in a patch, Distilly searches the existing document using an escaped regex search:

re.search(rf"(?m)^{re.escape(heading)}\s*$", merged)

If the heading exists, the algorithm locates the end of the original section (the next ## heading or the end of the file) and surgically replaces that slice with the new patch content. If the heading does not exist, the patch section is appended to the bottom of the file.

Post-Merge Cleanup

After processing all headings, merge_markdown_patch trims excess whitespace and ensures a single trailing newline. The resulting string is valid Markdown ready for artifact generation or further manipulation.

Rendering Final Skill Artifacts

Once patches are merged, Distilly renders the final deliverables—SKILL.md, work-only, and persona-only variants—through render_combined_skill, render_work_skill, and render_persona_skill (lines 58-75 in tools/skill_writer.py). These functions interpolate the merged content into language-specific templates that are themselves plain Markdown files potentially containing YAML front-matter.

Practical Code Examples

Merging a Markdown Patch

This example demonstrates how to invoke Distilly’s patch merging logic to update a skill’s persona file:

from pathlib import Path
from skill_writer import merge_markdown_patch

# Existing markdown content

existing = Path("persona.md").read_text(encoding="utf-8")

# Patch containing updated sections

patch = """

## Personality

- Friendly
- Curious

## Goals

- Deliver accurate answers
"""

# Merge replaces matching headings or appends new ones

updated = merge_markdown_patch(existing, patch)
print(updated)

Summarizing Research Files

To process raw research notes into a structured summary:

from merge_research import summarize_research_files, collect_markdown_files
from pathlib import Path

research_root = Path("my_skill/knowledge/research")
files = collect_markdown_files(research_root)
summary_md = summarize_research_files(files)

# Write the consolidated summary

(Path(research_root) / "merged" / "summary.md").write_text(
    summary_md, encoding="utf-8"
)

Summary

  • Section Detection: Distilly uses SECTION_HEADING_RE (^##\s+.+$) in tools/skill_writer.py (lines 14-15) to identify level-2 headings as logical content boundaries.
  • Idempotent Merging: The merge_markdown_patch function (lines 317-334) performs surgical updates by replacing existing sections or appending new ones, preventing duplicates.
  • Research Parsing: tools/research/merge_research.py (lines 24-25) leverages SECTION_PATTERN with capture groups to classify research sections and extract metadata via specialized regexes.
  • Artifact Generation: Final skills are rendered through render_combined_skill and related functions in tools/skill_writer.py (lines 58-75), which template the merged Markdown into distributable formats.

Frequently Asked Questions

What regex pattern does Distilly use to identify Markdown sections?

Distilly compiles the pattern ^##\s+.+$ with the re.MULTILINE flag to detect level-2 headings. This regex is stored as SECTION_HEADING_RE in tools/skill_writer.py and is used to split documents into manageable chunks for processing.

How does Distilly handle duplicate sections when merging Markdown patches?

The merge_markdown_patch function checks for exact heading matches using escaped regex searches. If a heading exists in the target document, it replaces the entire old section—from that heading to the next ## or EOF—with the new content. If the heading is absent, the section is appended to the file bottom, ensuring the operation remains idempotent and duplicate-free.

Can Distilly parse Markdown files with front-matter?

Yes. While the core parsing pipeline targets ## delimited content blocks, Distilly handles YAML front-matter in generated files through separate utilities in tools/install_generated_skill_common.py. The regex-based section detection operates on the body content after front-matter processing.

Where does Distilly store the parsing logic for research notes?

Research-specific parsing resides in tools/research/merge_research.py, which uses SECTION_PATTERN (^##\s+(.+)$) to extract lower-cased section names. It then walks bullet lines to classify content into contradictions, inferences, or gaps while extracting URLs, timestamps, and source-weight annotations for the final summary generation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →