# How Distilly Parses Markdown Files: Regex-Based Section Extraction Explained

> Distilly parses Markdown files with regex for efficient section extraction. Discover how this library enables structured research without external parsers.

- Repository: [Tianyi Zhou/distilly](https://github.com/titanwings/distilly)
- Tags: how-to-guide
- Published: 2026-09-10

---

**Distilly parses Markdown files using compiled regular expressions to detect level-2 headings, enabling idempotent section merging and structured research extraction without relying on external parsing libraries.**

The `titanwings/distilly` repository treats all skill-related content as Markdown documents, employing lightweight regex utilities to locate, extract, and replace structural elements. Understanding how Distilly parses Markdown files reveals a streamlined approach that prioritizes speed and determinism over heavy dependencies, making it ideal for AI skill generation workflows.

## Core Markdown Parsing Pipeline

Distilly’s parsing strategy centers on three distinct stages: detecting section boundaries, merging incremental updates, and extracting structured research data.

### Section Heading Detection with `SECTION_HEADING_RE`

In [`tools/skill_writer.py`](https://github.com/titanwings/distilly/blob/main/tools/skill_writer.py) (lines 14-15), Distilly defines a compiled multiline regex to identify logical content chunks:

```python
SECTION_HEADING_RE = re.compile(r'^##\s+.+$', re.MULTILINE)

```

This pattern matches level-2 headings (`##`) that serve as delimiters for skill sections. When processing a document, Distilly applies this regex using `re.finditer` to capture each heading’s text, start index, and the content slice ending just before the next heading (or EOF).

### Idempotent Patch Merging via `merge_markdown_patch`

The `merge_markdown_patch` function (lines 317-334 in [`tools/skill_writer.py`](https://github.com/titanwings/distilly/blob/main/tools/skill_writer.py)) handles incremental updates by treating patches as collections of heading-delimited sections. When a user supplies a Markdown fragment, the utility walks through each heading in the patch and either **replaces** the matching section in the existing document or **appends** the new section if no match exists.

This logic guarantees that updates are idempotent—running the same patch twice produces identical results without creating duplicate sections.

### Structured Research Extraction

For research notes, [`tools/research/merge_research.py`](https://github.com/titanwings/distilly/blob/main/tools/research/merge_research.py) (lines 24-25) implements a capture group pattern:

```python
SECTION_PATTERN = re.compile(r'^##\s+(.+)$', re.MULTILINE)

```

This extracts section names (e.g., "Evidence", "Contradictions", "Key Findings") to determine how subsequent bullet lines (`- `) should be classified. The parser then walks line-by-line, counting bullets into categories such as *contradictions*, *inferences*, *patterns*, or *gaps*, while additional regexes—`URL_PATTERN`, `SOURCE_WEIGHT_PATTERN`, and `TIMESTAMP_PATTERN`—gather metadata for the final Research Summary.

## Technical Implementation Details

### The Section Replacement Algorithm

For every heading detected in a patch, Distilly searches the existing document using an escaped regex search:

```python
re.search(rf"(?m)^{re.escape(heading)}\s*$", merged)

```

If the heading exists, the algorithm locates the end of the original section (the next `##` heading or the end of the file) and surgically replaces that slice with the new patch content. If the heading does **not** exist, the patch section is appended to the bottom of the file.

### Post-Merge Cleanup

After processing all headings, `merge_markdown_patch` trims excess whitespace and ensures a single trailing newline. The resulting string is valid Markdown ready for artifact generation or further manipulation.

### Rendering Final Skill Artifacts

Once patches are merged, Distilly renders the final deliverables—[`SKILL.md`](https://github.com/titanwings/distilly/blob/main/SKILL.md), work-only, and persona-only variants—through `render_combined_skill`, `render_work_skill`, and `render_persona_skill` (lines 58-75 in [`tools/skill_writer.py`](https://github.com/titanwings/distilly/blob/main/tools/skill_writer.py)). These functions interpolate the merged content into language-specific templates that are themselves plain Markdown files potentially containing YAML front-matter.

## Practical Code Examples

### Merging a Markdown Patch

This example demonstrates how to invoke Distilly’s patch merging logic to update a skill’s persona file:

```python
from pathlib import Path
from skill_writer import merge_markdown_patch

# Existing markdown content

existing = Path("persona.md").read_text(encoding="utf-8")

# Patch containing updated sections

patch = """

## Personality

- Friendly
- Curious

## Goals

- Deliver accurate answers
"""

# Merge replaces matching headings or appends new ones

updated = merge_markdown_patch(existing, patch)
print(updated)

```

### Summarizing Research Files

To process raw research notes into a structured summary:

```python
from merge_research import summarize_research_files, collect_markdown_files
from pathlib import Path

research_root = Path("my_skill/knowledge/research")
files = collect_markdown_files(research_root)
summary_md = summarize_research_files(files)

# Write the consolidated summary

(Path(research_root) / "merged" / "summary.md").write_text(
    summary_md, encoding="utf-8"
)

```

## Summary

- **Section Detection**: Distilly uses `SECTION_HEADING_RE` (`^##\s+.+$`) in [`tools/skill_writer.py`](https://github.com/titanwings/distilly/blob/main/tools/skill_writer.py) (lines 14-15) to identify level-2 headings as logical content boundaries.
- **Idempotent Merging**: The `merge_markdown_patch` function (lines 317-334) performs surgical updates by replacing existing sections or appending new ones, preventing duplicates.
- **Research Parsing**: [`tools/research/merge_research.py`](https://github.com/titanwings/distilly/blob/main/tools/research/merge_research.py) (lines 24-25) leverages `SECTION_PATTERN` with capture groups to classify research sections and extract metadata via specialized regexes.
- **Artifact Generation**: Final skills are rendered through `render_combined_skill` and related functions in [`tools/skill_writer.py`](https://github.com/titanwings/distilly/blob/main/tools/skill_writer.py) (lines 58-75), which template the merged Markdown into distributable formats.

## Frequently Asked Questions

### What regex pattern does Distilly use to identify Markdown sections?

Distilly compiles the pattern `^##\s+.+$` with the `re.MULTILINE` flag to detect level-2 headings. This regex is stored as `SECTION_HEADING_RE` in [`tools/skill_writer.py`](https://github.com/titanwings/distilly/blob/main/tools/skill_writer.py) and is used to split documents into manageable chunks for processing.

### How does Distilly handle duplicate sections when merging Markdown patches?

The `merge_markdown_patch` function checks for exact heading matches using escaped regex searches. If a heading exists in the target document, it replaces the entire old section—from that heading to the next `##` or EOF—with the new content. If the heading is absent, the section is appended to the file bottom, ensuring the operation remains idempotent and duplicate-free.

### Can Distilly parse Markdown files with front-matter?

Yes. While the core parsing pipeline targets `##` delimited content blocks, Distilly handles YAML front-matter in generated files through separate utilities in [`tools/install_generated_skill_common.py`](https://github.com/titanwings/distilly/blob/main/tools/install_generated_skill_common.py). The regex-based section detection operates on the body content after front-matter processing.

### Where does Distilly store the parsing logic for research notes?

Research-specific parsing resides in [`tools/research/merge_research.py`](https://github.com/titanwings/distilly/blob/main/tools/research/merge_research.py), which uses `SECTION_PATTERN` (`^##\s+(.+)$`) to extract lower-cased section names. It then walks bullet lines to classify content into contradictions, inferences, or gaps while extracting URLs, timestamps, and source-weight annotations for the final summary generation.