# How Humanizer's File Mode Preserves Code Blocks, Inline Code, Commands, and Link Targets When Editing Prose

> Discover how Humanizer's file mode preserves code blocks, inline code, commands, and link targets when editing prose. Learn how it selectively rewrites only natural language.

- Repository: [Siqi Chen/humanizer](https://github.com/blader/humanizer)
- Tags: deep-dive
- Published: 2026-09-12

---

**Humanizer's file mode uses a markdown-aware parser to rewrite only natural-language prose while detecting and excluding fenced code blocks, inline code spans, shell commands, file paths, YAML front-matter, and link targets from any modifications.**

Humanizer is an open-source tool from the `blader/humanizer` repository designed to refine markdown documents without breaking their technical functionality. When operating in **file mode**, the tool processes entire files specified by filename but applies its editing pipeline exclusively to natural-language sections. This targeted approach ensures that code remains compilable, commands remain executable, and hyperlinks remain functional after the surrounding prose has been humanized.

## How File Mode Detects and Protects Non-Prose Elements

The preservation system relies on a declarative specification combined with token-based parsing to identify protected content types. According to the source code in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) at lines 50-53, the tool explicitly categorizes specific markdown constructs as immutable during the rewrite process.

### Fenced Code Blocks and Inline Code

When Humanizer processes a document, it scans for **fenced code blocks** (delimited by triple backticks) and **inline code spans** (wrapped in single backticks). These elements are flagged as protected tokens and copied verbatim to the output buffer. This mechanism ensures that syntax highlighting, indentation, and internal string literals remain intact regardless of how extensively the surrounding prose is rewritten.

### Commands, Paths, and YAML Front-Matter

The preservation logic extends to **shell-style commands**, **file-system paths**, and **YAML front-matter blocks** that appear at the beginning of documents. These elements often contain precise syntax where minor textual changes would break functionality. The parser identifies these constructs using contextual patterns and excludes them from the rewrite algorithm entirely, maintaining their exact literal content.

### Link Targets and Raw Data Sections

**Markdown link targets** inside parentheses (e.g., `[text](url)`) and **raw data sections** such as JSON or CSV blocks receive specific protection. While the visible link text may be processed if it constitutes prose, the URL target itself remains unchanged. Similarly, structured data blocks are identified by their formatting and copied directly to the output without modification.

## The Technical Implementation

The core preservation logic is defined declaratively in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) with the instruction: *"Change prose only. Keep code blocks, inline code, commands, paths, YAML metadata, data, and link targets unchanged."* This specification drives the behavior of the markdown-aware parser implemented in the file mode pipeline.

The system architecture involves several key components:

- **[`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md)** contains the full description of file mode and the declarative preservation rules
- **[`agents/openai.yaml`](https://github.com/blader/humanizer/blob/main/agents/openai.yaml)** defines the OpenAI-compatible agent wrapper that loads and executes the skill specifications
- **[`scripts/validate-package.py`](https://github.com/blader/humanizer/blob/main/scripts/validate-package.py)** checks that skill files, including [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md), adhere to the required structure and that preservation rules remain correctly formatted

When invoked via `humanizer filename.md`, the tool executes a three-step process:

1. **Tokenization**: The markdown-aware parser loads the full document and classifies content into token types (prose, code, metadata, links)
2. **Selective Rewrite**: The algorithm applies humanization exclusively to prose tokens while copying protected tokens directly to the output buffer
3. **Atomic Write**: The modified content is written back to the original file, preserving all technical elements exactly as originally authored

## Practical Examples of Element Preservation

### Editing a Markdown File with Code Blocks

When processing a file containing mixed content, fenced code blocks remain syntactically identical:

```bash
humanizer article.md

```

Input:

```markdown
Here is a Python function:

```python
def calculate():
    return 42

```

Use this code to compute values.

```

After humanization, the prose sentences are refined, but the fenced Python block maintains its original indentation and syntax.

### Protecting Inline Commands

Inline code containing shell commands is preserved precisely:

```markdown
Check status with `git status` before proceeding.

```

After processing, the `git status` command remains unchanged within its backtick delimiters, ensuring the instruction remains executable.

### Maintaining YAML and Hyperlinks

Documents with metadata headers retain their functional components:

```yaml
---
title: API Documentation
version: 2.1
---

Read the [full specification](https://api.example.com/docs).

```

Humanizer leaves the YAML front-matter block and the URL `https://api.example.com/docs` untouched, potentially modifying only the visible text "full specification" if it requires stylistic improvement.

## Summary

- **Declarative rules** in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) (lines 50-53) explicitly define which elements to preserve during file mode processing
- The **markdown-aware parser** tokenizes documents and classifies content as either prose or protected technical elements
- **Protected elements** include fenced code blocks, inline code, shell commands, file paths, YAML front-matter, raw data sections, and link targets
- The **agent wrapper** in [`agents/openai.yaml`](https://github.com/blader/humanizer/blob/main/agents/openai.yaml) loads these specifications, while [`scripts/validate-package.py`](https://github.com/blader/humanizer/blob/main/scripts/validate-package.py) ensures structural integrity
- File mode guarantees that code remains compilable, commands remain runnable, and links remain functional after prose humanization

## Frequently Asked Questions

### Does Humanizer modify the visible text of markdown links while preserving the URL?

Yes, Humanizer may rewrite the visible link text (the content inside square brackets) if it constitutes natural-language prose, but it always preserves the URL target inside the parentheses unchanged. This ensures hyperlinks remain functional while their descriptive text can be refined for clarity.

### How does Humanizer distinguish between prose and inline code?

The parser uses **delimiter-based detection** to identify inline code spans—any text wrapped in single backticks is automatically classified as a protected token. According to the specification in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md), content within backtick delimiters is treated as immutable code or commands regardless of semantic context.

### Will YAML front-matter be reformatted during humanization?

No, YAML front-matter blocks at the beginning of markdown files are explicitly excluded from the rewrite step. The preservation system recognizes these metadata sections by their delimiter patterns (`---`) and copies them verbatim to the output, maintaining exact key-value formatting.

### Can file mode handle documents containing JSON or CSV data sections?

Yes, the preservation system explicitly includes **raw data sections** such as JSON or CSV blocks. When the parser encounters these structured data formats—whether fenced or inline—it copies them directly to the output buffer without modification, ensuring data integrity while humanizing surrounding explanatory text.