How Humanizer's File Mode Preserves Code Blocks, Inline Code, Commands, and Link Targets When Editing Prose

Humanizer's file mode uses a markdown-aware parser to rewrite only natural-language prose while detecting and excluding fenced code blocks, inline code spans, shell commands, file paths, YAML front-matter, and link targets from any modifications.

Humanizer is an open-source tool from the blader/humanizer repository designed to refine markdown documents without breaking their technical functionality. When operating in file mode, the tool processes entire files specified by filename but applies its editing pipeline exclusively to natural-language sections. This targeted approach ensures that code remains compilable, commands remain executable, and hyperlinks remain functional after the surrounding prose has been humanized.

How File Mode Detects and Protects Non-Prose Elements

The preservation system relies on a declarative specification combined with token-based parsing to identify protected content types. According to the source code in SKILL.md at lines 50-53, the tool explicitly categorizes specific markdown constructs as immutable during the rewrite process.

Fenced Code Blocks and Inline Code

When Humanizer processes a document, it scans for fenced code blocks (delimited by triple backticks) and inline code spans (wrapped in single backticks). These elements are flagged as protected tokens and copied verbatim to the output buffer. This mechanism ensures that syntax highlighting, indentation, and internal string literals remain intact regardless of how extensively the surrounding prose is rewritten.

Commands, Paths, and YAML Front-Matter

The preservation logic extends to shell-style commands, file-system paths, and YAML front-matter blocks that appear at the beginning of documents. These elements often contain precise syntax where minor textual changes would break functionality. The parser identifies these constructs using contextual patterns and excludes them from the rewrite algorithm entirely, maintaining their exact literal content.

Markdown link targets inside parentheses (e.g., [text](url)) and raw data sections such as JSON or CSV blocks receive specific protection. While the visible link text may be processed if it constitutes prose, the URL target itself remains unchanged. Similarly, structured data blocks are identified by their formatting and copied directly to the output without modification.

The Technical Implementation

The core preservation logic is defined declaratively in SKILL.md with the instruction: "Change prose only. Keep code blocks, inline code, commands, paths, YAML metadata, data, and link targets unchanged." This specification drives the behavior of the markdown-aware parser implemented in the file mode pipeline.

The system architecture involves several key components:

  • SKILL.md contains the full description of file mode and the declarative preservation rules
  • agents/openai.yaml defines the OpenAI-compatible agent wrapper that loads and executes the skill specifications
  • scripts/validate-package.py checks that skill files, including SKILL.md, adhere to the required structure and that preservation rules remain correctly formatted

When invoked via humanizer filename.md, the tool executes a three-step process:

  1. Tokenization: The markdown-aware parser loads the full document and classifies content into token types (prose, code, metadata, links)
  2. Selective Rewrite: The algorithm applies humanization exclusively to prose tokens while copying protected tokens directly to the output buffer
  3. Atomic Write: The modified content is written back to the original file, preserving all technical elements exactly as originally authored

Practical Examples of Element Preservation

Editing a Markdown File with Code Blocks

When processing a file containing mixed content, fenced code blocks remain syntactically identical:

humanizer article.md

Input:

Here is a Python function:

```python
def calculate():
    return 42

Use this code to compute values.


After humanization, the prose sentences are refined, but the fenced Python block maintains its original indentation and syntax.

### Protecting Inline Commands

Inline code containing shell commands is preserved precisely:

```markdown
Check status with `git status` before proceeding.

After processing, the git status command remains unchanged within its backtick delimiters, ensuring the instruction remains executable.

Documents with metadata headers retain their functional components:

---
title: API Documentation
version: 2.1
---

Read the [full specification](https://api.example.com/docs).

Humanizer leaves the YAML front-matter block and the URL https://api.example.com/docs untouched, potentially modifying only the visible text "full specification" if it requires stylistic improvement.

Summary

  • Declarative rules in SKILL.md (lines 50-53) explicitly define which elements to preserve during file mode processing
  • The markdown-aware parser tokenizes documents and classifies content as either prose or protected technical elements
  • Protected elements include fenced code blocks, inline code, shell commands, file paths, YAML front-matter, raw data sections, and link targets
  • The agent wrapper in agents/openai.yaml loads these specifications, while scripts/validate-package.py ensures structural integrity
  • File mode guarantees that code remains compilable, commands remain runnable, and links remain functional after prose humanization

Frequently Asked Questions

Yes, Humanizer may rewrite the visible link text (the content inside square brackets) if it constitutes natural-language prose, but it always preserves the URL target inside the parentheses unchanged. This ensures hyperlinks remain functional while their descriptive text can be refined for clarity.

How does Humanizer distinguish between prose and inline code?

The parser uses delimiter-based detection to identify inline code spans—any text wrapped in single backticks is automatically classified as a protected token. According to the specification in SKILL.md, content within backtick delimiters is treated as immutable code or commands regardless of semantic context.

Will YAML front-matter be reformatted during humanization?

No, YAML front-matter blocks at the beginning of markdown files are explicitly excluded from the rewrite step. The preservation system recognizes these metadata sections by their delimiter patterns (---) and copies them verbatim to the output, maintaining exact key-value formatting.

Can file mode handle documents containing JSON or CSV data sections?

Yes, the preservation system explicitly includes raw data sections such as JSON or CSV blocks. When the parser encounters these structured data formats—whether fenced or inline—it copies them directly to the output buffer without modification, ensuring data integrity while humanizing surrounding explanatory text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →