# How Humanizer Manages and Formats Hyphens, En Dashes, and Em Dashes

> Learn how Humanizer expertly manages hyphens, en dashes, and em dashes using specific pattern matching rules. Discover how it formats punctuation and handles compound adjectives for cleaner text.

- Repository: [Siqi Chen/humanizer](https://github.com/blader/humanizer)
- Tags: how-to-guide
- Published: 2026-09-13

---

**Humanizer manages hyphens, en dashes, and em dashes through two distinct pattern-matching rules defined in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md): Pattern 8 replaces dash connectors with commas, periods, or other punctuation, while Pattern 10 conditionally strips hyphens from compound adjectives based on grammatical position.**

The **blader/humanizer** repository is a pure-Markdown skill designed to rewrite AI-generated prose into natural, human-like text. Its approach to **managing and formatting hyphens, en dashes, and em dashes** relies on regex-based pattern scanning with strict context guards to prevent modifying technical syntax.

## The Core Dash Handling Patterns

Humanizer applies specific transformations through sequential pattern matching defined in the core skill definition file.

### Pattern 8: Dashes as Universal Connectors

Located at [[`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) lines 59-63](https://github.com/blader/humanizer/blob/main/SKILL.md#L59), this rule detects any dash character—including the hyphen (`-`), en dash (`–`), em dash (`—`), or double-hyphen (`--`)—when used as a connector between clauses. The pattern substitutes these dashes with grammatically appropriate punctuation such as periods, commas, colons, or parentheses, or rewrites the sentence entirely.

This rule executes **only outside of code blocks, inline code, commands, paths, and URLs** to preserve technical accuracy.

### Pattern 10: Hyphenated Pairs Everywhere

Defined at [[`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) lines 77-81](https://github.com/blader/humanizer/blob/main/SKILL.md#L77), this pattern targets compound adjectives like *cross-functional* or *high-quality* throughout the text. The logic applies a grammatical position test:

- **Before a noun**: The hyphen is retained (e.g., *a high-quality report*).
- **After a noun**: The hyphen is removed (e.g., *the report is high quality*).

## The Processing Pipeline

Humanizer processes text through a multi-stage pipeline to ensure accurate punctuation formatting.

### Text Tokenization and Pattern Scanning

The system receives raw text and tokenizes it line-by-line. It runs regular-expression checks corresponding to each pattern, matching the characters `—`, `–`, `-`, or the sequence `--` for dash detection.

### Context Guards to Protect Technical Syntax

Before applying any replacement, Humanizer verifies the match is **not** inside a fenced code block, inline backticks, URL, or file path. This **context guard** prevents accidental modification of command-line flags, file paths, or URL parameters that legitimately require hyphens.

### Replacement Strategy and Punctuation Selection

When Pattern 8 triggers, the engine analyzes surrounding words to select the most natural punctuation. Em dashes used as ornamental separators collapse to single spaces or disappear entirely. For Pattern 10, the engine consults an internal list of common compound adjectives to determine grammatical placement and applies the appropriate transformation.

## Practical Transformation Examples

The following examples demonstrate how Humanizer transforms dash-heavy AI output into clean prose while preserving technical syntax.

### Replacing Em Dashes with Commas

```markdown

# Before

The new policy — announced without warning — affects thousands of workers.

# After

The new policy, announced without warning, affects thousands of workers.

```

### Normalizing En Dashes and Double-Hyphens

```markdown

# Before

The changes – long overdue according to critics – will take effect immediately.

The updates -- critical for the launch -- were deployed yesterday.

# After

The changes, long overdue according to critics, will take effect immediately.

The updates, critical for the launch, were deployed yesterday.

```

### Handling Hyphenated Compounds Based on Position

```markdown

# Before (hyphen before noun - retained)

We need a high-quality report for the board.

# After

We need a high-quality report for the board.

# Before (hyphen after noun - removed)

The report is high-quality and data-driven.

# After

The report is high quality and data driven.

```

### Preserving Hyphens in Code Blocks

```markdown

# Before

Run `npm install --save-dev` to add the dependency.

# After

Run `npm install --save-dev` to add the dependency.

```

## Key Source Files and Implementation Details

The dash formatting logic resides in three primary files within the repository:

- **[`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md)**: Contains the core pattern definitions at lines 59-63 for dash connectors and lines 77-81 for hyphenated pairs, plus context guard specifications.
- **[`README.md`](https://github.com/blader/humanizer/blob/main/README.md)**: Provides user-facing documentation explaining how the skill processes punctuation in the "How it works" section.
- **[`.claude-plugin/plugin.json`](https://github.com/blader/humanizer/blob/main/.claude-plugin/plugin.json)**: Points agents to the correct skill file location, ensuring the patterns load properly.

## Summary

- **Pattern 8** at [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) lines 59-63 replaces em dashes, en dashes, and double-hyphens with commas, periods, colons, or parentheses when they connect clauses.
- **Pattern 10** at [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) lines 77-81 removes hyphens from compound adjectives when they appear after nouns but preserves them before nouns.
- **Context guards** prevent modifications inside code blocks, inline backticks, URLs, and file paths.
- The transformation process relies on line-by-line tokenization and regex-based pattern matching.
- Three key files define the behavior: [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) for rules, [`README.md`](https://github.com/blader/humanizer/blob/main/README.md) for documentation, and [`.claude-plugin/plugin.json`](https://github.com/blader/humanizer/blob/main/.claude-plugin/plugin.json) for plugin metadata.

## Frequently Asked Questions

### How does Humanizer distinguish between hyphens in code and prose?

Humanizer applies a **context guard** that checks whether a potential match occurs inside a fenced code block, inline backticks, URL, or file path. If the match falls within these technical contexts, the pattern skips the replacement and preserves the original punctuation.

### What punctuation does Humanizer use to replace em dashes and en dashes?

According to **Pattern 8** in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md), Humanizer substitutes connecting dashes with **commas, periods, colons, or parentheses** depending on the surrounding grammatical structure. Ornamental separators typically collapse to single spaces or are removed entirely.

### Why does Humanizer remove hyphens from compound adjectives after nouns?

**Pattern 10** implements standard English grammar rules where compound modifiers require hyphens only when they precede the noun they modify (attributive position). When they follow the noun (predicative position), as in *the report is high quality*, the hyphens are unnecessary and are removed to match natural human writing conventions.

### Where are the dash formatting rules defined in the Humanizer repository?

The specific regex patterns and replacement logic reside in **[`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md)** at lines 59-63 for universal dash connectors and lines 77-81 for hyphenated pairs. The [`.claude-plugin/plugin.json`](https://github.com/blader/humanizer/blob/main/.claude-plugin/plugin.json) file ensures agents load these rules correctly.