# How Humanizer Patterns Are Ordered by Strength and Frequency

> Discover how Humanizer orders its 25 detection patterns by strength and frequency. Learn how strong patterns trigger edits while weak ones require corroboration.

- Repository: [Siqi Chen/humanizer](https://github.com/blader/humanizer)
- Tags: deep-dive
- Published: 2026-09-10

---

**Humanizer arranges its 25 detection patterns from strongest (§1–§5) to weakest (§25), where high‑strength patterns trigger immediate edits on a single occurrence while weaker patterns require corroboration from multiple tells.**

The open‑source editing tool **blader/humanizer** embeds a strict hierarchy directly into its [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) file, ranking linguistic tells by their reliability and frequency. This systematic ordering ensures that editors act decisively on clear AI artifacts while avoiding over‑correction on ambiguous stylistic choices.

## The Five-Tier Strength Hierarchy in SKILL.md

The [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) file in the `blader/humanizer` repository organizes detection patterns into five graded sections, with **Section A** representing the most reliable indicators of AI‑generated text and **Section E** capturing residual artifacts that rarely affect meaning.

### Section A: Staging Instead of Stating (Patterns 1–5)

Section **A. Staging instead of stating** houses the five strongest and most frequent tells, explicitly numbered 1 through 5. According to the source documentation, these patterns “justify an edit on one sighting,” meaning a single occurrence is sufficient to trigger a revision. These include rhetorical constructions like “Not X but Y” framing and abrupt one‑line closers that signal automated staging rather than authentic narration.

### Section B: Rhythm by Rule (Moderate Strength)

Section **B. Rhythm by rule** introduces patterns that deviate from natural speech cadences, such as forced triads, repeated opening clauses, and dash over‑use. These tells are considered weaker than Section A markers; the documentation specifies they usually need “company from other tells” before an editor should act on them. A passage must exhibit multiple rhythmic violations before reaching the threshold for intervention.

### Section C: Inflation and Borrowed Authority

Section **C. Inflation and borrowed authority** groups linguistic tells that dress up ordinary observations with pseudo‑intellectual weight, including over‑used AI vocabulary, inflated significance markers, and vague bibliographic connections. These patterns rank below rhythmic violations in immediate priority but signal credibility issues when clustered.

### Section D: Formatting by Rule (Weak Alone)

Section **D. Formatting by rule** contains the lowest‑priority visual tells, such as bold decorations, decorative heading hierarchies, and curly quote usage. The source explicitly labels these as “weak alone,” meaning they justify edits only when appearing alongside stronger tells from Sections A or B.

### Section E: Chat Leftovers and Draft Artifacts

Section **E. Leftovers from the chat and the draft** lists residual chatbot artifacts that represent the “most certain tell” regarding AI generation but carry the least impact on prose quality. These include generic sign‑offs and draft‑mode markers that rarely alter meaning, making them the final priority in the editing workflow.

## How the Numeric Order Encodes Editing Priority

The pattern list’s numeric sequence from 1 to 25 directly encodes **relative strength and typical frequency**. Early numbers correspond to high‑confidence, high‑frequency AI tells, while later numbers indicate “weak alone” patterns that require contextual support. This numbering system allows automated tooling to sort matches by strength before presenting them to editors.

## Implementing Strength‑Based Processing

Consumer applications of the Humanizer skill should respect this hierarchy by sorting detected patterns according to their assigned strength tiers. The following Python implementation demonstrates how to load pattern definitions and prioritize matches from [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md):

```python
import re
from pathlib import Path

# Pattern definitions derived from SKILL.md sections

# Strength: 1=strongest (Section A), 5=weakest (Section E)

PATTERNS = [
    (1, r"\bnot\s+[^.]*\bbut\s+[^.]*", 1),          # Staging: Not X but Y

    (2, r"\b[A-Z][^.]*\.\s*$", 2),                  # One-line closers

    (3, r"\b(?:the\s+real\s+question|at\s+its\s+core)\b", 3),  # Deep sayings

    # ... patterns 4-24 ...

    (25, r"“[^”]*”", 5),                            # Curly quotes (Section D/E)

]

def find_patterns(text: str):
    """Return matches ordered by strength (strongest first)."""
    matches = []
    for pid, regex, strength in PATTERNS:
        for m in re.finditer(regex, text, flags=re.I):
            matches.append((strength, pid, m.span()))
    # Sort by strength ascending (1 before 5), then by position

    matches.sort()
    return matches

# Usage example

doc = Path("article.txt").read_text()
for strength, pid, span in find_patterns(doc):
    print(f"Pattern {pid} (strength tier {strength}) at {span}")

```

By sorting on the `strength` field before processing, the algorithm ensures that Section A violations receive immediate attention while Section D/E anomalies wait for corroborating evidence.

## Key Files Defining the Hierarchy

| File | Purpose |
|------|---------|
| [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) | Defines all 25 patterns, their numeric ordering, and the strength/frequency hierarchy governing edit decisions. |
| [`README.md`](https://github.com/blader/humanizer/blob/main/README.md) | Provides usage context and synchronizes with the pattern list structure. |
| [`scripts/validate-package.py`](https://github.com/blader/humanizer/blob/main/scripts/validate-package.py) | Validates that pattern numbers, section headings, and cross‑references remain consistent across repository updates. |

## Summary

- **blader/humanizer** ranks its 25 detection patterns in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) from strongest (1–5) to weakest (25).
- **Section A** patterns justify immediate edits on a single sighting, while **Sections B–E** require corroboration.
- The numeric order directly encodes editing priority, with lower numbers indicating higher frequency and reliability.
- Implementation code must sort pattern matches by strength tier to respect the documented hierarchy.
- **Section E** artifacts, though certain indicators of AI generation, have minimal impact on prose and lowest editing priority.

## Frequently Asked Questions

### What distinguishes a "strong" pattern from a "weak" one in Humanizer?

Strong patterns in **Section A** represent high‑frequency AI staging behaviors—such as forced “Not X but Y” constructions—that almost always indicate automated generation. Weak patterns, found in **Sections D and E**, include decorative formatting choices and residual chat artifacts that may appear in human writing and require multiple concurrent tells to justify editing.

### Why do weaker patterns require corroboration before triggering edits?

According to the [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) source, weaker patterns like forced triads or curly quotes are labeled “weak alone” because they frequently occur in legitimate human writing. Requiring “company from other tells” prevents over‑correction by ensuring that stylistic deviations appear in clusters characteristic of AI generation rather than isolated authorial quirks.

### How does the 1–25 numbering system correspond to editing workflow?

The **numeric order maps directly to action thresholds**: patterns 1–5 (Section A) trigger edits immediately upon detection, patterns 6–24 (Sections B–D) activate only when multiple tells appear in the same passage, and pattern 25+ (Section E) serves primarily as confirmation of AI origin without mandating stylistic changes.

### Can developers reorder the patterns in SKILL.md without breaking the logic?

Reordering the patterns would violate the **strength‑encoded numbering system** that consuming applications rely on for sorting. The [`scripts/validate-package.py`](https://github.com/blader/humanizer/blob/main/scripts/validate-package.py) file enforces this structure during CI, ensuring that pattern IDs consistently reflect their priority tiers and that automated tooling processes detection results correctly.