How Humanizer's Internal Pattern Detection Algorithm Works: A Rule-Based Approach to De-AI-ing Text
Humanizer's internal pattern detection algorithm is a deterministic, rule-based pipeline that scans input text against 25 catalogued linguistic patterns ordered by detection strength, resolves overlaps using priority ranking, and rewrites content while strictly preserving factual accuracy.
The blader/humanizer repository implements a transparent text-processing engine designed to transform AI-generated prose into human-like writing. Unlike black-box neural detection systems, Humanizer's internal pattern detection algorithm relies entirely on explicit rules encoded directly in markdown files, rendering the logic fully auditable, reproducible, and free of probabilistic uncertainty.
Internal Pattern Detection Algorithm: The Six-Step Pipeline
The algorithm executes as a sequential pipeline without hidden model inference. Each step references concrete definitions stored in the repository's skill files.
1. Loading the Pattern Catalogue from SKILL.md
When invoked, the engine loads 25 distinct patterns defined in [SKILL.md](https://github.com/blader/humanizer/blob/main/SKILL.md) at the repository root. Each pattern entry specifies:
- A concrete linguistic cue (e.g., "not-X-but-Y", forced triads, dash overuse)
- A watch list of synonyms and regular-expression matching rules
- Strength ranking (sections §1–§5 represent the strongest signals acting on single occurrences)
- A "weak alone" flag indicating whether the pattern triggers only when clustered with other weak patterns
2. Marking Tells Through Line-by-Line Scanning
The detector scans input text line-by-line, applying regex-style watch lists for every loaded pattern. Strong patterns—such as Not X but Y—generate immediate tells upon single detection. Weak patterns—including Dashes as the universal connector and Stacked qualifiers—only register when multiple instances appear within the same passage, satisfying the "weak-alone" guard condition.
3. Priority Resolution Using Strength Ordering
Patterns process strictly by their catalogued priority number (lowest § number = highest strength). When multiple patterns match overlapping text spans, the algorithm retains the first (strongest) match and discards lower-priority detections covering identical regions. This deterministic resolution adheres to the editorial principle that "every sentence you keep must add something the reader did not already have."
4. Draft Rewrite Generation
After marking all tells, the system generates a draft rewrite that preserves every factual claim—names, dates, and citations—while excising or rephrasing detected patterns. Specific rewrite logic applies per pattern type: "not X but Y" constructions convert to direct statements, forced triads merge into natural prose, and excessive dashes substitute with commas or parentheses.
5. Validation Pass
The draft undergoes a second scan using the complete pattern catalogue to verify no residual tells remain and that the rewrite introduced no factual errors. Weak-alone tells survive only if they continue to cluster in the revised text; isolated weak patterns are purged in this phase.
6. Final Output Emission
The pipeline emits the cleaned prose alongside metadata listing any surviving weak-alone patterns and optional original-vs-final comparisons for transparency.
Implementation Example: Emulating the Detection Logic
While the production algorithm runs inside the agent runtime, you can replicate the detection behavior using standard Python libraries. The following example demonstrates loading pattern definitions, detecting tells with strength ordering, and applying conditional rewrites:
import re
from pathlib import Path
# Load pattern definitions (simplified)
PATTERNS = [
{
"id": 1,
"name": "Not X but Y",
"regex": r"\bnot\s+(?:just|only|merely)\s+([^,.;]+),\s+but\s+([^,.;]+)\b",
"replace": lambda m: f"{m.group(2).strip().capitalize()}",
"weak_alone": False,
},
{
"id": 8,
"name": "Dashes as universal connector",
"regex": r"\s+—\s+|--",
"replace": ", ",
"weak_alone": True,
},
# … add remaining patterns similarly …
]
def detect_tells(text):
"""Return a list of (pattern_id, span) for all matches, strongest first."""
tells = []
for pat in PATTERNS:
for m in re.finditer(pat["regex"], text, flags=re.IGNORECASE):
tells.append((pat["id"], m.span()))
# Sort by pattern strength (lower id = stronger)
tells.sort(key=lambda t: t[0])
return tells
def rewrite(text):
"""Apply pattern replacements in strength order."""
for pat in PATTERNS:
text = re.sub(pat["regex"], pat["replace"], text, flags=re.IGNORECASE)
return text
# Example usage
sample = Path("example.txt").read_text()
tells = detect_tells(sample)
draft = rewrite(sample)
print("Detected tells:", tells)
print("Draft rewrite:\n", draft)
Repository Structure and Key Files
Understanding the algorithm requires familiarity with these specific paths in blader/humanizer:
| File | Purpose |
|---|---|
SKILL.md |
Contains the complete pattern catalogue, detection rules, and rewrite guidance for all 25 linguistic patterns. |
README.md |
Documents the high-level algorithmic flow and links to pattern definitions. |
scripts/validate-package.py |
Utility that verifies internal consistency, ensuring pattern counts align with SKILL.md section headings. |
agents/openai.yaml |
OpenAI-compatible wrapper configuration that forwards user input to the skill engine. |
Summary
- The algorithm is deterministic and rule-based, operating without neural network inference.
- Detection relies on 25 patterns ranked by strength in
SKILL.md, with §1–§5 acting as strong signals and §6+ requiring clustering. - Strength-ordered resolution handles overlapping matches by prioritizing the lowest section number.
- A two-pass validation system ensures factual preservation while eliminating AI-specific linguistic tells.
- The "weak-alone" guard prevents over-editing by requiring weak patterns to cluster before triggering rewrites.
Frequently Asked Questions
Is Humanizer's pattern detection algorithm based on machine learning?
No. According to the blader/humanizer source code, the algorithm is explicitly deterministic and rule-based. It matches text against regular-expression patterns defined in SKILL.md without invoking hidden neural models, embeddings, or probabilistic classifiers. This design choice ensures fully auditable and reproducible behavior.
How does the "weak alone" guard function in pattern detection?
The "weak alone" flag prevents weaker stylistic patterns—such as dash overuse or stacked qualifiers—from triggering edits unless they appear in clusters within the same passage. If a weak pattern occurs only once, the algorithm ignores it; if multiple weak patterns congregate, indicating systematic AI prose structure, the guard allows marking and subsequent rewriting.
What happens when multiple patterns detect the same text span?
The algorithm resolves conflicts through strict strength ordering. It processes patterns sequentially by their catalogued priority (§1 strongest to §25 weakest), retaining the first match and discarding any subsequent lower-priority patterns that overlap the same text span. This guarantees deterministic output for identical inputs.
Where are the pattern definitions stored in the repository?
All 25 detection patterns, including their watch lists, regex-style matching rules, strength rankings, and rewrite guidance, reside in [SKILL.md](https://github.com/blader/humanizer/blob/main/SKILL.md) at the repository root. The [scripts/validate-package.py](https://github.com/blader/humanizer/blob/main/scripts/validate-package.py) utility continuously verifies this file's internal consistency against the expected pattern count.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →