How Humanizer Matches Writing Style When a Sample Is Provided
Humanizer determines a writer’s personal style by first reading a supplied writing sample and then mirroring its characteristics during the rewrite process, treating the sample as ground truth that overrides default style rules.
The blader/humanizer project enables users to humanize AI-generated text while preserving their authentic voice. When you provide a writing sample, the system analyzes specific linguistic markers and applies them throughout the transformation process, ensuring the output reflects your actual writing habits rather than generic humanization patterns.
Detecting the Writing Sample
Humanizer checks for the presence of a writing sample at the start of its execution flow. According to the skill definition in SKILL.md (lines 30-38), when a user includes a sample of their previous writing, the engine immediately jumps to the "Match the writer's voice" section. This detection triggers a specialized workflow that prioritizes sample analysis over standard style templates.
The sample serves as a reference document that instructs the engine on how to reshape the AI-generated content. Without a sample, Humanizer relies on generic anti-AI pattern rules; with a sample, those rules become conditional based on the user's demonstrated preferences.
Analyzing Stylistic Signals
Before any rewriting begins, Humanizer extracts a comprehensive set of stylistic signals from the provided sample. The analysis focuses on measurable characteristics that define personal voice:
- Sentence length and rhythm – Identifying patterns in short-long sentence variation and cadence.
- Word choice – Cataloging favored adjectives, adverbs, and specific vocabulary terms.
- Paragraph structure – Noting opening styles and transition preferences.
- Punctuation density – Measuring use of em dashes, en dashes, commas, colons, and semicolons.
- Idiosyncratic patterns – Recording repeated phrases, deliberate quirks, and unique stylistic fingerprints.
This upfront analysis ensures the engine understands the writer's habits before applying any transformations to the target text.
Prioritizing Sample Over Generic Rules
The skill explicitly establishes hierarchy in SKILL.md (lines 38-39): "A writing sample takes priority over these style rules." This means any default prohibitions or style guidelines are overridden if the sample demonstrates those characteristics.
For example, section 14 of the skill definition normally bans em dashes (—) and en dashes (–). However, as noted in SKILL.md (lines 81-84), this rule is conditionally relaxed when the sample contains these punctuation marks. If the analysis flags uses_em_dash = True based on the sample, the engine preserves dash usage at the same frequency rather than replacing them with commas or periods.
Applying Sample-Driven Transformations
During the rewrite phase, Humanizer applies the extracted metrics to shape the output:
- Rhythm matching – The engine restructures sentences to mirror the sample's cadence, maintaining the same patterns of short and long constructions.
- Vocabulary retention – Specific word choices and phrasing patterns identified in the sample are incorporated into the rewritten text.
- Punctuation mirroring – The density and types of punctuation match the sample's ratios, including conditional allowance of otherwise-prohibited characters.
- Quirk preservation – Intentional repetitions or unique stylistic quirks remain untouched, as mandated by the sample-priority clause.
Validating Against the Source Material
After producing the draft, Humanizer performs a final validation check to ensure compliance with the sample's characteristics. As specified in SKILL.md (lines 93-94), the engine verifies that prohibited characters like em dashes only appear in the output if the original sample used them. This validation step ensures the rewrite authentically represents the user's voice while maintaining the anti-AI pattern requirements that don't conflict with the sample.
Implementation Example
Below is a conceptual representation of how to invoke Humanizer with a writing sample:
# Pseudo-code: feeding Humanizer a sample and AI-generated text
sample = """
I love wandering through the city's narrow alleys, coffee in hand,
watching the sunrise paint the rooftops orange. It's messy, but it feels right.
"""
ai_text = """
The AI-generated paragraph describes a city tour with many attractions and a pleasant experience.
"""
# Humanizer call (the skill parses the sample first)
result = humanizer.run(
sample=sample,
text=ai_text,
)
print(result.rewrite) # Output matches the sample's cadence and dash usage
The internal logic processes the sample before rewriting:
1. If `sample` is present:
- Extract metrics (sentence length, punctuation, lexical tokens).
- Store flags (e.g., `uses_em_dash = True`).
2. When rewriting:
- Apply the 35 AI-pattern removals.
- For each style rule (e.g., dash ban §14), check the sample flag:
- If `uses_em_dash` → allow em dashes at the same rate.
- Otherwise → replace dashes with commas/periods.
- Preserve repeated phrases or quirky wording found in the sample.
3. Return the final text with a pattern summary.
Key Source Files
The style-matching functionality is defined across three primary files in the blader/humanizer repository:
SKILL.md– Contains the procedural definition, including sample detection (lines 30-38), the priority clause (lines 38-39), conditional dash rules (lines 81-84), and validation logic (lines 93-94).README.md– Provides user-facing documentation on supplying samples for voice matching.AGENTS.md– Describes deployment packaging for various AI agents.
Summary
- Humanizer detects writing samples in
SKILL.md(lines 30-38) and redirects to voice-matching workflows. - The engine extracts rhythm, vocabulary, punctuation, and quirks from samples before rewriting.
- Samples take priority over default style rules, overriding prohibitions like the dash ban in §14.
- Validation ensures prohibited characters only appear if the sample uses them (lines 93-94).
- The logic is implemented entirely within the markdown skill definition, requiring no external configuration files.
Frequently Asked Questions
How does Humanizer detect when a writing sample is provided?
According to SKILL.md (lines 30-38), the skill checks for the presence of a sample at the start of execution. When detected, the engine immediately jumps to the "Match the writer's voice" section, triggering the stylistic analysis workflow before any rewriting occurs.
What specific elements does Humanizer analyze in a writing sample?
The system analyzes sentence length and rhythm, word choice including favored adjectives and adverbs, paragraph opening patterns, punctuation density (particularly em/en dashes and commas), and any repeated phrases or deliberate quirks that define the writer's unique voice.
Can a writing sample override default prohibitions like dash usage?
Yes. The skill explicitly states that "A writing sample takes priority over these style rules" (lines 38-39). For example, section 14 normally bans em dashes, but this prohibition is conditionally relaxed if the sample contains them (lines 81-84).
Where is the style-matching logic defined in the codebase?
The entire workflow is defined in SKILL.md within the blader/humanizer repository. This file contains the detection logic (lines 30-38), priority hierarchy (lines 38-39), conditional rule implementation (lines 81-84), and validation checks (lines 93-94).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →