How Humanizer Uses Voice Matching When You Provide a Writing Sample

Humanizer analyzes stylistic elements from your writing sample—including sentence length, punctuation habits, and word choice—to establish a voice baseline that overrides generic editing rules, ensuring the rewrite mirrors your unique authorial style.

The blader/humanizer repository provides an AI-powered text transformation tool that goes beyond standard grammar correction. When you supply a writing sample, the system performs voice matching, a process that captures your specific linguistic fingerprint to maintain consistency across generated content.

How Voice Matching Works in Humanizer

Voice matching operates by examining the structural and stylistic DNA of your provided text. Unlike generic AI rewrites that apply universal rules, Humanizer treats your sample as the authoritative standard for the output.

Analyzing Stylistic Fingerprints

When you provide a writing sample, Humanizer extracts several key characteristics to build a voice profile:

  • Sentence length and rhythm – Short, punchy constructions versus elaborate, flowing syntax
  • Word choice preferences – Specific vocabulary patterns and formality levels
  • Punctuation patterns – Distinctive use of em-dashes, semicolons, or comma placement
  • Opening and transition styles – How you typically begin sentences and connect ideas

These extracted features become the constraints for the rewrite generation process.

Overriding Default Pattern Rules

Normally, Humanizer applies standard pattern-based rules defined in the skill configuration. However, as documented in SKILL.md at line 40, when a writing sample is present, voice matching overrides these generic rules.

For example, while the default configuration in section 8 might standardize dash usage across all outputs, your sample's specific dash styling takes precedence. The system prioritizes preserving your original punctuation habits over applying generic editorial standards.

Providing a Writing Sample to Humanizer

You can invoke voice matching by including your reference text when calling the Humanizer skill. The examples below demonstrate how different samples produce distinct outputs.

Preserving Punctuation Style

When you want to maintain specific punctuation habits like em-dash spacing:

User sample:
"The new feature — announced yesterday — improves performance."

Humanizer's rewrite (voice-matched):
"The new feature — announced yesterday — improves performance."

The system recognizes the spaced em-dash pattern from your sample and preserves it, rather than converting to standard closed em-dashes or commas.

Matching Sentence Rhythm

For content with distinctive rhythmic patterns:

User sample:
"Quick insights, bold moves, lasting impact."

Humanizer's rewrite (voice-matched):
"Quick insights, bold moves, lasting impact."

Humanizer maintains the parallel structure and staccato cadence, ensuring the rewrite feels like it came from the same author.

Fallback to Genre-Based Voice

If you do not provide a sample, Humanizer defaults to genre-based voice selection as defined in the Voice section of SKILL.md:

User request (no sample):
"Explain how caching works in Next.js."

Humanizer's rewrite (genre-based):
"Next.js caches data at multiple layers, including request memoization, the data cache, and the router cache."

Without sample-based voice matching, the system selects a technical blog voice from predefined genres (blog, technical, legal, etc.).

Technical Implementation

The voice matching functionality spans several key files in the repository.

Skill Definition

The core logic resides in SKILL.md, which defines the Humanizer skill's architecture. The Voice section (around line 40) explicitly describes how writing samples trigger voice matching mode and how this mode takes precedence over default pattern handlers.

Usage Integration

The README.md documents how end users invoke Humanizer with writing samples, explaining the interface for passing reference text to the skill during execution.

Agent Configuration

In agents/openai.yaml, the packaging configuration for OpenAI-compatible agents shows how the writing sample parameter is passed through the agent framework to the underlying Humanizer processing engine.

Summary

  • Voice matching analyzes your writing sample's stylistic DNA—sentence structure, punctuation, and vocabulary—to create a rewrite baseline.
  • When active, voice matching overrides default pattern rules (such as standard dash handling) to preserve your unique style.
  • Without a sample, Humanizer falls back to genre selection (blog, technical, legal) from predefined voice categories.
  • The functionality is configured in SKILL.md, with usage documented in README.md and agent integration handled in agents/openai.yaml.

Frequently Asked Questions

How does Humanizer analyze my writing sample for voice matching?

Humanizer examines specific stylistic markers including sentence length distribution, word choice patterns, punctuation habits (particularly em-dash and comma usage), and transition structures. According to the SKILL.md source code, these characteristics form a baseline profile that constrains the rewrite generation, ensuring the output mirrors your original voice rather than applying generic editorial standards.

What happens if I don't provide a writing sample?

If no sample is supplied, Humanizer defaults to genre-based voice selection. As implemented in the blader/humanizer repository, the system selects from predefined categories such as blog, technical, or legal writing styles. This fallback mechanism applies standard pattern-based rules rather than personalized voice matching.

Can voice matching override specific grammar rules?

Yes. The voice matching process explicitly overrides generic pattern-based rules when they conflict with your sample's style. For instance, if your sample uses spaced em-dashes (like "word — word") while the default rules specify closed em-dashes, the sample's styling takes precedence. This override behavior is documented in the Voice section of SKILL.md.

Where is the voice matching logic configured in the repository?

The primary configuration resides in SKILL.md, specifically in the Voice section around line 40, which defines how writing samples trigger voice matching mode. The README.md provides user-facing documentation for supplying samples, while agents/openai.yaml handles the technical interface for passing samples through OpenAI-compatible agent frameworks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →