# What Are the "Leftovers from the Chat" Patterns (22-25) That Humanizer Removes?

> Discover how Humanizer's Leftovers from the Chat patterns (22-25) clean copied text by removing speaker labels, code fences, timestamps, and excess whitespace for better readability.

- Repository: [Siqi Chen/humanizer](https://github.com/blader/humanizer)
- Tags: deep-dive
- Published: 2026-09-09

---

**Humanizer's "Leftovers from the Chat" patterns (22-25) strip speaker labels, markdown block quotes, stray code fences, horizontal rules, timestamps, and redundant whitespace from text copied out of conversational UIs.**

The open-source **blader/humanizer** repository defines a specialized cleanup system in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) that targets conversational debris surviving copy-and-paste operations. These regex-based patterns, cataloged under section 22 (specifically sub-pattern 22-2S), normalize chat exports into clean, speaker-agnostic text by removing eight distinct categories of artifacts.

## Speaker Label Removal (User:, Assistant:)

The first patterns in the 22-25 sequence target explicit speaker identifiers that prefix dialogue turns in chat logs.

- **`User:`** and **`You:`** — Deletes explicit labels marking human input lines that appear at the start of user turns.
- **`Assistant:`**, **`AI:`**, and **`ChatGPT:`** — Removes assistant-speaker labels that prefix the model’s replies.

According to the source code analysis, these tokens are identified by regular expressions in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) section 22 and are nullified before further processing begins.

## Markdown Artifact Cleanup

### Block Quote Markers

Pattern 22-2S handles Markdown block-quote remnants that persist in exported chat threads.

- **Leading `>`** — Strips single-level block-quote markers common in threaded views.
- **Leading `>>>`** — Removes deeper nesting markers that appear in multi-level conversation exports.

### Stray Code Fence Delimiters

The system identifies isolated triple backticks (`` ``` ``) that are not part of actual code blocks. These stray delimiters are cleaned to prevent formatting errors when the text is rendered in other Markdown processors.

## Structural Formatting Normalization

### Horizontal Rule Separators

Lines containing only **`---`** or **`***`** that function as turn separators in chat logs are deleted entirely. These horizontal rules serve no semantic purpose outside the conversational UI context.

### Whitespace Collapse

Redundant line-breaks are collapsed to single breaks, and trailing whitespace is trimmed from all lines. This prevents the "staircase" formatting often caused by copying nested chat replies.

## Metadata and Timestamp Stripping

The final patterns target non-content metadata lines such as **`[10:23 PM]`** or similar timestamp formats. These regex patterns match generic time annotations that are not part of the substantive text, ensuring the output contains only the actual dialogue content.

## Implementation in SKILL.md

As implemented in the `blader/humanizer` repository, these rules live in `SKILL.md` under the "Leftovers from the Chat" designation. The following Python example demonstrates the conceptual regex patterns derived from section 22, sub-pattern 22-2S:

```python
import re

# Patterns corresponding to Humanizer's 22-25 cleanup rules

chat_cleanup_patterns = [
    r'^User:\s*',                    # Pattern 22 variant: User labels

    r'^(?:Assistant|AI|ChatGPT):\s*', # Pattern 23 variant: Assistant labels

    r'^>{1,3}\s?',                   # Pattern 24: Block quote markers

    r'^\s*```\s*$',                  # Stray code fences

    r'^(?:---+|\*\*\*+)\s*$',        # Pattern 25 variant: Horizontal rules

    r'\[\d{1,2}:\d{2}\s*[AP]M\]',    # Timestamp metadata

    r'[ \t]+$',                      # Trailing whitespace

]

def clean_chat_leftovers(text):
    lines = text.split('\n')
    cleaned = []
    for line in lines:
        # Skip horizontal rules and timestamps

        if any(re.match(p, line) for p in chat_cleanup_patterns[3:6]):
            continue
        # Remove speaker labels and block quotes

        for pattern in chat_cleanup_patterns[:4]:
            line = re.sub(pattern, '', line)
        # Trim trailing whitespace

        line = re.sub(chat_cleanup_patterns[6], '', line)
        if line or (cleaned and cleaned[-1] != ''):
            cleaned.append(line)
    return '\n'.join(cleaned)

```

## Before and After Example

The following demonstrates how the Leftovers from the Chat patterns transform raw copied content:

```markdown

# Input (raw chat copy):

User: Explain recursion.
> Assistant: Here is an example:
> ```

> def factorial(n):
>     return n * factorial(n-1)
> ```

> 
> ---
> [2:45 PM]
> Does that help?

# Output (after Humanizer):

Explain recursion.
Here is an example:

```

def factorial(n):
    return n * factorial(n-1)

```

Does that help?

```

## Summary

- **Leftovers from the Chat** patterns (22-25) are defined in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) section 22, specifically sub-pattern 22-2S.
- **Speaker labels** including `User:`, `You:`, `Assistant:`, `AI:`, and `ChatGPT:` are completely removed.
- **Markdown artifacts** such as leading `>`, `>>>`, and stray code fences (`` ``` ``) are stripped.
- **Structural separators** like `---` and `***` are deleted to eliminate turn boundaries.
- **Metadata lines** including timestamps formatted as `[10:23 PM]` are filtered out.
- **Whitespace normalization** collapses redundant line breaks and trims trailing spaces.

## Frequently Asked Questions

### Where are the "Leftovers from the Chat" patterns defined in the Humanizer repository?

According to the source analysis, these patterns are defined in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md) under section 22, specifically within sub-pattern 22-2S. This section contains the regular expressions that identify and remove chat-specific artifacts.

### Does Humanizer remove all Markdown formatting or only chat-specific artifacts?

Humanizer specifically targets **chat-specific artifacts** such as speaker labels, nested block quotes from threaded views, and separator lines. It preserves legitimate Markdown formatting like intentional code blocks and headers that are part of the actual content.

### What does the "S" in pattern 22-2S stand for?

While the repository documentation does not explicitly define the "S" notation, pattern 22-2S refers to the comprehensive cleanup sub-pattern within section 22 that handles the full set of speaker labels and formatting artifacts listed in the "Leftovers from the Chat" category.

### Are these patterns applied automatically when using Humanizer?

Yes, as part of the core skill configuration in [`SKILL.md`](https://github.com/blader/humanizer/blob/main/SKILL.md), the "Leftovers from the Chat" patterns (22-25) are applied during the text normalization phase to ensure clean output from conversational sources.