# How HoshinoBot Performs Text Normalization and Chinese Character Conversion

> Discover how HoshinoBot performs text normalization and Chinese character conversion using Unicode NFKC and zhconv. Learn how it prepares messages for efficient matching.

- Repository: [ice9coffee/hoshinobot](https://github.com/ice9coffee/hoshinobot)
- Tags: internals
- Published: 2026-03-03

---

**HoshinoBot normalizes incoming messages by applying Unicode NFKC normalization, lowercasing, and converting Traditional Chinese characters to Simplified Chinese using the zhconv library, storing the result in `event.norm_text` for downstream keyword and regex matching.**

The `ice9coffee/hoshinobot` repository implements a robust text processing pipeline that ensures consistent character representation across different input methods and regional variants. This system allows bot plugins to match keywords reliably regardless of whether users type in full-width or half-width characters, uppercase or lowercase, or Traditional versus Simplified Chinese scripts.

## The Two-Step Normalization Pipeline

The core normalization logic resides in [`hoshino/util/__init__.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/util/__init__.py) within the `normalize_str()` function (lines 89-96). This helper method processes all incoming text through a strict sequence of transformations before the bot attempts any pattern matching.

### Step 1: Unicode Normalization and Case Folding

First, the system standardizes Unicode compatibility characters and eliminates visual look-alikes. The implementation calls `unicodedata.normalize('NFKC', string)` to perform **Compatibility Decomposition** followed by **Canonical Composition**. This conversion ensures that full-width Latin characters, circled numbers, and compatibility ideographs are reduced to their standard equivalents. Immediately following this, the code applies `string.lower()` to enforce case-insensitive matching across all ASCII and Unicode alphabetic characters.

### Step 2: Simplified Chinese Conversion

After Unicode stabilization, the pipeline addresses Chinese script variants using the external `zhconv` library. The code invokes `zhconv.convert(string, 'zh-hans')` to transform any Traditional Chinese characters (zh-hant) or mixed-script input into Simplified Chinese (zh-hans). This step ensures that keywords written in Simplified Chinese will match user input regardless of whether the user types in Traditional characters, regional variants, or mixed forms.

```python

# hoshino/util/__init__.py (lines 89-96)

import unicodedata
import zhconv

def normalize_str(string) -> str:
    """
    规范化unicode字符串 并 转为小写 并 转为简体
    """
    # Unicode NFKC normalisation

    string = unicodedata.normalize('NFKC', string)
    # Lower‑case

    string = string.lower()
    # Simplified Chinese conversion (zh‑hans)

    string = zhconv.convert(string, 'zh-hans')
    return string

```

## Architecture and Data Flow

The normalization pipeline integrates into HoshinoBot's event processing chain through a specialized trigger class that enriches message events before they reach service handlers.

### The _TextNormalizer Class

Located in [`hoshino/trigger.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/trigger.py) (lines 57-62), the `_TextNormalizer` class inherits from `_PlainTextExtractor` and implements the `find_handler` method. This class sits early in the trigger processing chain and performs two critical operations: it first extracts the plain text content from the CQ event into `event.plain_text`, then populates `event.norm_text` by passing the extracted string through `util.normalize_str()`. The class returns an empty handler list, serving purely as a data enrichment layer that prepares the normalized representation for subsequent triggers.

```python

# hoshino/trigger.py (lines 57-62)

class _TextNormalizer(_PlainTextExtractor):
    def find_handler(self, event: CQEvent):
        # Extract plain text from the event

        super().find_handler(event)          # sets event.plain_text

        # Normalize and store in norm_text

        event.norm_text = util.normalize_str(event.plain_text)
        return []            # No handler produced here – just enriches the event

```

### Integration with Service Decorators

When plugin authors register keyword or regex triggers using `sv.on_keyword()` or `sv.on_rex()`, the `Service` class in [`hoshino/service.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/service.py) (lines 59-70) creates a `ServiceFunc` instance with `normalize=True` by default. This flag signals the corresponding `KeywordTrigger` or `RexTrigger` to reference `event.norm_text` rather than the raw message text when performing pattern matching. The normalized text ensures that regex patterns and keyword lists written against Simplified Chinese and standard ASCII will match user input regardless of the original script variant or character width.

### Prefix and Suffix Trigger Handling

For prefix-based and suffix-based triggers, HoshinoBot implements additional conversion logic to maintain bidirectional compatibility. In [`hoshino/trigger.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/trigger.py) (lines 31-42 and 72-84), the `PrefixTrigger` and `SuffixTrigger` classes convert their configured trigger strings using `zhconv.convert(..., "zh-hant")` when building their internal trie structures. This allows users to trigger commands using Traditional Chinese prefixes even when the bot's internal command registry stores the prefix in Simplified Chinese, ensuring consistent behavior across different input methods.

## Practical Usage for Plugin Developers

By default, all keyword and regex triggers in HoshinoBot operate on normalized text. Plugin developers can access the normalized content through `event.norm_text` when handling messages.

```python
from hoshino import Service

sv = Service('weather')

@sv.on_keyword('天气')  # normalize=True is the default

async def weather_handler(bot, ev):
    # ev.norm_text contains NFKC-normalized, lowercased, Simplified Chinese text

    if '北京' in ev.norm_text:
        await bot.send(ev, '北京天气查询结果...')

```

When exact character preservation is required, developers can opt out of normalization by passing `normalize=False` to the decorator. This configuration forces the trigger to match against `event.plain_text` instead, bypassing the Unicode normalization and Chinese conversion pipeline entirely.

```python
@sv.on_keyword('特殊字符', normalize=False)
async def exact_match_handler(bot, ev):
    # Access raw text without NFKC or zhconv processing

    raw_content = ev.plain_text
    await bot.send(ev, f'原始输入: {raw_content}')

```

## Summary

- **HoshinoBot implements a three-phase normalization pipeline** combining Unicode NFKC normalization, ASCII lowercasing, and Simplified Chinese conversion via the `zhconv` library.
- **Normalized text is stored in `event.norm_text`** by the `_TextNormalizer` class in [`hoshino/trigger.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/trigger.py) before any keyword or regex matching occurs.
- **Keyword and regex triggers use normalized text by default**, ensuring consistent matching across full-width/half-width variants, case differences, and Traditional/Simplified Chinese scripts.
- **Prefix and suffix triggers** implement additional Traditional Chinese conversion logic to support trigger strings written in either script variant.
- **Plugins can access raw text** through `event.plain_text` or disable normalization entirely by setting `normalize=False` in service decorators.

## Frequently Asked Questions

### What is the difference between `event.plain_text` and `event.norm_text`?

The `event.plain_text` attribute contains the raw text extracted from the message event without any modifications, preserving the original character width, case, and Chinese script variant. The `event.norm_text` attribute contains the processed version that has passed through `normalize_str()` in [`hoshino/util/__init__.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/util/__init__.py), meaning it has been NFKC-normalized, lowercased, and converted to Simplified Chinese. Standard keyword and regex triggers reference `event.norm_text` to ensure consistent matching behavior.

### Why does HoshinoBot convert Traditional Chinese to Simplified Chinese?

The conversion ensures maximum compatibility for bot plugins that typically register keywords and commands in Simplified Chinese. By normalizing all input to Simplified Chinese using `zhconv.convert(string, 'zh-hans')` in the `normalize_str()` function, the system allows users typing in Traditional Chinese (common in Taiwan and Hong Kong) or mixed scripts to trigger the same commands without requiring plugin authors to maintain duplicate keyword lists for both script variants.

### How can I disable text normalization for a specific keyword trigger?

Pass `normalize=False` as an argument to the `on_keyword()` or `on_rex()` decorator when registering your service function. According to the implementation in [`hoshino/service.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/service.py), this parameter prevents the trigger from accessing `event.norm_text` and forces it to use `event.plain_text` instead, preserving the exact Unicode characters, case sensitivity, and original Chinese script variant entered by the user.

### What Unicode normalization form does HoshinoBot use and why?

HoshinoBot uses **NFKC** (Normalization Form KC, or Compatibility Composition) as implemented in `unicodedata.normalize('NFKC', string)` within [`hoshino/util/__init__.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/util/__init__.py). This form was chosen because it decomposes compatibility characters (such as full-width ASCII, subscripts, and circled digits) into their canonical equivalents while recomposing combining characters where possible. This eliminates visual look-alikes that could otherwise bypass keyword filters or fail to match expected patterns.