How HoshinoBot Performs Text Normalization and Chinese Character Conversion
HoshinoBot normalizes incoming messages by applying Unicode NFKC normalization, lowercasing, and converting Traditional Chinese characters to Simplified Chinese using the zhconv library, storing the result in event.norm_text for downstream keyword and regex matching.
The ice9coffee/hoshinobot repository implements a robust text processing pipeline that ensures consistent character representation across different input methods and regional variants. This system allows bot plugins to match keywords reliably regardless of whether users type in full-width or half-width characters, uppercase or lowercase, or Traditional versus Simplified Chinese scripts.
The Two-Step Normalization Pipeline
The core normalization logic resides in hoshino/util/__init__.py within the normalize_str() function (lines 89-96). This helper method processes all incoming text through a strict sequence of transformations before the bot attempts any pattern matching.
Step 1: Unicode Normalization and Case Folding
First, the system standardizes Unicode compatibility characters and eliminates visual look-alikes. The implementation calls unicodedata.normalize('NFKC', string) to perform Compatibility Decomposition followed by Canonical Composition. This conversion ensures that full-width Latin characters, circled numbers, and compatibility ideographs are reduced to their standard equivalents. Immediately following this, the code applies string.lower() to enforce case-insensitive matching across all ASCII and Unicode alphabetic characters.
Step 2: Simplified Chinese Conversion
After Unicode stabilization, the pipeline addresses Chinese script variants using the external zhconv library. The code invokes zhconv.convert(string, 'zh-hans') to transform any Traditional Chinese characters (zh-hant) or mixed-script input into Simplified Chinese (zh-hans). This step ensures that keywords written in Simplified Chinese will match user input regardless of whether the user types in Traditional characters, regional variants, or mixed forms.
# hoshino/util/__init__.py (lines 89-96)
import unicodedata
import zhconv
def normalize_str(string) -> str:
"""
规范化unicode字符串 并 转为小写 并 转为简体
"""
# Unicode NFKC normalisation
string = unicodedata.normalize('NFKC', string)
# Lower‑case
string = string.lower()
# Simplified Chinese conversion (zh‑hans)
string = zhconv.convert(string, 'zh-hans')
return string
Architecture and Data Flow
The normalization pipeline integrates into HoshinoBot's event processing chain through a specialized trigger class that enriches message events before they reach service handlers.
The _TextNormalizer Class
Located in hoshino/trigger.py (lines 57-62), the _TextNormalizer class inherits from _PlainTextExtractor and implements the find_handler method. This class sits early in the trigger processing chain and performs two critical operations: it first extracts the plain text content from the CQ event into event.plain_text, then populates event.norm_text by passing the extracted string through util.normalize_str(). The class returns an empty handler list, serving purely as a data enrichment layer that prepares the normalized representation for subsequent triggers.
# hoshino/trigger.py (lines 57-62)
class _TextNormalizer(_PlainTextExtractor):
def find_handler(self, event: CQEvent):
# Extract plain text from the event
super().find_handler(event) # sets event.plain_text
# Normalize and store in norm_text
event.norm_text = util.normalize_str(event.plain_text)
return [] # No handler produced here – just enriches the event
Integration with Service Decorators
When plugin authors register keyword or regex triggers using sv.on_keyword() or sv.on_rex(), the Service class in hoshino/service.py (lines 59-70) creates a ServiceFunc instance with normalize=True by default. This flag signals the corresponding KeywordTrigger or RexTrigger to reference event.norm_text rather than the raw message text when performing pattern matching. The normalized text ensures that regex patterns and keyword lists written against Simplified Chinese and standard ASCII will match user input regardless of the original script variant or character width.
Prefix and Suffix Trigger Handling
For prefix-based and suffix-based triggers, HoshinoBot implements additional conversion logic to maintain bidirectional compatibility. In hoshino/trigger.py (lines 31-42 and 72-84), the PrefixTrigger and SuffixTrigger classes convert their configured trigger strings using zhconv.convert(..., "zh-hant") when building their internal trie structures. This allows users to trigger commands using Traditional Chinese prefixes even when the bot's internal command registry stores the prefix in Simplified Chinese, ensuring consistent behavior across different input methods.
Practical Usage for Plugin Developers
By default, all keyword and regex triggers in HoshinoBot operate on normalized text. Plugin developers can access the normalized content through event.norm_text when handling messages.
from hoshino import Service
sv = Service('weather')
@sv.on_keyword('天气') # normalize=True is the default
async def weather_handler(bot, ev):
# ev.norm_text contains NFKC-normalized, lowercased, Simplified Chinese text
if '北京' in ev.norm_text:
await bot.send(ev, '北京天气查询结果...')
When exact character preservation is required, developers can opt out of normalization by passing normalize=False to the decorator. This configuration forces the trigger to match against event.plain_text instead, bypassing the Unicode normalization and Chinese conversion pipeline entirely.
@sv.on_keyword('特殊字符', normalize=False)
async def exact_match_handler(bot, ev):
# Access raw text without NFKC or zhconv processing
raw_content = ev.plain_text
await bot.send(ev, f'原始输入: {raw_content}')
Summary
- HoshinoBot implements a three-phase normalization pipeline combining Unicode NFKC normalization, ASCII lowercasing, and Simplified Chinese conversion via the
zhconvlibrary. - Normalized text is stored in
event.norm_textby the_TextNormalizerclass inhoshino/trigger.pybefore any keyword or regex matching occurs. - Keyword and regex triggers use normalized text by default, ensuring consistent matching across full-width/half-width variants, case differences, and Traditional/Simplified Chinese scripts.
- Prefix and suffix triggers implement additional Traditional Chinese conversion logic to support trigger strings written in either script variant.
- Plugins can access raw text through
event.plain_textor disable normalization entirely by settingnormalize=Falsein service decorators.
Frequently Asked Questions
What is the difference between event.plain_text and event.norm_text?
The event.plain_text attribute contains the raw text extracted from the message event without any modifications, preserving the original character width, case, and Chinese script variant. The event.norm_text attribute contains the processed version that has passed through normalize_str() in hoshino/util/__init__.py, meaning it has been NFKC-normalized, lowercased, and converted to Simplified Chinese. Standard keyword and regex triggers reference event.norm_text to ensure consistent matching behavior.
Why does HoshinoBot convert Traditional Chinese to Simplified Chinese?
The conversion ensures maximum compatibility for bot plugins that typically register keywords and commands in Simplified Chinese. By normalizing all input to Simplified Chinese using zhconv.convert(string, 'zh-hans') in the normalize_str() function, the system allows users typing in Traditional Chinese (common in Taiwan and Hong Kong) or mixed scripts to trigger the same commands without requiring plugin authors to maintain duplicate keyword lists for both script variants.
How can I disable text normalization for a specific keyword trigger?
Pass normalize=False as an argument to the on_keyword() or on_rex() decorator when registering your service function. According to the implementation in hoshino/service.py, this parameter prevents the trigger from accessing event.norm_text and forces it to use event.plain_text instead, preserving the exact Unicode characters, case sensitivity, and original Chinese script variant entered by the user.
What Unicode normalization form does HoshinoBot use and why?
HoshinoBot uses NFKC (Normalization Form KC, or Compatibility Composition) as implemented in unicodedata.normalize('NFKC', string) within hoshino/util/__init__.py. This form was chosen because it decomposes compatibility characters (such as full-width ASCII, subscripts, and circled digits) into their canonical equivalents while recomposing combining characters where possible. This eliminates visual look-alikes that could otherwise bypass keyword filters or fail to match expected patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →