# How the Character Name Alias Conversion System Works in HoshinoBot

> Uncover the HoshinoBot character name alias conversion system. Learn how its trie-based Roster and fuzzy matching accurately identify characters like Princess Connect! Re:Dive for seamless interaction.

- Repository: [ice9coffee/hoshinobot](https://github.com/ice9coffee/hoshinobot)
- Tags: internals
- Published: 2026-03-03

---

**The character name alias conversion system in HoshinoBot uses a trie-based `Roster` class combined with fuzzy string matching to map user-input names, nicknames, and misspellings to unique character IDs for Princess Connect! Re:Dive.**

The [`chara.py`](https://github.com/ice9coffee/hoshinobot/blob/main/chara.py) module in the [ice9coffee/hoshinobot](https://github.com/ice9coffee/hoshinobot) repository powers character identification across the bot's Princess Connect functionality. According to the source code in [`hoshino/modules/priconne/chara.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/modules/priconne/chara.py), the system handles hundreds of multilingual aliases by normalizing input and performing fast prefix-tree lookups.

## Core Architecture of the Roster Class

The alias system centers on the **`Roster`** class, which maintains a `pygtrie.CharTrie` data structure to enable **longest-prefix matching**. This design allows the bot to efficiently parse concatenated team strings and resolve individual character names in a single pass.

### Trie Initialization and Data Loading

When the module imports, it instantiates a global `roster = Roster()` object. The constructor automatically invokes **`update()`** (lines 34-47), which performs the following steps:

1. Reloads the `_pcr_data` module containing the **`CHARA_NAME`** dictionary
2. Iterates over every character ID and its associated alias list
3. Normalizes each alias using `util.normalize_str` (removes whitespace, converts to lowercase, strips special characters)
4. Inserts normalized aliases into the trie with their corresponding character IDs as values
5. Logs warnings for any duplicate aliases encountered during population

This initialization happens at lines 30-33 in [`hoshino/modules/priconne/chara.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/modules/priconne/chara.py), ensuring the trie remains synchronized with the underlying data source.

## Lookup Mechanisms and Conversion Methods

The system provides three distinct search strategies to accommodate exact matches, typos, and complex team compositions.

### Exact Matching with get_id()

The **`get_id(name)`** method (lines 49-52) performs direct trie lookups for normalized input:

- Returns the exact character ID if the normalized alias exists
- Returns **`1000`** (the `UNKNOWN` constant) if no match is found
- Example: `name2id('日和')`, `name2id('hiyori')`, and `name2id('猫拳')` all return `1001`

### Fuzzy Matching for Misspellings

When exact lookup fails, the **`guess_id(name)`** method (lines 53-57) employs `fuzzywuzzy.process.extractOne` to find the closest matching alias:

- Returns a tuple containing `(character_id, matched_alias, similarity_score)`
- Scores range from 0-100 based on Levenshtein distance
- Useful for handling typos like "hiy0ri" or ambiguous inputs like "日和莉"

### Parsing Concatenated Team Strings

The **`parse_team(namestr)`** method (lines 58-71) handles bulk character identification from run-on strings such as "日和优衣怜":

- Repeatedly calls `self._roster.longest_prefix(namestr)` to extract the longest matching alias from the remaining string
- Appends the corresponding ID to the result list and removes the matched prefix
- Collects unrecognized characters into a separate string for error reporting
- Example: `roster.parse_team('日和优衣怜')` returns `([1001, 1002, 1003], '')`

## Helper Functions and Public API

The module exposes convenient wrappers around the `Roster` methods:

- **`name2id(name)`** (lines 77-79): Direct wrapper for `get_id()`
- **`fromname(name)`** (lines 85-88): Returns a `Chara` object instance from a name string
- **`is_npc(id)`** (lines 95-99): Checks if the resolved ID corresponds to a non-player character

## Practical Code Examples

```python

# Exact conversion using official and variant names

>>> from hoshino.modules.priconne.chara import name2id
>>> name2id('日和')        # Chinese official name

1001
>>> name2id('猫拳')        # Nickname variant

1001

# Handling typos through fuzzy matching

>>> from hoshino.modules.priconne.chara import guess_id
>>> guess_id('日和莉')
(1001, '日和', 90)

# Parsing a concatenated team composition

>>> from hoshino.modules.priconne.chara import roster
>>> roster.parse_team('日和优衣怜')
([1001, 1002, 1003], '')

# Identifying unknown characters

>>> roster.parse_team('日和X优衣')
([1001, 1002], 'X')

```

## Summary

- The **`Roster`** class in [`hoshino/modules/priconne/chara.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/modules/priconne/chara.py) uses a **`pygtrie.CharTrie`** to store normalized character aliases mapped to integer IDs.
- **`update()`** rebuilds the trie from `_pcr_data.CHARA_NAME`, normalizing all aliases via `util.normalize_str`.
- **`get_id()`** provides exact O(1) lookups, while **`guess_id()`** uses `fuzzywuzzy` for typo-tolerant matching.
- **`parse_team()`** enables extraction of multiple character IDs from concatenated strings using longest-prefix matching.
- Adding new aliases requires only updating the `CHARA_NAME` dictionary in [`_pcr_data.py`](https://github.com/ice9coffee/hoshinobot/blob/main/_pcr_data.py); the trie rebuilds automatically on the next module reload.

## Frequently Asked Questions

### How does HoshinoBot handle misspelled character names?

When exact lookup fails, the system falls back to **`guess_id()`**, which uses `fuzzywuzzy.process.extractOne` to calculate similarity scores against all stored aliases. This catches typos like "hiy0ri" (normalized to "hiyori") and returns the best match with a confidence score ranging from 0 to 100.

### What data structure enables fast alias lookups in chara.py?

The **`Roster`** class maintains a **`pygtrie.CharTrie`** (character trie) that stores normalized alias strings as keys and character IDs as values. This prefix-tree structure supports O(m) longest-prefix lookups where m is the length of the input string, making it efficient for parsing both individual names and concatenated team compositions.

### How does the system parse multiple characters from a single string?

The **`parse_team(namestr)`** method iteratively applies `longest_prefix()` to the input string, extracting the longest matching alias at each step. It removes matched prefixes from the remaining string and continues until the string is empty or only unknown characters remain, returning a list of character IDs and any unrecognized text.

### Where are the character aliases defined in the repository?

All aliases reside in **[`hoshino/modules/priconne/_pcr_data.py`](https://github.com/ice9coffee/hoshinobot/blob/main/hoshino/modules/priconne/_pcr_data.py)** within the `CHARA_NAME` dictionary. Each entry maps a character ID to a list of strings containing official names, Japanese names, English transliterations, nicknames, and common typo variants. The `Roster.update()` method reads this dictionary to populate the trie at runtime.