How the Character Name Alias Conversion System Works in HoshinoBot
The character name alias conversion system in HoshinoBot uses a trie-based Roster class combined with fuzzy string matching to map user-input names, nicknames, and misspellings to unique character IDs for Princess Connect! Re:Dive.
The chara.py module in the ice9coffee/hoshinobot repository powers character identification across the bot's Princess Connect functionality. According to the source code in hoshino/modules/priconne/chara.py, the system handles hundreds of multilingual aliases by normalizing input and performing fast prefix-tree lookups.
Core Architecture of the Roster Class
The alias system centers on the Roster class, which maintains a pygtrie.CharTrie data structure to enable longest-prefix matching. This design allows the bot to efficiently parse concatenated team strings and resolve individual character names in a single pass.
Trie Initialization and Data Loading
When the module imports, it instantiates a global roster = Roster() object. The constructor automatically invokes update() (lines 34-47), which performs the following steps:
- Reloads the
_pcr_datamodule containing theCHARA_NAMEdictionary - Iterates over every character ID and its associated alias list
- Normalizes each alias using
util.normalize_str(removes whitespace, converts to lowercase, strips special characters) - Inserts normalized aliases into the trie with their corresponding character IDs as values
- Logs warnings for any duplicate aliases encountered during population
This initialization happens at lines 30-33 in hoshino/modules/priconne/chara.py, ensuring the trie remains synchronized with the underlying data source.
Lookup Mechanisms and Conversion Methods
The system provides three distinct search strategies to accommodate exact matches, typos, and complex team compositions.
Exact Matching with get_id()
The get_id(name) method (lines 49-52) performs direct trie lookups for normalized input:
- Returns the exact character ID if the normalized alias exists
- Returns
1000(theUNKNOWNconstant) if no match is found - Example:
name2id('日和'),name2id('hiyori'), andname2id('猫拳')all return1001
Fuzzy Matching for Misspellings
When exact lookup fails, the guess_id(name) method (lines 53-57) employs fuzzywuzzy.process.extractOne to find the closest matching alias:
- Returns a tuple containing
(character_id, matched_alias, similarity_score) - Scores range from 0-100 based on Levenshtein distance
- Useful for handling typos like "hiy0ri" or ambiguous inputs like "日和莉"
Parsing Concatenated Team Strings
The parse_team(namestr) method (lines 58-71) handles bulk character identification from run-on strings such as "日和优衣怜":
- Repeatedly calls
self._roster.longest_prefix(namestr)to extract the longest matching alias from the remaining string - Appends the corresponding ID to the result list and removes the matched prefix
- Collects unrecognized characters into a separate string for error reporting
- Example:
roster.parse_team('日和优衣怜')returns([1001, 1002, 1003], '')
Helper Functions and Public API
The module exposes convenient wrappers around the Roster methods:
name2id(name)(lines 77-79): Direct wrapper forget_id()fromname(name)(lines 85-88): Returns aCharaobject instance from a name stringis_npc(id)(lines 95-99): Checks if the resolved ID corresponds to a non-player character
Practical Code Examples
# Exact conversion using official and variant names
>>> from hoshino.modules.priconne.chara import name2id
>>> name2id('日和') # Chinese official name
1001
>>> name2id('猫拳') # Nickname variant
1001
# Handling typos through fuzzy matching
>>> from hoshino.modules.priconne.chara import guess_id
>>> guess_id('日和莉')
(1001, '日和', 90)
# Parsing a concatenated team composition
>>> from hoshino.modules.priconne.chara import roster
>>> roster.parse_team('日和优衣怜')
([1001, 1002, 1003], '')
# Identifying unknown characters
>>> roster.parse_team('日和X优衣')
([1001, 1002], 'X')
Summary
- The
Rosterclass inhoshino/modules/priconne/chara.pyuses apygtrie.CharTrieto store normalized character aliases mapped to integer IDs. update()rebuilds the trie from_pcr_data.CHARA_NAME, normalizing all aliases viautil.normalize_str.get_id()provides exact O(1) lookups, whileguess_id()usesfuzzywuzzyfor typo-tolerant matching.parse_team()enables extraction of multiple character IDs from concatenated strings using longest-prefix matching.- Adding new aliases requires only updating the
CHARA_NAMEdictionary in_pcr_data.py; the trie rebuilds automatically on the next module reload.
Frequently Asked Questions
How does HoshinoBot handle misspelled character names?
When exact lookup fails, the system falls back to guess_id(), which uses fuzzywuzzy.process.extractOne to calculate similarity scores against all stored aliases. This catches typos like "hiy0ri" (normalized to "hiyori") and returns the best match with a confidence score ranging from 0 to 100.
What data structure enables fast alias lookups in chara.py?
The Roster class maintains a pygtrie.CharTrie (character trie) that stores normalized alias strings as keys and character IDs as values. This prefix-tree structure supports O(m) longest-prefix lookups where m is the length of the input string, making it efficient for parsing both individual names and concatenated team compositions.
How does the system parse multiple characters from a single string?
The parse_team(namestr) method iteratively applies longest_prefix() to the input string, extracting the longest matching alias at each step. It removes matched prefixes from the remaining string and continues until the string is empty or only unknown characters remain, returning a list of character IDs and any unrecognized text.
Where are the character aliases defined in the repository?
All aliases reside in hoshino/modules/priconne/_pcr_data.py within the CHARA_NAME dictionary. Each entry maps a character ID to a list of strings containing official names, Japanese names, English transliterations, nicknames, and common typo variants. The Roster.update() method reads this dictionary to populate the trie at runtime.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →