How the Character Name Alias Conversion System Works in HoshinoBot

The character name alias conversion system in HoshinoBot uses a trie-based Roster class combined with fuzzy string matching to map user-input names, nicknames, and misspellings to unique character IDs for Princess Connect! Re:Dive.

The chara.py module in the ice9coffee/hoshinobot repository powers character identification across the bot's Princess Connect functionality. According to the source code in hoshino/modules/priconne/chara.py, the system handles hundreds of multilingual aliases by normalizing input and performing fast prefix-tree lookups.

Core Architecture of the Roster Class

The alias system centers on the Roster class, which maintains a pygtrie.CharTrie data structure to enable longest-prefix matching. This design allows the bot to efficiently parse concatenated team strings and resolve individual character names in a single pass.

Trie Initialization and Data Loading

When the module imports, it instantiates a global roster = Roster() object. The constructor automatically invokes update() (lines 34-47), which performs the following steps:

  1. Reloads the _pcr_data module containing the CHARA_NAME dictionary
  2. Iterates over every character ID and its associated alias list
  3. Normalizes each alias using util.normalize_str (removes whitespace, converts to lowercase, strips special characters)
  4. Inserts normalized aliases into the trie with their corresponding character IDs as values
  5. Logs warnings for any duplicate aliases encountered during population

This initialization happens at lines 30-33 in hoshino/modules/priconne/chara.py, ensuring the trie remains synchronized with the underlying data source.

Lookup Mechanisms and Conversion Methods

The system provides three distinct search strategies to accommodate exact matches, typos, and complex team compositions.

Exact Matching with get_id()

The get_id(name) method (lines 49-52) performs direct trie lookups for normalized input:

  • Returns the exact character ID if the normalized alias exists
  • Returns 1000 (the UNKNOWN constant) if no match is found
  • Example: name2id('日和'), name2id('hiyori'), and name2id('猫拳') all return 1001

Fuzzy Matching for Misspellings

When exact lookup fails, the guess_id(name) method (lines 53-57) employs fuzzywuzzy.process.extractOne to find the closest matching alias:

  • Returns a tuple containing (character_id, matched_alias, similarity_score)
  • Scores range from 0-100 based on Levenshtein distance
  • Useful for handling typos like "hiy0ri" or ambiguous inputs like "日和莉"

Parsing Concatenated Team Strings

The parse_team(namestr) method (lines 58-71) handles bulk character identification from run-on strings such as "日和优衣怜":

  • Repeatedly calls self._roster.longest_prefix(namestr) to extract the longest matching alias from the remaining string
  • Appends the corresponding ID to the result list and removes the matched prefix
  • Collects unrecognized characters into a separate string for error reporting
  • Example: roster.parse_team('日和优衣怜') returns ([1001, 1002, 1003], '')

Helper Functions and Public API

The module exposes convenient wrappers around the Roster methods:

  • name2id(name) (lines 77-79): Direct wrapper for get_id()
  • fromname(name) (lines 85-88): Returns a Chara object instance from a name string
  • is_npc(id) (lines 95-99): Checks if the resolved ID corresponds to a non-player character

Practical Code Examples


# Exact conversion using official and variant names

>>> from hoshino.modules.priconne.chara import name2id
>>> name2id('日和')        # Chinese official name

1001
>>> name2id('猫拳')        # Nickname variant

1001

# Handling typos through fuzzy matching

>>> from hoshino.modules.priconne.chara import guess_id
>>> guess_id('日和莉')
(1001, '日和', 90)

# Parsing a concatenated team composition

>>> from hoshino.modules.priconne.chara import roster
>>> roster.parse_team('日和优衣怜')
([1001, 1002, 1003], '')

# Identifying unknown characters

>>> roster.parse_team('日和X优衣')
([1001, 1002], 'X')

Summary

  • The Roster class in hoshino/modules/priconne/chara.py uses a pygtrie.CharTrie to store normalized character aliases mapped to integer IDs.
  • update() rebuilds the trie from _pcr_data.CHARA_NAME, normalizing all aliases via util.normalize_str.
  • get_id() provides exact O(1) lookups, while guess_id() uses fuzzywuzzy for typo-tolerant matching.
  • parse_team() enables extraction of multiple character IDs from concatenated strings using longest-prefix matching.
  • Adding new aliases requires only updating the CHARA_NAME dictionary in _pcr_data.py; the trie rebuilds automatically on the next module reload.

Frequently Asked Questions

How does HoshinoBot handle misspelled character names?

When exact lookup fails, the system falls back to guess_id(), which uses fuzzywuzzy.process.extractOne to calculate similarity scores against all stored aliases. This catches typos like "hiy0ri" (normalized to "hiyori") and returns the best match with a confidence score ranging from 0 to 100.

What data structure enables fast alias lookups in chara.py?

The Roster class maintains a pygtrie.CharTrie (character trie) that stores normalized alias strings as keys and character IDs as values. This prefix-tree structure supports O(m) longest-prefix lookups where m is the length of the input string, making it efficient for parsing both individual names and concatenated team compositions.

How does the system parse multiple characters from a single string?

The parse_team(namestr) method iteratively applies longest_prefix() to the input string, extracting the longest matching alias at each step. It removes matched prefixes from the remaining string and continues until the string is empty or only unknown characters remain, returning a list of character IDs and any unrecognized text.

Where are the character aliases defined in the repository?

All aliases reside in hoshino/modules/priconne/_pcr_data.py within the CHARA_NAME dictionary. Each entry maps a character ID to a list of strings containing official names, Japanese names, English transliterations, nicknames, and common typo variants. The Roster.update() method reads this dictionary to populate the trie at runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →