How Karukan Supports Emoji Input: Kana Readings and Slack-Style :trigger: Queries

Karukan handles emoji input through a unified EmojiRewriter that matches both Hiragana readings and Slack-style :trigger queries against pre-computed lookup tables loaded from Mozc data.

The karukan repository implements a Rust-based Input Method Editor (IME) that supports intelligent emoji completion through two distinct input paths. According to the source code in karukan-engine/src/rewriter/emoji.rs, the system uses compile-time loaded YAML data to map Japanese readings and ASCII triggers to Unicode emoji characters.

Emoji Rewriter Architecture

The emoji functionality centers on a single struct implementing the generic Rewriter trait. This design allows the engine to treat emoji candidates as first-class rewrite suggestions alongside other conversion candidates.

The EmojiRewriter Implementation

The EmojiRewriter struct serves as the entry point for all emoji-related transformations. When the candidate generation pipeline processes user input (as seen in karukan-im/src/core/engine/conversion.rs), it invokes EmojiRewriter::rewrite with the current candidate string. This method branches based on whether the input starts with a colon character, routing to either the Hiragana reading lookup or the Slack-style trigger matcher.

Data Sources and Lookup Tables

The rewriter relies on static data compiled from Mozc’s emoji_data.tsv and processed by scripts/emoji_porter.py. At compile time, the system embeds karukan-engine/data/emoji.yml using include_str! (line 56). Runtime initialization creates two primary lookup structures within a LazyLock:

  • by_reading: HashMap<String, Vec<String>> – Maps Hiragana strings like "わらい" to emoji characters
  • triggers: Vec<(String, String)> – Stores ASCII trigger pairs for subsequence matching

The porting script generates triggers from three sources: manual aliases, CLDR snake-case names, and romaji transliterations of Hiragana readings (lines 33-41).

Two Input Methods for Emoji Retrieval

Karukan supports complementary input styles to accommodate both native Japanese typing and international keyboard workflows.

Hiragana Reading Lookup

When a candidate string does not begin with a colon, the rewriter queries EMOJI_TABLE.by_reading.get(candidate). This path connects traditional Japanese input to emoji selection. For example, typing "ぴえん" matches the reading for 🥺 (pleading face).

The system constructs descriptive labels using format_description (lines 176-182), appending metadata like "絵文字 笑顔" to help users identify the correct character.

Slack-Style :trigger Queries

Inputs beginning with : activate the fuzzy trigger matcher. The rewriter strips the leading colon and performs subsequence matching against all entries in EMOJI_TABLE.triggers. This supports familiar shortcuts like :smile for 😄 and :pien (romaji derived from the Hiragana reading) for 🥺.

The format_trigger_description function (lines 84-94) builds UI labels showing the full trigger, such as ":smile 笑顔".

Fuzzy Matching Algorithm

The trigger matching employs a heuristic ranking system borrowed from peco. The best_match_score function (lines 20-27) evaluates candidates based on:

  1. Longest contiguous run – calculated by longest_run_from (lines 49-73)
  2. Earliest start position – preferring matches at the beginning of triggers
  3. Shortest trigger length – favoring more specific triggers over generic ones

This algorithm enables partial matches: typing :hlo matches "halo" inside "smiling_face_with_halo" (😇).

Both lookup paths share a seen HashSet to deduplicate results when multiple triggers resolve to the same emoji character.

Code Examples

use karukan_engine::rewriter::emoji::EmojiRewriter;

// Hiragana reading lookup
let rewriter = EmojiRewriter::new();
let candidates = rewriter.rewrite("ぴえん");
assert!(candidates.iter().any(|(emoji, _)| emoji == "🥺"));
// Slack-style trigger matching
let candidates = rewriter.rewrite(":smile");
assert!(candidates.iter().any(|(emoji, desc)| {
    emoji == "😄" && desc.as_ref().unwrap().contains(":smile")
}));
// Romaji trigger generated from Japanese reading
let candidates = rewriter.rewrite(":pien");
assert!(candidates.iter().any(|(emoji, _)| emoji == "🥺"));
// Fuzzy subsequence matching
let candidates = rewriter.rewrite(":hlo");
assert!(candidates.iter().any(|(emoji, _)| emoji == "😇"));

Summary

  • Single rewriter interface: EmojiRewriter in karukan-engine/src/rewriter/emoji.rs implements the Rewriter trait for seamless integration.
  • Dual lookup strategy: EMOJI_TABLE.by_reading handles Hiragana input while EMOJI_TABLE.triggers supports ASCII shortcuts.
  • Compile-time data: emoji.yml embeds Mozc data via include_str!, eliminating runtime file I/O.
  • Peco-style fuzzy matching: The best_match_score and longest_run_from functions rank triggers by contiguous match length and position.
  • Deduplication: A shared seen set prevents duplicate emoji entries when multiple triggers match.

Frequently Asked Questions

What data source does Karukan use for emoji metadata?

Karukan derives its emoji data from Mozc’s emoji_data.tsv, processed offline by scripts/emoji_porter.py into karukan-engine/data/emoji.yml. The script augments the base data with CLDR snake-case names, manual aliases, and romaji transliterations of Japanese readings.

How does the fuzzy matching algorithm work for Slack-style triggers?

The system uses best_match_score (lines 20-27) to rank triggers by three criteria: longest contiguous character run (computed by longest_run_from), earliest match position, and shortest overall trigger length. This allows partial matches like :hlo to find "smiling_face_with_halo".

Can I use romaji readings to find Japanese emoji names?

Yes. The emoji_porter.py script automatically generates romaji triggers from Hiragana readings, enabling inputs like :pien to match ぴえん (🥺). These synthetic triggers coexist with manual aliases and CLDR names in the trigger table.

Where does the emoji data reside in the compiled binary?

The emoji.yml file is embedded at compile time using include_str! (line 56) and parsed into static LazyLock tables. This approach ensures zero runtime file-system dependencies while providing fast HashMap lookups for both reading and trigger queries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →