# How Karukan Supports Emoji Input: Kana Readings and Slack-Style :trigger: Queries

> Discover how Karukan supports emoji input with kana readings and Slack style :trigger: queries via its unified EmojiRewriter and pre-computed lookup tables.

- Repository: [Hitoshi Togasaki/karukan](https://github.com/togatoga/karukan)
- Tags: deep-dive
- Published: 2026-07-03

---

**Karukan handles emoji input through a unified `EmojiRewriter` that matches both Hiragana readings and Slack-style `:trigger` queries against pre-computed lookup tables loaded from Mozc data.**

The `karukan` repository implements a Rust-based Input Method Editor (IME) that supports intelligent emoji completion through two distinct input paths. According to the source code in [`karukan-engine/src/rewriter/emoji.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/rewriter/emoji.rs), the system uses compile-time loaded YAML data to map Japanese readings and ASCII triggers to Unicode emoji characters.

## Emoji Rewriter Architecture

The emoji functionality centers on a single struct implementing the generic `Rewriter` trait. This design allows the engine to treat emoji candidates as first-class rewrite suggestions alongside other conversion candidates.

### The EmojiRewriter Implementation

The `EmojiRewriter` struct serves as the entry point for all emoji-related transformations. When the candidate generation pipeline processes user input (as seen in [`karukan-im/src/core/engine/conversion.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/conversion.rs)), it invokes `EmojiRewriter::rewrite` with the current candidate string. This method branches based on whether the input starts with a colon character, routing to either the Hiragana reading lookup or the Slack-style trigger matcher.

### Data Sources and Lookup Tables

The rewriter relies on static data compiled from Mozc’s `emoji_data.tsv` and processed by [`scripts/emoji_porter.py`](https://github.com/togatoga/karukan/blob/main/scripts/emoji_porter.py). At compile time, the system embeds [`karukan-engine/data/emoji.yml`](https://github.com/togatoga/karukan/blob/main/karukan-engine/data/emoji.yml) using `include_str!` (line 56). Runtime initialization creates two primary lookup structures within a `LazyLock`:

- **`by_reading: HashMap<String, Vec<String>>`** – Maps Hiragana strings like `"わらい"` to emoji characters
- **`triggers: Vec<(String, String)>`** – Stores ASCII trigger pairs for subsequence matching

The porting script generates triggers from three sources: manual aliases, CLDR snake-case names, and romaji transliterations of Hiragana readings (lines 33-41).

## Two Input Methods for Emoji Retrieval

Karukan supports complementary input styles to accommodate both native Japanese typing and international keyboard workflows.

### Hiragana Reading Lookup

When a candidate string does **not** begin with a colon, the rewriter queries `EMOJI_TABLE.by_reading.get(candidate)`. This path connects traditional Japanese input to emoji selection. For example, typing `"ぴえん"` matches the reading for 🥺 (pleading face).

The system constructs descriptive labels using `format_description` (lines 176-182), appending metadata like `"絵文字 笑顔"` to help users identify the correct character.

### Slack-Style :trigger Queries

Inputs beginning with `:` activate the fuzzy trigger matcher. The rewriter strips the leading colon and performs subsequence matching against all entries in `EMOJI_TABLE.triggers`. This supports familiar shortcuts like `:smile` for 😄 and `:pien` (romaji derived from the Hiragana reading) for 🥺.

The `format_trigger_description` function (lines 84-94) builds UI labels showing the full trigger, such as `"`:smile 笑顔`"`.

## Fuzzy Matching Algorithm

The trigger matching employs a heuristic ranking system borrowed from peco. The `best_match_score` function (lines 20-27) evaluates candidates based on:

1. **Longest contiguous run** – calculated by `longest_run_from` (lines 49-73)
2. **Earliest start position** – preferring matches at the beginning of triggers
3. **Shortest trigger length** – favoring more specific triggers over generic ones

This algorithm enables partial matches: typing `:hlo` matches `"halo"` inside `"smiling_face_with_halo"` (😇).

Both lookup paths share a `seen` HashSet to deduplicate results when multiple triggers resolve to the same emoji character.

## Code Examples

```rust
use karukan_engine::rewriter::emoji::EmojiRewriter;

// Hiragana reading lookup
let rewriter = EmojiRewriter::new();
let candidates = rewriter.rewrite("ぴえん");
assert!(candidates.iter().any(|(emoji, _)| emoji == "🥺"));

```

```rust
// Slack-style trigger matching
let candidates = rewriter.rewrite(":smile");
assert!(candidates.iter().any(|(emoji, desc)| {
    emoji == "😄" && desc.as_ref().unwrap().contains(":smile")
}));

```

```rust
// Romaji trigger generated from Japanese reading
let candidates = rewriter.rewrite(":pien");
assert!(candidates.iter().any(|(emoji, _)| emoji == "🥺"));

```

```rust
// Fuzzy subsequence matching
let candidates = rewriter.rewrite(":hlo");
assert!(candidates.iter().any(|(emoji, _)| emoji == "😇"));

```

## Summary

- **Single rewriter interface**: `EmojiRewriter` in [`karukan-engine/src/rewriter/emoji.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/rewriter/emoji.rs) implements the `Rewriter` trait for seamless integration.
- **Dual lookup strategy**: `EMOJI_TABLE.by_reading` handles Hiragana input while `EMOJI_TABLE.triggers` supports ASCII shortcuts.
- **Compile-time data**: [`emoji.yml`](https://github.com/togatoga/karukan/blob/main/emoji.yml) embeds Mozc data via `include_str!`, eliminating runtime file I/O.
- **Peco-style fuzzy matching**: The `best_match_score` and `longest_run_from` functions rank triggers by contiguous match length and position.
- **Deduplication**: A shared `seen` set prevents duplicate emoji entries when multiple triggers match.

## Frequently Asked Questions

### What data source does Karukan use for emoji metadata?

Karukan derives its emoji data from Mozc’s `emoji_data.tsv`, processed offline by [`scripts/emoji_porter.py`](https://github.com/togatoga/karukan/blob/main/scripts/emoji_porter.py) into [`karukan-engine/data/emoji.yml`](https://github.com/togatoga/karukan/blob/main/karukan-engine/data/emoji.yml). The script augments the base data with CLDR snake-case names, manual aliases, and romaji transliterations of Japanese readings.

### How does the fuzzy matching algorithm work for Slack-style triggers?

The system uses `best_match_score` (lines 20-27) to rank triggers by three criteria: longest contiguous character run (computed by `longest_run_from`), earliest match position, and shortest overall trigger length. This allows partial matches like `:hlo` to find `"smiling_face_with_halo"`.

### Can I use romaji readings to find Japanese emoji names?

Yes. The [`emoji_porter.py`](https://github.com/togatoga/karukan/blob/main/emoji_porter.py) script automatically generates romaji triggers from Hiragana readings, enabling inputs like `:pien` to match ぴえん (🥺). These synthetic triggers coexist with manual aliases and CLDR names in the trigger table.

### Where does the emoji data reside in the compiled binary?

The [`emoji.yml`](https://github.com/togatoga/karukan/blob/main/emoji.yml) file is embedded at compile time using `include_str!` (line 56) and parsed into static `LazyLock` tables. This approach ensures zero runtime file-system dependencies while providing fast HashMap lookups for both reading and trigger queries.