# Candidate Priority Order in Karukan's IME: Learning, User Dict, Model, System Dict, Rewriter

> Understand Karukan's IME candidate priority order: Learning, User Dict, Model, System Dict, and Rewriter. Discover how your input is prioritized for efficient text prediction.

- Repository: [Hitoshi Togasaki/karukan](https://github.com/togatoga/karukan)
- Tags: internals
- Published: 2026-07-03

---

**The candidate priority order in Karukan's IME follows a strict five-level hierarchy: Learning cache takes precedence over User Dictionary entries, which in turn override Model predictions, followed by System Dictionary entries, and finally Rewriter variants.**

The `togatoga/karukan` input method engine assembles conversion candidates using a deterministic priority chain defined in the core conversion logic. This ordering ensures that personal learning history and explicit user customizations surface above generic model predictions, creating a personalized typing experience.

## The Five-Level Candidate Priority Hierarchy

When Karukan's engine generates candidates for a given pre-edit buffer, it iterates through sources in the following strict sequence. Each candidate is tagged with a `CandidateSource` enum variant that determines its display rank.

### 1. Learning Cache (Highest Priority)

The **learning cache** holds candidates derived from the user's own conversion history and most-recent selections. This source receives top priority so the IME can instantly recall what the user previously selected for the same reading. When the engine processes input in [`karukan-im/src/core/engine/input.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/input.rs), it first injects learning candidates before querying other sources.

### 2. User Dictionary

Entries from the **user dictionary** follow immediately after the learning cache. These include custom words and phrases the user has added manually. By placing explicit user-added entries above model and system suggestions, Karukan ensures that personal terminology overrides generic predictions.

### 3. Model Predictions

**Model predictions** generated by the neural kana-kanji model (LLM) occupy the third tier. This placement gives the neural model the opportunity to surface context-aware suggestions while still respecting the user's explicit history and custom dictionary entries. The model candidates are produced via the engine's `model.predict()` method.

### 4. System Dictionary

The **system dictionary**—the built-in dictionary that ships with Karukan—provides a solid fallback when the model lacks confident output. These entries are queried after the learning cache, user dictionary, and model have been exhausted.

### 5. Rewriter and Fallback (Lowest Priority)

Finally, the **rewriter** chain supplies variant forms such as full-width/half-width conversions, case changes, and symbol substitutions. These fallback candidates appear last in the list, ensuring that primary linguistic conversions always precede formatting variants.

## Implementation in the Conversion Engine

The priority order is hard-coded in [`karukan-im/src/core/engine/conversion.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/conversion.rs), where a source comment explicitly documents the sequence:

```rust
/// Priority: Learning → User Dictionary → Model → System Dictionary → Fallback

```

The `CandidateSource` enum defined in [`karukan-im/src/core/engine/mod.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/mod.rs) provides the variant labels used during candidate construction. The following simplified illustration from the conversion pipeline demonstrates how the engine assembles candidates in priority order:

```rust
fn build_candidates(buf: &InputBuffer, engine: &IMEEngine) -> Vec<AnnotatedCandidate> {
    let mut candidates = Vec::new();

    // 1️⃣ Learning cache
    if let Some(learn) = engine.learning.as_ref() {
        for c in learn.lookup(&buf.reading) {
            candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::Learning));
        }
    }

    // 2️⃣ User dictionary
    for c in engine.user_dict.lookup(&buf.reading) {
        candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::UserDict));
    }

    // 3️⃣ Model predictions
    for c in engine.model.predict(&buf.reading) {
        candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::Model));
    }

    // 4️⃣ System dictionary
    for c in engine.system_dict.lookup(&buf.reading) {
        candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::SystemDict));
    }

    // 5️⃣ Rewriter / fallback
    for c in engine.rewriter.apply(&candidates) {
        candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::Rewriter));
    }

    candidates
}

```

The `AnnotatedCandidate` struct in [`karukan-im/src/core/candidate.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/candidate.rs) carries these source labels, allowing the UI to display the highest-priority candidates first while maintaining the ability to scroll to lower-priority alternatives.

## Why This Priority Order Matters

This strict hierarchy ensures **predictability** and **personalization** in the typing experience. By checking the learning cache first, Karukan minimizes cognitive load by surfacing recently-used conversions instantly. User dictionary entries take precedence over model predictions, guaranteeing that explicit customizations are never buried under neural suggestions. The model predictions rank above the system dictionary, allowing context-aware neural outputs to supplement—but not override—personal data. Finally, the rewriter stage handles orthographic variants without cluttering the primary candidate list.

## Summary

- **Learning cache** ranks highest to provide instant recall of recent user selections.
- **User Dictionary** entries follow, ensuring explicit customizations override generic suggestions.
- **Model predictions** come third, offering neural context-aware candidates after personal sources.
- **System Dictionary** provides fallback coverage for standard vocabulary.
- **Rewriter** variants appear last, handling formatting and symbol substitutions.

## Frequently Asked Questions

### Where is the candidate priority order defined in Karukan's source code?

The priority order is explicitly documented in a comment within [`karukan-im/src/core/engine/conversion.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/conversion.rs). The implementation iterates through `CandidateSource` variants in the sequence Learning → User Dictionary → Model → System Dictionary → Rewriter, tagging each candidate accordingly.

### How does the Learning cache get priority over the Model predictions?

During candidate construction in the conversion pipeline, the engine queries the learning cache (implemented in [`karukan-engine/src/learning.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/learning.rs)) before invoking the model's prediction methods. This hard-coded sequence ensures that `CandidateSource::Learning` candidates are appended to the list before `CandidateSource::Model` candidates are considered.

### Can I customize the candidate priority order in Karukan?

No, the priority order is currently hard-coded in the conversion engine as of the latest source code. The sequence is fixed in the `build_candidates` logic within [`conversion.rs`](https://github.com/togatoga/karukan/blob/main/conversion.rs), with no configuration options exposed to reorder the precedence of Learning, User Dictionary, Model, System Dictionary, or Rewriter sources.

### What happens if no candidates are found in the first four sources?

If the Learning cache, User Dictionary, Model, and System Dictionary all return empty results, the engine falls back to the Rewriter chain. The Rewriter generates variant forms such as width and case conversions, ensuring that the candidate list is never empty even when no dictionary matches exist.