Candidate Priority Order in Karukan's IME: Learning, User Dict, Model, System Dict, Rewriter
The candidate priority order in Karukan's IME follows a strict five-level hierarchy: Learning cache takes precedence over User Dictionary entries, which in turn override Model predictions, followed by System Dictionary entries, and finally Rewriter variants.
The togatoga/karukan input method engine assembles conversion candidates using a deterministic priority chain defined in the core conversion logic. This ordering ensures that personal learning history and explicit user customizations surface above generic model predictions, creating a personalized typing experience.
The Five-Level Candidate Priority Hierarchy
When Karukan's engine generates candidates for a given pre-edit buffer, it iterates through sources in the following strict sequence. Each candidate is tagged with a CandidateSource enum variant that determines its display rank.
1. Learning Cache (Highest Priority)
The learning cache holds candidates derived from the user's own conversion history and most-recent selections. This source receives top priority so the IME can instantly recall what the user previously selected for the same reading. When the engine processes input in karukan-im/src/core/engine/input.rs, it first injects learning candidates before querying other sources.
2. User Dictionary
Entries from the user dictionary follow immediately after the learning cache. These include custom words and phrases the user has added manually. By placing explicit user-added entries above model and system suggestions, Karukan ensures that personal terminology overrides generic predictions.
3. Model Predictions
Model predictions generated by the neural kana-kanji model (LLM) occupy the third tier. This placement gives the neural model the opportunity to surface context-aware suggestions while still respecting the user's explicit history and custom dictionary entries. The model candidates are produced via the engine's model.predict() method.
4. System Dictionary
The system dictionary—the built-in dictionary that ships with Karukan—provides a solid fallback when the model lacks confident output. These entries are queried after the learning cache, user dictionary, and model have been exhausted.
5. Rewriter and Fallback (Lowest Priority)
Finally, the rewriter chain supplies variant forms such as full-width/half-width conversions, case changes, and symbol substitutions. These fallback candidates appear last in the list, ensuring that primary linguistic conversions always precede formatting variants.
Implementation in the Conversion Engine
The priority order is hard-coded in karukan-im/src/core/engine/conversion.rs, where a source comment explicitly documents the sequence:
/// Priority: Learning → User Dictionary → Model → System Dictionary → Fallback
The CandidateSource enum defined in karukan-im/src/core/engine/mod.rs provides the variant labels used during candidate construction. The following simplified illustration from the conversion pipeline demonstrates how the engine assembles candidates in priority order:
fn build_candidates(buf: &InputBuffer, engine: &IMEEngine) -> Vec<AnnotatedCandidate> {
let mut candidates = Vec::new();
// 1️⃣ Learning cache
if let Some(learn) = engine.learning.as_ref() {
for c in learn.lookup(&buf.reading) {
candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::Learning));
}
}
// 2️⃣ User dictionary
for c in engine.user_dict.lookup(&buf.reading) {
candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::UserDict));
}
// 3️⃣ Model predictions
for c in engine.model.predict(&buf.reading) {
candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::Model));
}
// 4️⃣ System dictionary
for c in engine.system_dict.lookup(&buf.reading) {
candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::SystemDict));
}
// 5️⃣ Rewriter / fallback
for c in engine.rewriter.apply(&candidates) {
candidates.push(AnnotatedCandidate::new(c.text, CandidateSource::Rewriter));
}
candidates
}
The AnnotatedCandidate struct in karukan-im/src/core/candidate.rs carries these source labels, allowing the UI to display the highest-priority candidates first while maintaining the ability to scroll to lower-priority alternatives.
Why This Priority Order Matters
This strict hierarchy ensures predictability and personalization in the typing experience. By checking the learning cache first, Karukan minimizes cognitive load by surfacing recently-used conversions instantly. User dictionary entries take precedence over model predictions, guaranteeing that explicit customizations are never buried under neural suggestions. The model predictions rank above the system dictionary, allowing context-aware neural outputs to supplement—but not override—personal data. Finally, the rewriter stage handles orthographic variants without cluttering the primary candidate list.
Summary
- Learning cache ranks highest to provide instant recall of recent user selections.
- User Dictionary entries follow, ensuring explicit customizations override generic suggestions.
- Model predictions come third, offering neural context-aware candidates after personal sources.
- System Dictionary provides fallback coverage for standard vocabulary.
- Rewriter variants appear last, handling formatting and symbol substitutions.
Frequently Asked Questions
Where is the candidate priority order defined in Karukan's source code?
The priority order is explicitly documented in a comment within karukan-im/src/core/engine/conversion.rs. The implementation iterates through CandidateSource variants in the sequence Learning → User Dictionary → Model → System Dictionary → Rewriter, tagging each candidate accordingly.
How does the Learning cache get priority over the Model predictions?
During candidate construction in the conversion pipeline, the engine queries the learning cache (implemented in karukan-engine/src/learning.rs) before invoking the model's prediction methods. This hard-coded sequence ensures that CandidateSource::Learning candidates are appended to the list before CandidateSource::Model candidates are considered.
Can I customize the candidate priority order in Karukan?
No, the priority order is currently hard-coded in the conversion engine as of the latest source code. The sequence is fixed in the build_candidates logic within conversion.rs, with no configuration options exposed to reorder the precedence of Learning, User Dictionary, Model, System Dictionary, or Rewriter sources.
What happens if no candidates are found in the first four sources?
If the Learning cache, User Dictionary, Model, and System Dictionary all return empty results, the engine falls back to the Rewriter chain. The Rewriter generates variant forms such as width and case conversions, ensuring that the candidate list is never empty even when no dictionary matches exist.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →