How Karukan Determines Adaptive Model Selection Based on Input
Karukan selects between its main Kana-Kanji model and light model by evaluating runtime token count, requested candidate count, and a latency flag that triggers when conversions exceed max_latency_ms.
Karukan is an open-source Japanese input method editor that dynamically balances accuracy and performance through intelligent adaptive model selection. The engine determines whether to use its large main model or a smaller light model by analyzing real-time input characteristics and historical conversion latency. This decision logic is implemented in the karukan-im/src/core/engine/strategy.rs file and controlled by the StrategyMode configuration enum defined in karukan-im/src/config/settings.rs.
The Three Strategy Modes
Karukan supports three distinct conversion strategies defined in karukan-im/src/config/settings.rs that govern how model selection occurs:
- Light Mode: Forces the use of the light model exclusively. The light model is loaded into the main slot, and both auto-suggest and explicit conversion (Space key) use this smaller model.
- Main Mode: Uses only the large main model for all conversions, regardless of input length or latency.
- Adaptive Mode: Dynamically switches between models based on runtime conditions, input token count, and measured latency.
The adaptive mode is where the sophisticated decision logic resides, enabling the engine to optimize for both speed and accuracy.
How Adaptive Mode Makes Decisions
When conversion.strategy is set to Adaptive, Karukan evaluates three critical runtime inputs in the determine_adaptive_strategy function:
- Token count of the current reading – measured using the main model's tokenizer
- Number of candidate suggestions requested – one candidate indicates auto-suggest; more than one indicates explicit beam-search conversion
- Adaptive latency flag – a boolean set to
truewhen the most recent main-model conversion exceededEngineConfig.max_latency_ms
The Decision Function Implementation
The complete decision logic resides in karukan-im/src/core/engine/strategy.rs:
fn determine_adaptive_strategy(
reading_tokens: usize,
num_candidates: usize,
has_light_model: bool,
adaptive_use_light_model: bool,
config: &EngineConfig,
) -> ConversionStrategy {
if !has_light_model {
return ConversionStrategy::MainModelOnly;
}
if num_candidates == 1 {
// auto‑suggest
if adaptive_use_light_model {
ConversionStrategy::LightModelOnly
} else {
ConversionStrategy::MainModelOnly
}
} else {
// user pressed Space
if adaptive_use_light_model {
ConversionStrategy::LightModelOnly
} else if reading_tokens <= config.short_input_threshold {
ConversionStrategy::ParallelBeam {
beam_width: num_candidates.min(config.beam_width),
}
} else {
ConversionStrategy::LightModelOnly
}
}
}
Decision Logic for Auto-Suggest
For automatic suggestions where num_candidates == 1, the logic is straightforward:
- If
adaptive_use_light_modelistrue(previous conversion was too slow), the engine returnsConversionStrategy::LightModelOnly - Otherwise, it returns
ConversionStrategy::MainModelOnly
Decision Logic for Explicit Conversion
When the user presses Space to request multiple candidates (num_candidates > 1), the strategy becomes more nuanced:
- If the adaptive flag is
true, the engine immediately selectsLightModelOnlyto avoid another slow conversion - If the adaptive flag is
falseandreading_tokens <= short_input_threshold, the engine usesParallelBeamto run both models simultaneously and beam-search the best results - If the input exceeds the threshold (
reading_tokens > short_input_threshold), the engine proactively selectsLightModelOnlyfor long inputs
The short_input_threshold (default 40 tokens) and beam_width (default 9) are user-configurable values that control when parallel execution occurs.
The Adaptive Latency Flag Mechanism
After each conversion, the engine calls update_adaptive_model_flag to adjust the adaptive state based on performance metrics:
match strategy {
ConversionStrategy::MainModelOnly | ConversionStrategy::ParallelBeam { .. } => {
self.metrics.adaptive_use_light_model =
self.metrics.conversion_ms > self.config.max_latency_ms;
}
_ => {} // LightModelOnly or MainModelBeam do not affect the flag
}
When a conversion using the main model or parallel beam exceeds max_latency_ms, the flag is set to true. This flag is automatically cleared when a new word starts (when reading_tokens resets to 0), ensuring that latency adaptations are contextually relevant to the current input session.
Code Examples
Example 1: Auto-Suggest with Fast Main Model
// Settings: Adaptive mode, max_latency_ms = 200 ms
engine.metrics.adaptive_use_light_model = false; // last conversion was fast
let strategy = engine.determine_strategy("かんじ", 1);
// → ConversionStrategy::MainModelOnly
Example 2: Auto-Suggest After Slow Conversion
engine.metrics.adaptive_use_light_model = true; // previous conversion took >200 ms
let strategy = engine.determine_strategy("かんじ", 1);
// → ConversionStrategy::LightModelOnly
Example 3: Short Input with Space Key
// reading length = 12 tokens, short_input_threshold = 20
let strategy = engine.determine_strategy("にほんご", 5);
// → ConversionStrategy::ParallelBeam { beam_width: 5 }
Example 4: Long Input with Space Key
// reading length = 45 tokens (exceeds short_input_threshold)
let strategy = engine.determine_strategy("おおさかこうえんであそんでいた", 5);
// → ConversionStrategy::LightModelOnly
Configuration and Key Files
The adaptive model selection system spans several core files in the Karukan repository:
karukan-im/src/core/engine/strategy.rs: Implements thedetermine_adaptive_strategyfunction and the decision logic for all conversion strategies.karukan-im/src/config/settings.rs: Defines theStrategyModeenum and configuration fields includingmax_latency_ms,short_input_threshold, andbeam_width.karukan-im/src/core/engine/mod.rs: Contains theInputMethodEnginestruct that holds runtime metrics (conversion_ms,adaptive_use_light_model) and invokes the strategy determination.karukan-engine/src/kanji/backend.rs: Provides the token-counting functionality used to measure input length for strategy decisions.
Summary
Karukan's adaptive model selection strategy intelligently balances accuracy and latency by:
- Monitoring conversion latency in real-time and setting an adaptive flag when the main model exceeds
max_latency_ms - Selecting the light model automatically for auto-suggest when previous conversions were slow
- Running both models in parallel via
ParallelBeamfor short explicit conversions under the threshold - Proactively using the light model for long inputs that exceed
short_input_threshold - Resetting the adaptive state when input context changes (new word starts)
This approach ensures fast auto-suggestions on short inputs while maintaining low-latency fallback options for complex or lengthy conversions.
Frequently Asked Questions
What triggers Karukan to switch from the main model to the light model?
The switch occurs when the adaptive_use_light_model flag is set to true, which happens after a main-model conversion exceeds the user-configured max_latency_ms. Additionally, for long inputs in adaptive mode where reading_tokens exceeds short_input_threshold, the engine proactively selects the light model even without a latency violation.
How does the short input threshold affect conversion strategy?
The short_input_threshold (default 40 tokens) determines whether Karukan runs both models in parallel. When a user requests explicit conversion (Space key) on a short input below this threshold, and the adaptive flag is false, the engine uses ParallelBeam to execute both models simultaneously and return the best candidate. Inputs above this threshold bypass parallel execution and use the light model directly.
What is the difference between Light Mode and Adaptive Mode?
Light Mode forces exclusive use of the small model for all conversions, sacrificing accuracy for consistent speed. Adaptive Mode dynamically selects between the main and light models based on real-time input length, candidate count, and measured latency, providing high accuracy for fast queries and automatic fallback to the light model when performance degrades.
How does Karukan measure conversion latency?
The engine records the duration of each conversion in milliseconds within the InputMethodEngine metrics struct. After each conversion using the main model or parallel beam strategy, it compares metrics.conversion_ms against config.max_latency_ms. If the conversion time exceeds this threshold, the adaptive_use_light_model flag is set to true for subsequent operations until the input context resets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →