How Karukan Determines Adaptive Model Selection Based on Input

Karukan selects between its main Kana-Kanji model and light model by evaluating runtime token count, requested candidate count, and a latency flag that triggers when conversions exceed max_latency_ms.

Karukan is an open-source Japanese input method editor that dynamically balances accuracy and performance through intelligent adaptive model selection. The engine determines whether to use its large main model or a smaller light model by analyzing real-time input characteristics and historical conversion latency. This decision logic is implemented in the karukan-im/src/core/engine/strategy.rs file and controlled by the StrategyMode configuration enum defined in karukan-im/src/config/settings.rs.

The Three Strategy Modes

Karukan supports three distinct conversion strategies defined in karukan-im/src/config/settings.rs that govern how model selection occurs:

  • Light Mode: Forces the use of the light model exclusively. The light model is loaded into the main slot, and both auto-suggest and explicit conversion (Space key) use this smaller model.
  • Main Mode: Uses only the large main model for all conversions, regardless of input length or latency.
  • Adaptive Mode: Dynamically switches between models based on runtime conditions, input token count, and measured latency.

The adaptive mode is where the sophisticated decision logic resides, enabling the engine to optimize for both speed and accuracy.

How Adaptive Mode Makes Decisions

When conversion.strategy is set to Adaptive, Karukan evaluates three critical runtime inputs in the determine_adaptive_strategy function:

  1. Token count of the current reading – measured using the main model's tokenizer
  2. Number of candidate suggestions requested – one candidate indicates auto-suggest; more than one indicates explicit beam-search conversion
  3. Adaptive latency flag – a boolean set to true when the most recent main-model conversion exceeded EngineConfig.max_latency_ms

The Decision Function Implementation

The complete decision logic resides in karukan-im/src/core/engine/strategy.rs:

fn determine_adaptive_strategy(
    reading_tokens: usize,
    num_candidates: usize,
    has_light_model: bool,
    adaptive_use_light_model: bool,
    config: &EngineConfig,
) -> ConversionStrategy {
    if !has_light_model {
        return ConversionStrategy::MainModelOnly;
    }

    if num_candidates == 1 {
        // auto‑suggest
        if adaptive_use_light_model {
            ConversionStrategy::LightModelOnly
        } else {
            ConversionStrategy::MainModelOnly
        }
    } else {
        // user pressed Space
        if adaptive_use_light_model {
            ConversionStrategy::LightModelOnly
        } else if reading_tokens <= config.short_input_threshold {
            ConversionStrategy::ParallelBeam {
                beam_width: num_candidates.min(config.beam_width),
            }
        } else {
            ConversionStrategy::LightModelOnly
        }
    }
}

Decision Logic for Auto-Suggest

For automatic suggestions where num_candidates == 1, the logic is straightforward:

  • If adaptive_use_light_model is true (previous conversion was too slow), the engine returns ConversionStrategy::LightModelOnly
  • Otherwise, it returns ConversionStrategy::MainModelOnly

Decision Logic for Explicit Conversion

When the user presses Space to request multiple candidates (num_candidates > 1), the strategy becomes more nuanced:

  • If the adaptive flag is true, the engine immediately selects LightModelOnly to avoid another slow conversion
  • If the adaptive flag is false and reading_tokens <= short_input_threshold, the engine uses ParallelBeam to run both models simultaneously and beam-search the best results
  • If the input exceeds the threshold (reading_tokens > short_input_threshold), the engine proactively selects LightModelOnly for long inputs

The short_input_threshold (default 40 tokens) and beam_width (default 9) are user-configurable values that control when parallel execution occurs.

The Adaptive Latency Flag Mechanism

After each conversion, the engine calls update_adaptive_model_flag to adjust the adaptive state based on performance metrics:

match strategy {
    ConversionStrategy::MainModelOnly | ConversionStrategy::ParallelBeam { .. } => {
        self.metrics.adaptive_use_light_model =
            self.metrics.conversion_ms > self.config.max_latency_ms;
    }
    _ => {} // LightModelOnly or MainModelBeam do not affect the flag
}

When a conversion using the main model or parallel beam exceeds max_latency_ms, the flag is set to true. This flag is automatically cleared when a new word starts (when reading_tokens resets to 0), ensuring that latency adaptations are contextually relevant to the current input session.

Code Examples

Example 1: Auto-Suggest with Fast Main Model

// Settings: Adaptive mode, max_latency_ms = 200 ms
engine.metrics.adaptive_use_light_model = false; // last conversion was fast
let strategy = engine.determine_strategy("かんじ", 1);
// → ConversionStrategy::MainModelOnly

Example 2: Auto-Suggest After Slow Conversion

engine.metrics.adaptive_use_light_model = true; // previous conversion took >200 ms
let strategy = engine.determine_strategy("かんじ", 1);
// → ConversionStrategy::LightModelOnly

Example 3: Short Input with Space Key

// reading length = 12 tokens, short_input_threshold = 20
let strategy = engine.determine_strategy("にほんご", 5);
// → ConversionStrategy::ParallelBeam { beam_width: 5 }

Example 4: Long Input with Space Key

// reading length = 45 tokens (exceeds short_input_threshold)
let strategy = engine.determine_strategy("おおさかこうえんであそんでいた", 5);
// → ConversionStrategy::LightModelOnly

Configuration and Key Files

The adaptive model selection system spans several core files in the Karukan repository:

Summary

Karukan's adaptive model selection strategy intelligently balances accuracy and latency by:

  • Monitoring conversion latency in real-time and setting an adaptive flag when the main model exceeds max_latency_ms
  • Selecting the light model automatically for auto-suggest when previous conversions were slow
  • Running both models in parallel via ParallelBeam for short explicit conversions under the threshold
  • Proactively using the light model for long inputs that exceed short_input_threshold
  • Resetting the adaptive state when input context changes (new word starts)

This approach ensures fast auto-suggestions on short inputs while maintaining low-latency fallback options for complex or lengthy conversions.

Frequently Asked Questions

What triggers Karukan to switch from the main model to the light model?

The switch occurs when the adaptive_use_light_model flag is set to true, which happens after a main-model conversion exceeds the user-configured max_latency_ms. Additionally, for long inputs in adaptive mode where reading_tokens exceeds short_input_threshold, the engine proactively selects the light model even without a latency violation.

How does the short input threshold affect conversion strategy?

The short_input_threshold (default 40 tokens) determines whether Karukan runs both models in parallel. When a user requests explicit conversion (Space key) on a short input below this threshold, and the adaptive flag is false, the engine uses ParallelBeam to execute both models simultaneously and return the best candidate. Inputs above this threshold bypass parallel execution and use the light model directly.

What is the difference between Light Mode and Adaptive Mode?

Light Mode forces exclusive use of the small model for all conversions, sacrificing accuracy for consistent speed. Adaptive Mode dynamically selects between the main and light models based on real-time input length, candidate count, and measured latency, providing high accuracy for fast queries and automatic fallback to the light model when performance degrades.

How does Karukan measure conversion latency?

The engine records the duration of each conversion in milliseconds within the InputMethodEngine metrics struct. After each conversion using the main model or parallel beam strategy, it compares metrics.conversion_ms against config.max_latency_ms. If the conversion time exceeds this threshold, the adaptive_use_light_model flag is set to true for subsequent operations until the input context resets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →