# How Karukan Determines Adaptive Model Selection Based on Input

> Learn how Karukan dynamically selects adaptive models based on runtime token count, candidate count, and latency flags for efficient conversions.

- Repository: [Hitoshi Togasaki/karukan](https://github.com/togatoga/karukan)
- Tags: deep-dive
- Published: 2026-07-03

---

**Karukan selects between its main Kana-Kanji model and light model by evaluating runtime token count, requested candidate count, and a latency flag that triggers when conversions exceed `max_latency_ms`.**

Karukan is an open-source Japanese input method editor that dynamically balances accuracy and performance through intelligent **adaptive model selection**. The engine determines whether to use its large main model or a smaller light model by analyzing real-time input characteristics and historical conversion latency. This decision logic is implemented in the [`karukan-im/src/core/engine/strategy.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/strategy.rs) file and controlled by the `StrategyMode` configuration enum defined in [`karukan-im/src/config/settings.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/config/settings.rs).

## The Three Strategy Modes

Karukan supports three distinct conversion strategies defined in [`karukan-im/src/config/settings.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/config/settings.rs) that govern how model selection occurs:

- **Light Mode**: Forces the use of the light model exclusively. The light model is loaded into the main slot, and both auto-suggest and explicit conversion (Space key) use this smaller model.
- **Main Mode**: Uses only the large main model for all conversions, regardless of input length or latency.
- **Adaptive Mode**: Dynamically switches between models based on runtime conditions, input token count, and measured latency.

The adaptive mode is where the sophisticated decision logic resides, enabling the engine to optimize for both speed and accuracy.

## How Adaptive Mode Makes Decisions

When `conversion.strategy` is set to `Adaptive`, Karukan evaluates three critical runtime inputs in the `determine_adaptive_strategy` function:

1. **Token count of the current reading** – measured using the main model's tokenizer
2. **Number of candidate suggestions requested** – one candidate indicates auto-suggest; more than one indicates explicit beam-search conversion
3. **Adaptive latency flag** – a boolean set to `true` when the most recent main-model conversion exceeded `EngineConfig.max_latency_ms`

### The Decision Function Implementation

The complete decision logic resides in [`karukan-im/src/core/engine/strategy.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/strategy.rs):

```rust
fn determine_adaptive_strategy(
    reading_tokens: usize,
    num_candidates: usize,
    has_light_model: bool,
    adaptive_use_light_model: bool,
    config: &EngineConfig,
) -> ConversionStrategy {
    if !has_light_model {
        return ConversionStrategy::MainModelOnly;
    }

    if num_candidates == 1 {
        // auto‑suggest
        if adaptive_use_light_model {
            ConversionStrategy::LightModelOnly
        } else {
            ConversionStrategy::MainModelOnly
        }
    } else {
        // user pressed Space
        if adaptive_use_light_model {
            ConversionStrategy::LightModelOnly
        } else if reading_tokens <= config.short_input_threshold {
            ConversionStrategy::ParallelBeam {
                beam_width: num_candidates.min(config.beam_width),
            }
        } else {
            ConversionStrategy::LightModelOnly
        }
    }
}

```

### Decision Logic for Auto-Suggest

For automatic suggestions where `num_candidates == 1`, the logic is straightforward:

- If `adaptive_use_light_model` is `true` (previous conversion was too slow), the engine returns `ConversionStrategy::LightModelOnly`
- Otherwise, it returns `ConversionStrategy::MainModelOnly`

### Decision Logic for Explicit Conversion

When the user presses Space to request multiple candidates (`num_candidates > 1`), the strategy becomes more nuanced:

- If the adaptive flag is `true`, the engine immediately selects `LightModelOnly` to avoid another slow conversion
- If the adaptive flag is `false` and `reading_tokens <= short_input_threshold`, the engine uses `ParallelBeam` to run both models simultaneously and beam-search the best results
- If the input exceeds the threshold (`reading_tokens > short_input_threshold`), the engine proactively selects `LightModelOnly` for long inputs

The `short_input_threshold` (default 40 tokens) and `beam_width` (default 9) are user-configurable values that control when parallel execution occurs.

## The Adaptive Latency Flag Mechanism

After each conversion, the engine calls `update_adaptive_model_flag` to adjust the adaptive state based on performance metrics:

```rust
match strategy {
    ConversionStrategy::MainModelOnly | ConversionStrategy::ParallelBeam { .. } => {
        self.metrics.adaptive_use_light_model =
            self.metrics.conversion_ms > self.config.max_latency_ms;
    }
    _ => {} // LightModelOnly or MainModelBeam do not affect the flag
}

```

When a conversion using the main model or parallel beam exceeds `max_latency_ms`, the flag is set to `true`. This flag is automatically cleared when a new word starts (when `reading_tokens` resets to 0), ensuring that latency adaptations are contextually relevant to the current input session.

## Code Examples

### Example 1: Auto-Suggest with Fast Main Model

```rust
// Settings: Adaptive mode, max_latency_ms = 200 ms
engine.metrics.adaptive_use_light_model = false; // last conversion was fast
let strategy = engine.determine_strategy("かんじ", 1);
// → ConversionStrategy::MainModelOnly

```

### Example 2: Auto-Suggest After Slow Conversion

```rust
engine.metrics.adaptive_use_light_model = true; // previous conversion took >200 ms
let strategy = engine.determine_strategy("かんじ", 1);
// → ConversionStrategy::LightModelOnly

```

### Example 3: Short Input with Space Key

```rust
// reading length = 12 tokens, short_input_threshold = 20
let strategy = engine.determine_strategy("にほんご", 5);
// → ConversionStrategy::ParallelBeam { beam_width: 5 }

```

### Example 4: Long Input with Space Key

```rust
// reading length = 45 tokens (exceeds short_input_threshold)
let strategy = engine.determine_strategy("おおさかこうえんであそんでいた", 5);
// → ConversionStrategy::LightModelOnly

```

## Configuration and Key Files

The adaptive model selection system spans several core files in the Karukan repository:

- **[`karukan-im/src/core/engine/strategy.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/strategy.rs)**: Implements the `determine_adaptive_strategy` function and the decision logic for all conversion strategies.
- **[`karukan-im/src/config/settings.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/config/settings.rs)**: Defines the `StrategyMode` enum and configuration fields including `max_latency_ms`, `short_input_threshold`, and `beam_width`.
- **[`karukan-im/src/core/engine/mod.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/mod.rs)**: Contains the `InputMethodEngine` struct that holds runtime metrics (`conversion_ms`, `adaptive_use_light_model`) and invokes the strategy determination.
- **[`karukan-engine/src/kanji/backend.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/kanji/backend.rs)**: Provides the token-counting functionality used to measure input length for strategy decisions.

## Summary

Karukan's **adaptive model selection** strategy intelligently balances accuracy and latency by:

- Monitoring conversion latency in real-time and setting an adaptive flag when the main model exceeds `max_latency_ms`
- Selecting the light model automatically for auto-suggest when previous conversions were slow
- Running both models in parallel via `ParallelBeam` for short explicit conversions under the threshold
- Proactively using the light model for long inputs that exceed `short_input_threshold`
- Resetting the adaptive state when input context changes (new word starts)

This approach ensures fast auto-suggestions on short inputs while maintaining low-latency fallback options for complex or lengthy conversions.

## Frequently Asked Questions

### What triggers Karukan to switch from the main model to the light model?

The switch occurs when the `adaptive_use_light_model` flag is set to `true`, which happens after a main-model conversion exceeds the user-configured `max_latency_ms`. Additionally, for long inputs in adaptive mode where `reading_tokens` exceeds `short_input_threshold`, the engine proactively selects the light model even without a latency violation.

### How does the short input threshold affect conversion strategy?

The `short_input_threshold` (default 40 tokens) determines whether Karukan runs both models in parallel. When a user requests explicit conversion (Space key) on a short input below this threshold, and the adaptive flag is false, the engine uses `ParallelBeam` to execute both models simultaneously and return the best candidate. Inputs above this threshold bypass parallel execution and use the light model directly.

### What is the difference between Light Mode and Adaptive Mode?

Light Mode forces exclusive use of the small model for all conversions, sacrificing accuracy for consistent speed. Adaptive Mode dynamically selects between the main and light models based on real-time input length, candidate count, and measured latency, providing high accuracy for fast queries and automatic fallback to the light model when performance degrades.

### How does Karukan measure conversion latency?

The engine records the duration of each conversion in milliseconds within the `InputMethodEngine` metrics struct. After each conversion using the main model or parallel beam strategy, it compares `metrics.conversion_ms` against `config.max_latency_ms`. If the conversion time exceeds this threshold, the `adaptive_use_light_model` flag is set to `true` for subsequent operations until the input context resets.