# How Incremental Rechunking with ChunkPlan Optimizes Live Conversion Performance in Karukan

> Discover how Karukans ChunkPlan optimizes live conversion performance with incremental rechunking. Minimize reconversion work and keep latency low for faster edits.

- Repository: [Hitoshi Togasaki/karukan](https://github.com/togatoga/karukan)
- Tags: performance
- Published: 2026-07-03

---

**ChunkPlan minimizes reconversion work by reusing unchanged leading and trailing chunks while recomputing only the edited middle section, keeping latency proportional to edit size rather than buffer length.**

Karukan is an open-source Japanese input method engine that delivers real-time text conversion using neural models. When users edit the composing buffer, the system must invalidate only the affected regions without reconverting the entire text from scratch. **Incremental rechunking** via the `ChunkPlan` data structure in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs) provides the deterministic logic to identify these minimal diff regions.

## Understanding Chunk Boundaries

Before incremental optimization can occur, Karukan must split the composing buffer into discrete chunks. The `group_chunks` function in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs) (lines 66-84) handles this segmentation.

A chunk boundary is triggered by two conditions:

- The character count reaches the configured maximum length (`config.composing_chunk_len`)
- The script type transitions between Japanese and non-Japanese characters (detected via `is_japanese`)

This approach respects linguistic boundaries while ensuring model inputs remain bounded and predictable.

## Computing the Edit Span with ChunkPlan

When the buffer content changes, `ChunkPlan::compute` calculates exactly which portions of the previous conversion remain valid. Implemented in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs) (lines 11-55), this method performs a bidirectional diff:

1. **Common prefix**: It compares `old_text` and `text` to find `cp`, the length of unchanged characters at the start.
2. **Common suffix**: It calculates `cs`, the length of unchanged characters at the end.
3. **Chunk alignment**: It maps these character offsets to `lead_count` (unchanged leading chunks) and `trail_count` (unchanged trailing chunks).

The resulting `mid_start..mid_end` range identifies the only region requiring fresh chunking and conversion.

### Generating a ChunkPlan

```rust
// Compute a plan for a new buffer.
// old_lens: lengths of previous chunks (in chars)
// old_text: concatenated old reading
// text: new reading as Vec<char>
let plan = ChunkPlan::compute(&old_lens, &old_text, &text, chunk_len);
println!(
    "reuse {} leading, {} trailing; recompute chars {}..{}",
    plan.lead_count, plan.trail_count, plan.mid_start, plan.mid_end
);

```

## Executing Incremental Rechunking

The `chunked_auto_suggest` function applies the plan computed by `ChunkPlan`. Located in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs) (lines 60-87), this function executes a three-phase reconstruction strategy:

**Reuse Leading Chunks**

The engine drains `plan.lead_count` chunks from the old buffer. These chunks retain their cached conversions and left context, requiring no model calls.

**Reconvert the Middle**

The engine slices the new text using `&text[plan.mid_start..plan.mid_end]` and passes it to `group_chunks`. Each resulting sub-chunk is converted via `convert_new_chunk`, with the accumulated text serving as left context for subsequent chunks.

**Append Trailing Chunks**

Finally, the engine appends the remaining `trail_count` chunks from the old buffer. While their absolute position shifts, their conversions remain valid until explicitly edited.

### Applying the Plan in Practice

```rust
// Inside chunked_auto_suggest – applying the plan
let mut chunks = Vec::new();
let mut combined = String::new();

// 1️⃣ Keep leading chunks unchanged
for chunk in old.drain(..plan.lead_count) {
    combined.push_str(&chunk.converted);
    chunks.push(chunk);
}

// 2️⃣ Re‑chunk the edited middle part and convert each new chunk
let middle = &text[plan.mid_start..plan.mid_end];
for sub in group_chunks(middle, chunk_len) {
    let reading: String = sub.iter().collect();
    let new = self.convert_new_chunk(reading, &base_ctx, &combined);
    combined.push_str(&new.converted);
    chunks.push(new);
}

// 3️⃣ Append trailing cached chunks
for chunk in old.drain(trail_start..) {
    combined.push_str(&chunk.converted);
    chunks.push(chunk);
}

```

## Performance Benefits and Trade-offs

**Incremental rechunking** ensures that conversion latency scales with the size of the edit, not the total buffer length. When typing at the end of a long composition, only the final chunk triggers reconversion. In-buffer insertions or deletions affect at most the intersecting chunks—typically one or two—leaving the rest of the buffer untouched.

This optimization keeps the input method responsive even for lengthy Japanese sentences, avoiding unnecessary neural model invocations. The system accepts a controlled trade-off: trailing chunks may experience left-context drift after middle edits, but they remain cached until the user modifies them directly. This prioritizes low latency over immediate context perfection, reconverting trailing chunks only when necessary.

## Summary

- **ChunkPlan** provides deterministic, unit-testable computation of minimal reconversion work in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs).
- The algorithm identifies common prefixes and suffixes to isolate the edited middle span, translating character offsets into reusable chunk counts.
- `chunked_auto_suggest` executes the plan by preserving leading chunks, converting the middle span, and appending trailing chunks.
- Performance scales with edit size rather than buffer length, enabling responsive live conversion for long inputs.
- The system tolerates temporary context drift in trailing chunks to maintain low latency, reconverting them only upon subsequent edits.

## Frequently Asked Questions

### What determines chunk boundaries in Karukan?

Chunk boundaries are determined by the `group_chunks` function in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs) (lines 66-84). A chunk ends when it reaches the configured maximum length (`config.composing_chunk_len`) or when the character type switches between Japanese and non-Japanese scripts. This ensures chunks respect linguistic boundaries while maintaining predictable sizes for the neural model.

### How does ChunkPlan handle mid-buffer edits?

`ChunkPlan::compute` compares the old and new text to find the longest common prefix and suffix. Chunks entirely within the unchanged prefix or suffix are marked for reuse, while the overlapping middle region is flagged for reconversion. This means inserting text in the middle only triggers reconversion for chunks that intersect the edit, not the entire buffer.

### Why does Karukan tolerate context drift in trailing chunks?

Trailing chunks are kept cached even after middle edits to avoid reconverting the entire suffix of the buffer on every keystroke. While their absolute position and left context technically change, they remain visually consistent until the user edits them directly. This trade-off prioritizes low latency and responsiveness over immediate context accuracy, reconverting trailing chunks only when they are next modified.

### Where is the ChunkPlan implementation located?

The `ChunkPlan` struct and its `compute` method are implemented in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs) (lines 11-55). The `chunked_auto_suggest` function that applies the plan resides in the same file (lines 60-87), alongside the `group_chunks` utility function (lines 66-84) that handles boundary detection.