# How Karukan's Live Conversion Chunking Handles Japanese and Non-Japanese Text Boundaries

> Discover how Karukan's live conversion chunking accurately separates Japanese and non-Japanese text boundaries, preserving original characters for superior neural model processing.

- Repository: [Hitoshi Togasaki/karukan](https://github.com/togatoga/karukan)
- Tags: internals
- Published: 2026-07-03

---

**Karukan's live conversion chunking classifies every character as either Japanese (hiragana, katakana, or kanji) or non-Japanese, creating separate chunks that keep ASCII digits, symbols, and punctuation verbatim while isolating Japanese text for neural model processing.**

Karukan is an input method engine that processes real-time text composition through intelligent segmentation. The **live conversion chunking** system in the `togatoga/karukan` repository ensures that mixed-language input is handled efficiently by splitting the composing buffer into discrete units based on character type, preventing unnecessary neural model calls for non-convertible text.

## Binary Character Classification in `is_japanese`

At the core of the boundary detection is the `is_japanese` function defined in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs) (lines 47-57). This function performs a single, inexpensive Unicode range check to categorize each character.

The function recognizes three specific Japanese Unicode blocks:

- **Hiragana**: `\u{3040}` to `\u{309F}`
- **Katakana**: `\u{30A0}` to `\u{30FF}` (including the prolonged sound mark `ー` at U+30FC)
- **CJK Ideographs (Kanji)**: `\u{3400}` to `\u{9FFF}`

Crucially, the function explicitly excludes the katakana middle dot `・` (U+30FB) despite its presence in the katakana block:

```rust
fn is_japanese(c: char) -> bool {
    // 中黒 (・): a katakana‑block separator, treated as a non‑Japanese symbol.
    if c == '\u{30FB}' { return false; }
    matches!(c,
        '\u{3040}'..='\u{309F}'   // hiragana
        | '\u{30A0}'..='\u{30FF}' // katakana (incl. ー U+30FC)
        | '\u{3400}'..='\u{9FFF}' // CJK ideographs (kanji)
    )
}

```

## Chunk Boundary Algorithm with `group_chunks`

The `group_chunks` function (lines 66-84 in [`chunk.rs`](https://github.com/togatoga/karukan/blob/main/chunk.rs)) implements the boundary algorithm that walks the character vector and creates new chunks whenever either:

1. The **maximum chunk length** (`config.composing_chunk_len`) is reached, or
2. The **character group** (Japanese versus non-Japanese) changes.

A run of Japanese characters becomes one chunk (or several if it exceeds the length cap), while a run of non-Japanese characters becomes its own chunk and is passed through verbatim to the pre-edit without being sent to the model.

```rust
fn group_chunks(chars: &[char], max: usize) -> Vec<&[char]> {
    let mut out = Vec::new();
    let mut start = 0;
    while start < chars.len() {
        let limit = (start + max).min(chars.len());
        let japanese = is_japanese(chars[start]);
        let mut i = start;
        while i < limit && is_japanese(chars[i]) == japanese { i += 1; }
        out.push(&chars[start..i]);
        start = i;
    }
    out
}

```

## Handling Punctuation and Mixed Input

Because punctuation marks fall outside the Japanese Unicode ranges, they automatically trigger chunk boundaries. This means a sentence like `今日は。明日` produces three chunks: `["今日は", "。", "明日"]`. The period remains untouched by the neural converter, while the Japanese segments are processed separately.

The explicit exclusion of the middle dot `・` ensures that foreign names written in katakana, such as `ジョン・スミス`, split into three distinct chunks: Japanese, separator, Japanese. This prevents the neural model from attempting to convert the entire string as a single Japanese word.

## Incremental Re-chunking After Edits

When the user edits the composing buffer, Karukan avoids re-processing the entire text. The `ChunkPlan::compute` method (lines 87-107) calculates common prefix and suffix lengths between the old and new text. Unchanged leading and trailing chunks are reused, while only the modified middle span is re-chunked via `group_chunks`. This optimization appears in the `chunked_auto_suggest` function (lines 121-130 in [`chunk.rs`](https://github.com/togatoga/karukan/blob/main/chunk.rs)).

## Conversion Pipeline for New Chunks

When processing new chunks in `convert_new_chunk` (lines 48-66 in [`chunk.rs`](https://github.com/togatoga/karukan/blob/main/chunk.rs)), the engine checks the first character to determine the chunk type. If the chunk is Japanese, it is sent to the neural converter with the appropriate left context; otherwise, it is returned unchanged.

The following example demonstrates how to manually split a string using Karukan's logic:

```rust
// Example: manually split a string using the same logic as Karukan.
fn split_like_karukan(s: &str, max_len: usize) -> Vec<String> {
    let chars: Vec<char> = s.chars().collect();
    karukan_im::engine::group_chunks(&chars, max_len)
        .into_iter()
        .map(|c| c.iter().collect())
        .collect()
}

fn main() {
    let txt = "あ123い、スーパー・マーケット";
    let chunks = split_like_karukan(txt, 40);
    // => ["あ", "123", "い", "、", "スーパー", "・", "マーケット"]
    println!("{:?}", chunks);
}

```

## Summary

- **Binary classification**: Karukan uses a single `is_japanese` check based on Unicode ranges to separate text types, with explicit handling of the middle dot separator.
- **Automatic punctuation boundaries**: Non-Japanese characters (including punctuation and symbols) create natural chunk boundaries without special casing.
- **Length-constrained grouping**: The `group_chunks` function respects `config.composing_chunk_len` to ensure neural model calls remain bounded.
- **Incremental processing**: The `ChunkPlan::compute` algorithm minimizes re-conversion by reusing unchanged chunks after edits.
- **Verbatim pass-through**: Non-Japanese chunks bypass the neural model entirely, improving performance for mixed-language input.

## Frequently Asked Questions

### Why does Karukan treat the middle dot ・ as non-Japanese?

The middle dot (U+30FB) is explicitly excluded in the `is_japanese` function to act as a word separator. This ensures that foreign names written in katakana, such as `ジョン・スミス`, split into three chunks (`ジョン`, `・`, `スミス`) for better conversion accuracy, preventing the engine from treating the entire string as a single Japanese word.

### What happens to ASCII characters during live conversion?

ASCII characters are classified as non-Japanese and grouped into separate chunks. According to the `convert_new_chunk` implementation in [`chunk.rs`](https://github.com/togatoga/karukan/blob/main/chunk.rs), these chunks are passed through to the pre-edit buffer verbatim without being sent to the neural conversion model, preserving exact spelling while isolating Japanese segments for processing.

### How does Karukan handle long continuous Japanese text?

The `group_chunks` function enforces a maximum chunk length defined by `config.composing_chunk_len`. When a Japanese run exceeds this limit, it splits into multiple Japanese chunks at the length boundary, ensuring each neural model call remains bounded and responsive even for lengthy composing buffers.

### Where does the chunking logic reside in the codebase?

The core implementation lives in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs), which contains the `is_japanese` and `group_chunks` functions, the `ChunkPlan` struct for incremental updates (lines 87-107), and the `chunked_auto_suggest` method (lines 121-130) that coordinates the conversion process.