How Karukan's Live Conversion Chunking Handles Japanese and Non-Japanese Text Boundaries
Karukan's live conversion chunking classifies every character as either Japanese (hiragana, katakana, or kanji) or non-Japanese, creating separate chunks that keep ASCII digits, symbols, and punctuation verbatim while isolating Japanese text for neural model processing.
Karukan is an input method engine that processes real-time text composition through intelligent segmentation. The live conversion chunking system in the togatoga/karukan repository ensures that mixed-language input is handled efficiently by splitting the composing buffer into discrete units based on character type, preventing unnecessary neural model calls for non-convertible text.
Binary Character Classification in is_japanese
At the core of the boundary detection is the is_japanese function defined in karukan-im/src/core/engine/chunk.rs (lines 47-57). This function performs a single, inexpensive Unicode range check to categorize each character.
The function recognizes three specific Japanese Unicode blocks:
- Hiragana:
\u{3040}to\u{309F} - Katakana:
\u{30A0}to\u{30FF}(including the prolonged sound markーat U+30FC) - CJK Ideographs (Kanji):
\u{3400}to\u{9FFF}
Crucially, the function explicitly excludes the katakana middle dot ・ (U+30FB) despite its presence in the katakana block:
fn is_japanese(c: char) -> bool {
// 中黒 (・): a katakana‑block separator, treated as a non‑Japanese symbol.
if c == '\u{30FB}' { return false; }
matches!(c,
'\u{3040}'..='\u{309F}' // hiragana
| '\u{30A0}'..='\u{30FF}' // katakana (incl. ー U+30FC)
| '\u{3400}'..='\u{9FFF}' // CJK ideographs (kanji)
)
}
Chunk Boundary Algorithm with group_chunks
The group_chunks function (lines 66-84 in chunk.rs) implements the boundary algorithm that walks the character vector and creates new chunks whenever either:
- The maximum chunk length (
config.composing_chunk_len) is reached, or - The character group (Japanese versus non-Japanese) changes.
A run of Japanese characters becomes one chunk (or several if it exceeds the length cap), while a run of non-Japanese characters becomes its own chunk and is passed through verbatim to the pre-edit without being sent to the model.
fn group_chunks(chars: &[char], max: usize) -> Vec<&[char]> {
let mut out = Vec::new();
let mut start = 0;
while start < chars.len() {
let limit = (start + max).min(chars.len());
let japanese = is_japanese(chars[start]);
let mut i = start;
while i < limit && is_japanese(chars[i]) == japanese { i += 1; }
out.push(&chars[start..i]);
start = i;
}
out
}
Handling Punctuation and Mixed Input
Because punctuation marks fall outside the Japanese Unicode ranges, they automatically trigger chunk boundaries. This means a sentence like 今日は。明日 produces three chunks: ["今日は", "。", "明日"]. The period remains untouched by the neural converter, while the Japanese segments are processed separately.
The explicit exclusion of the middle dot ・ ensures that foreign names written in katakana, such as ジョン・スミス, split into three distinct chunks: Japanese, separator, Japanese. This prevents the neural model from attempting to convert the entire string as a single Japanese word.
Incremental Re-chunking After Edits
When the user edits the composing buffer, Karukan avoids re-processing the entire text. The ChunkPlan::compute method (lines 87-107) calculates common prefix and suffix lengths between the old and new text. Unchanged leading and trailing chunks are reused, while only the modified middle span is re-chunked via group_chunks. This optimization appears in the chunked_auto_suggest function (lines 121-130 in chunk.rs).
Conversion Pipeline for New Chunks
When processing new chunks in convert_new_chunk (lines 48-66 in chunk.rs), the engine checks the first character to determine the chunk type. If the chunk is Japanese, it is sent to the neural converter with the appropriate left context; otherwise, it is returned unchanged.
The following example demonstrates how to manually split a string using Karukan's logic:
// Example: manually split a string using the same logic as Karukan.
fn split_like_karukan(s: &str, max_len: usize) -> Vec<String> {
let chars: Vec<char> = s.chars().collect();
karukan_im::engine::group_chunks(&chars, max_len)
.into_iter()
.map(|c| c.iter().collect())
.collect()
}
fn main() {
let txt = "あ123い、スーパー・マーケット";
let chunks = split_like_karukan(txt, 40);
// => ["あ", "123", "い", "、", "スーパー", "・", "マーケット"]
println!("{:?}", chunks);
}
Summary
- Binary classification: Karukan uses a single
is_japanesecheck based on Unicode ranges to separate text types, with explicit handling of the middle dot separator. - Automatic punctuation boundaries: Non-Japanese characters (including punctuation and symbols) create natural chunk boundaries without special casing.
- Length-constrained grouping: The
group_chunksfunction respectsconfig.composing_chunk_lento ensure neural model calls remain bounded. - Incremental processing: The
ChunkPlan::computealgorithm minimizes re-conversion by reusing unchanged chunks after edits. - Verbatim pass-through: Non-Japanese chunks bypass the neural model entirely, improving performance for mixed-language input.
Frequently Asked Questions
Why does Karukan treat the middle dot ・ as non-Japanese?
The middle dot (U+30FB) is explicitly excluded in the is_japanese function to act as a word separator. This ensures that foreign names written in katakana, such as ジョン・スミス, split into three chunks (ジョン, ・, スミス) for better conversion accuracy, preventing the engine from treating the entire string as a single Japanese word.
What happens to ASCII characters during live conversion?
ASCII characters are classified as non-Japanese and grouped into separate chunks. According to the convert_new_chunk implementation in chunk.rs, these chunks are passed through to the pre-edit buffer verbatim without being sent to the neural conversion model, preserving exact spelling while isolating Japanese segments for processing.
How does Karukan handle long continuous Japanese text?
The group_chunks function enforces a maximum chunk length defined by config.composing_chunk_len. When a Japanese run exceeds this limit, it splits into multiple Japanese chunks at the length boundary, ensuring each neural model call remains bounded and responsive even for lengthy composing buffers.
Where does the chunking logic reside in the codebase?
The core implementation lives in karukan-im/src/core/engine/chunk.rs, which contains the is_japanese and group_chunks functions, the ChunkPlan struct for incremental updates (lines 87-107), and the chunked_auto_suggest method (lines 121-130) that coordinates the conversion process.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →