How Preedit Text Rendering Works with Cursor Position Tracking in Karukan
Karukan calculates cursor position using character-based indexing within the Preedit struct, combining the input buffer and Romaji buffer to generate underlined preedit text with precise caret placement during Japanese text composition.
Karukan is an open-source Japanese input method engine (IME) that manages real-time preedit text rendering while users convert Romaji to hiragana, katakana, or kanji. Understanding how preedit text rendering works with cursor position tracking in Karukan requires examining the interplay between the core data structures in karukan-im/src/core/preedit.rs and the display logic in karukan-im/src/core/engine/display.rs. The implementation ensures accurate cursor positioning across multibyte Unicode characters through a character-offset system rather than byte indexing.
Preedit Data Model and Caret Management
The foundation of preedit rendering resides in the Preedit struct defined in karukan-im/src/core/preedit.rs. This structure encapsulates the complete state of the composition string:
pub struct Preedit {
text: String, // Full pre‑edit text
caret: usize, // Caret position in characters
attributes: Vec<PreeditAttribute>,// Styling (underline, highlight, …)
}
The caret field stores the cursor position as a character count, not a byte offset. This distinction is critical for Japanese text processing, where hiragana, katakana, and kanji require multiple bytes in UTF-8 encoding. The Preedit::set_caret() method clamps the supplied value to self.len(), preventing out-of-range cursor positions that could crash the UI or misplace the insertion point.
Styling attributes attach to specific text ranges via the PreeditAttribute vector. During active composition, Karukan applies a single underline attribute covering the entire string using PreeditAttribute::underline(0, len), signaling to the frontend that the text is still being edited.
Building the Display String and Calculating Caret Offsets
When the engine enters the composing state, InputMethodEngine::build_composing_preedit() in karukan-im/src/core/engine/display.rs constructs the display representation. This method handles two distinct input paths and computes the appropriate caret offset through display_caret_position().
Live Conversion vs. Buffer Fallback
The display logic first determines the source text based on the current input state:
- Live conversion mode: If
self.live.textexists, the display string becomeslive.text + romaji_buffer, with the caret positioned atdisplay.chars().count(). - Buffer fallback: Without live conversion, the engine calls
build_input_display()to concatenate the input buffer and Romaji buffer.
The build_input_display() implementation splits the input buffer at the cursor position, handles katakana conversion, and inserts the Romaji buffer:
let before = self.input_buf.text.chars().take(self.input_buf.cursor_pos).collect::<String>();
let after = self.input_buf.text.chars().skip(self.input_buf.cursor_pos).collect::<String>();
let buffer = self.converters.romaji.buffer();
let display_before = if katakana { hiragana_to_katakana(&before) } else { before };
let display_after = if katakana { hiragana_to_katakana(&after) } else { after };
format!("{}{}{}", display_before, buffer, display_after)
This approach inserts the Romaji buffer exactly at the cursor location, ensuring the visual caret appears immediately after the unconverted Latin characters.
Character-Based Positioning for Multibyte Support
For the buffer fallback path, the caret offset calculation sums the input buffer cursor position with the Romaji buffer length: self.input_buf.cursor_pos + self.converters.romaji.buffer().chars().count(). By operating on character counts rather than byte indices, Karukan prevents cursor drift when displaying mixed scripts or converting between hiragana and katakana.
The final Preedit construction follows this pattern:
let mut preedit = Preedit::with_text(&display);
preedit.set_caret(caret);
preedit.set_attributes(vec![PreeditAttribute::underline(0, len)]);
UI Integration and Protocol Transmission
The frontend—whether macOS or fcitx5—receives the preedit data via JSON-RPC as defined in karukan-im/src/server/engine_protocol.rs. The protocol transmits the text, caret index, and attributes vector to the client, which translates these into platform-specific rendering commands.
Because the caret uses character indices, the UI can correctly position the cursor even when the preedit string contains multibyte characters. This ensures consistent behavior across different operating systems and display frameworks.
The following pseudocode demonstrates how the engine sends preedit data to the client:
// Inside the IME engine (composing state)
let preedit = engine.build_composing_preedit();
// Send to the client (pseudocode)
json_rpc.send("preedit", {
"text": preedit.text(),
"caret": preedit.caret(),
"attributes": preedit.attributes().iter().map(|a| ({
"start": a.start,
"end": a.end,
"type": match a.attr_type {
AttributeType::Underline => "underline",
AttributeType::UnderlineDouble => "underline_double",
AttributeType::Highlight => "highlight",
AttributeType::Reverse => "reverse",
}
})).collect()
});
Testing validates the caret positioning logic through assertions:
#[test]
fn preedit_caret_is_correct() {
let mut engine = make_engine(); // set up a default engine
engine.input_buf.text = "かな".into();
engine.input_buf.cursor_pos = 1; // after the first kana
engine.converters.romaji.push_buffer('d'); // user typed an 'd'
let preedit = engine.build_composing_preedit();
assert_eq!(preedit.text(), "かだな"); // "か" + romaji buffer "d" + "な"
assert_eq!(preedit.caret(), 2); // caret after the inserted 'd'
}
Summary
- Character-based indexing: The
Preeditstruct stores the caret as ausizecharacter count, not byte offset, ensuring correct handling of multibyte Japanese characters. - Dual rendering paths: The engine chooses between live conversion text and the conventional input buffer, calculating caret position differently for each path in
display_caret_position(). - Buffer insertion logic:
build_input_display()splits the input buffer at the cursor, converts segments to katakana when necessary, and inserts the Romaji buffer at the insertion point. - Safety clamping:
set_caret()prevents invalid cursor positions by clamping to the text length. - Protocol abstraction: The JSON-RPC layer in
engine_protocol.rstransmits character indices to frontends, enabling accurate cursor rendering across platforms.
Frequently Asked Questions
How does Karukan prevent cursor drift with multibyte Japanese characters?
Karukan stores the caret position as a character index (scalar value count) rather than a byte offset. The Preedit struct in karukan-im/src/core/preedit.rs uses usize for the caret field, and calculations in display.rs use chars().count() to determine offsets. This ensures that hiragana, katakana, and kanji characters—each multiple bytes in UTF-8—occupy a single cursor position, preventing the drift that would occur with byte-based indexing.
What is the difference between live conversion and buffer-based preedit rendering?
Live conversion mode displays the already-converted live.text appended with the active Romaji buffer, placing the caret at the end of the combined string. Buffer-based rendering constructs the display from build_input_display(), which splits the input buffer at the cursor position, optionally converts segments to katakana, and inserts the Romaji buffer between the split segments. The latter path is used when live conversion is not active, and it requires summing the input buffer cursor position with the Romaji buffer length to determine the final caret offset.
How does the underline attribute get applied to preedit text?
During composition, build_composing_preedit() creates a PreeditAttribute::underline(0, len) covering the entire text length. This attribute is stored in the attributes vector of the Preedit struct. When transmitted via JSON-RPC to the frontend, the UI layer translates this attribute into platform-specific underlining commands, providing the visual cue that the text is still being composed.
Why does Karukan clamp the caret position to the text length?
The Preedit::set_caret() method clamps the provided value to self.len() (the character length of the text) to prevent out-of-range errors that could crash the UI frontend or place the cursor in invalid positions. This defensive programming ensures that even if internal calculations produce an offset beyond the current text boundary, the caret remains within the valid range of 0 to text length.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →