# How Preedit Text Rendering Works with Cursor Position Tracking in Karukan

> Understand preedit text rendering and cursor tracking in Karukan. Learn how Japanese text composition achieves precise caret placement using character-based indexing.

- Repository: [Hitoshi Togasaki/karukan](https://github.com/togatoga/karukan)
- Tags: internals
- Published: 2026-07-03

---

**Karukan calculates cursor position using character-based indexing within the `Preedit` struct, combining the input buffer and Romaji buffer to generate underlined preedit text with precise caret placement during Japanese text composition.**

Karukan is an open-source Japanese input method engine (IME) that manages real-time preedit text rendering while users convert Romaji to hiragana, katakana, or kanji. Understanding how preedit text rendering works with cursor position tracking in Karukan requires examining the interplay between the core data structures in [`karukan-im/src/core/preedit.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/preedit.rs) and the display logic in [`karukan-im/src/core/engine/display.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/display.rs). The implementation ensures accurate cursor positioning across multibyte Unicode characters through a character-offset system rather than byte indexing.

## Preedit Data Model and Caret Management

The foundation of preedit rendering resides in the `Preedit` struct defined in [`karukan-im/src/core/preedit.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/preedit.rs). This structure encapsulates the complete state of the composition string:

```rust
pub struct Preedit {
    text: String,                     // Full pre‑edit text
    caret: usize,                     // Caret position in characters
    attributes: Vec<PreeditAttribute>,// Styling (underline, highlight, …)
}

```

The `caret` field stores the cursor position as a **character count**, not a byte offset. This distinction is critical for Japanese text processing, where hiragana, katakana, and kanji require multiple bytes in UTF-8 encoding. The `Preedit::set_caret()` method clamps the supplied value to `self.len()`, preventing out-of-range cursor positions that could crash the UI or misplace the insertion point.

Styling attributes attach to specific text ranges via the `PreeditAttribute` vector. During active composition, Karukan applies a single underline attribute covering the entire string using `PreeditAttribute::underline(0, len)`, signaling to the frontend that the text is still being edited.

## Building the Display String and Calculating Caret Offsets

When the engine enters the composing state, `InputMethodEngine::build_composing_preedit()` in [`karukan-im/src/core/engine/display.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/display.rs) constructs the display representation. This method handles two distinct input paths and computes the appropriate caret offset through `display_caret_position()`.

### Live Conversion vs. Buffer Fallback

The display logic first determines the source text based on the current input state:

1. **Live conversion mode**: If `self.live.text` exists, the display string becomes `live.text + romaji_buffer`, with the caret positioned at `display.chars().count()`.
2. **Buffer fallback**: Without live conversion, the engine calls `build_input_display()` to concatenate the input buffer and Romaji buffer.

The `build_input_display()` implementation splits the input buffer at the cursor position, handles katakana conversion, and inserts the Romaji buffer:

```rust
let before = self.input_buf.text.chars().take(self.input_buf.cursor_pos).collect::<String>();
let after  = self.input_buf.text.chars().skip(self.input_buf.cursor_pos).collect::<String>();
let buffer = self.converters.romaji.buffer();

let display_before = if katakana { hiragana_to_katakana(&before) } else { before };
let display_after  = if katakana { hiragana_to_katakana(&after) } else { after };

format!("{}{}{}", display_before, buffer, display_after)

```

This approach inserts the Romaji buffer exactly at the cursor location, ensuring the visual caret appears immediately after the unconverted Latin characters.

### Character-Based Positioning for Multibyte Support

For the buffer fallback path, the caret offset calculation sums the input buffer cursor position with the Romaji buffer length: `self.input_buf.cursor_pos + self.converters.romaji.buffer().chars().count()`. By operating on character counts rather than byte indices, Karukan prevents cursor drift when displaying mixed scripts or converting between hiragana and katakana.

The final `Preedit` construction follows this pattern:

```rust
let mut preedit = Preedit::with_text(&display);
preedit.set_caret(caret);
preedit.set_attributes(vec![PreeditAttribute::underline(0, len)]);

```

## UI Integration and Protocol Transmission

The frontend—whether macOS or fcitx5—receives the preedit data via JSON-RPC as defined in [`karukan-im/src/server/engine_protocol.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/server/engine_protocol.rs). The protocol transmits the `text`, `caret` index, and `attributes` vector to the client, which translates these into platform-specific rendering commands.

Because the caret uses **character indices**, the UI can correctly position the cursor even when the preedit string contains multibyte characters. This ensures consistent behavior across different operating systems and display frameworks.

The following pseudocode demonstrates how the engine sends preedit data to the client:

```rust
// Inside the IME engine (composing state)
let preedit = engine.build_composing_preedit();

// Send to the client (pseudocode)
json_rpc.send("preedit", {
    "text": preedit.text(),
    "caret": preedit.caret(),
    "attributes": preedit.attributes().iter().map(|a| ({
        "start": a.start,
        "end": a.end,
        "type": match a.attr_type {
            AttributeType::Underline => "underline",
            AttributeType::UnderlineDouble => "underline_double",
            AttributeType::Highlight => "highlight",
            AttributeType::Reverse => "reverse",
        }
    })).collect()
});

```

Testing validates the caret positioning logic through assertions:

```rust
#[test]
fn preedit_caret_is_correct() {
    let mut engine = make_engine(); // set up a default engine
    engine.input_buf.text = "かな".into();
    engine.input_buf.cursor_pos = 1; // after the first kana
    engine.converters.romaji.push_buffer('d'); // user typed an 'd'
    let preedit = engine.build_composing_preedit();

    assert_eq!(preedit.text(), "かだな"); // "か" + romaji buffer "d" + "な"
    assert_eq!(preedit.caret(), 2);       // caret after the inserted 'd'
}

```

## Summary

- **Character-based indexing**: The `Preedit` struct stores the caret as a `usize` character count, not byte offset, ensuring correct handling of multibyte Japanese characters.
- **Dual rendering paths**: The engine chooses between live conversion text and the conventional input buffer, calculating caret position differently for each path in `display_caret_position()`.
- **Buffer insertion logic**: `build_input_display()` splits the input buffer at the cursor, converts segments to katakana when necessary, and inserts the Romaji buffer at the insertion point.
- **Safety clamping**: `set_caret()` prevents invalid cursor positions by clamping to the text length.
- **Protocol abstraction**: The JSON-RPC layer in [`engine_protocol.rs`](https://github.com/togatoga/karukan/blob/main/engine_protocol.rs) transmits character indices to frontends, enabling accurate cursor rendering across platforms.

## Frequently Asked Questions

### How does Karukan prevent cursor drift with multibyte Japanese characters?

Karukan stores the caret position as a character index (scalar value count) rather than a byte offset. The `Preedit` struct in [`karukan-im/src/core/preedit.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/preedit.rs) uses `usize` for the caret field, and calculations in [`display.rs`](https://github.com/togatoga/karukan/blob/main/display.rs) use `chars().count()` to determine offsets. This ensures that hiragana, katakana, and kanji characters—each multiple bytes in UTF-8—occupy a single cursor position, preventing the drift that would occur with byte-based indexing.

### What is the difference between live conversion and buffer-based preedit rendering?

**Live conversion** mode displays the already-converted `live.text` appended with the active Romaji buffer, placing the caret at the end of the combined string. **Buffer-based** rendering constructs the display from `build_input_display()`, which splits the input buffer at the cursor position, optionally converts segments to katakana, and inserts the Romaji buffer between the split segments. The latter path is used when live conversion is not active, and it requires summing the input buffer cursor position with the Romaji buffer length to determine the final caret offset.

### How does the underline attribute get applied to preedit text?

During composition, `build_composing_preedit()` creates a `PreeditAttribute::underline(0, len)` covering the entire text length. This attribute is stored in the `attributes` vector of the `Preedit` struct. When transmitted via JSON-RPC to the frontend, the UI layer translates this attribute into platform-specific underlining commands, providing the visual cue that the text is still being composed.

### Why does Karukan clamp the caret position to the text length?

The `Preedit::set_caret()` method clamps the provided value to `self.len()` (the character length of the text) to prevent out-of-range errors that could crash the UI frontend or place the cursor in invalid positions. This defensive programming ensures that even if internal calculations produce an offset beyond the current text boundary, the caret remains within the valid range of 0 to text length.