# How Karukan Handles Surrounding Text for Context-Aware Conversion with Left Context (lctx)

> Discover how Karukan leverages left context (lctx) to enhance context-aware text conversion. Learn how it builds and processes surrounding text for smarter language models.

- Repository: [Hitoshi Togasaki/karukan](https://github.com/togatoga/karukan)
- Tags: deep-dive
- Published: 2026-07-03

---

**Karukan builds a left context (lctx) for each conversion chunk by concatenating the editor's surrounding left text with the converted output of all preceding chunks, then truncating the combined string to `max_api_context_len` before passing it to the language model.**

The `togatoga/karukan` IME engine provides context-aware conversion by feeding the language model a carefully constructed **left context (lctx)** derived from two distinct sources: the text already present in the editor to the left of the cursor, and the converted results of earlier chunks in the current composing session. This architecture allows the model to understand narrative flow even when processing long documents incrementally.

## Understanding the Left Context (lctx) Architecture

Karukan processes user input in discrete chunks rather than sending the entire buffer to the model at once. For each chunk, the engine synthesizes a left context that represents exactly what the model would see if the whole document were being processed sequentially.

### What Makes Up the Left Context

The `lctx` for any given chunk consists of two concatenated components:

1. **Editor surrounding text on the left** – The text that already exists in the application before the IME's composing buffer begins.
2. **Converted text of all preceding chunks** – The model-generated output from earlier chunks within the same composing session.

These two strings are joined together and then truncated to respect the language model API's context window limits, defined by the `max_api_context_len` parameter.

### The Truncation Mechanism

Before any API call, the engine trims the context to fit within model constraints. In [`karukan-im/src/core/engine/display.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/display.rs), the `truncate_context_for_api()` method extracts the left side of the stored surrounding context and truncates it to `max_api_context_len` (lines 51‑60). This ensures that the model never receives more context than it can process, keeping latency predictable regardless of document length.

## Implementation in the Codebase

The surrounding text handling spans multiple files in the `karukan-im/src/core/engine/` directory, with clear separation between storage, processing, and display logic.

### Storing Surrounding Text from the Editor

When the frontend detects cursor movement or text changes, it notifies the engine via the JSON‑RPC method `set_surrounding_text`. This is defined in [`karukan-im/src/server/protocol.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/server/protocol.rs) and stores the data in the `SurroundingContext` struct (defined in [`types.rs`](https://github.com/togatoga/karukan/blob/main/types.rs)), which holds separate `left` and `right` strings:

```rust
// JSON-RPC method called by the frontend
fn set_surrounding_text(&mut self, left: String, right: String) {
    self.surrounding_context = Some(SurroundingContext { left, right });
}

```

The `left` field becomes the foundation for the `base_ctx` used in subsequent conversions.

### Building lctx for New Chunks

The core logic for constructing the left context resides in [`karukan-im/src/core/engine/chunk.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/chunk.rs). The `lctx_for()` method (lines 74‑81) concatenates the trimmed editor left context (`base`) with the already‑converted text of previous chunks (`preceding_converted`), then applies the same truncation rules used for the API limit.

When converting a new chunk, the `convert_new_chunk()` function (lines 60‑66) checks if the reading contains Japanese characters. If so, it calls `self.lctx_for(&base_ctx, &combined)` to generate the context-aware input:

```rust
fn convert_new_chunk(
    &mut self,
    reading: String,
    base_ctx: &str,          // trimmed editor left context
    combined: &str,          // converted text of previous chunks
) -> ComposingChunk {
    let converted = if reading.chars().next().is_some_and(is_japanese) {
        // Build the left-context for this chunk
        let lctx = self.lctx_for(base_ctx, combined);
        self.convert_chunk(&reading, &lctx)
    } else {
        // Non-Japanese runs are passed through verbatim
        reading.clone()
    };
    ComposingChunk { reading, converted }
}

```

### Displaying Context in the UI

To help users understand what context the model is using, the engine exposes the current chunk's left context through the auxiliary text line. In [`display.rs`](https://github.com/togatoga/karukan/blob/main/display.rs), the `display_context_chunked()` method (lines 16‑23) retrieves the lctx for the chunk containing the cursor via `chunk_lctx()`, then formats it using `context_line()`:

```rust
pub(super) fn display_context_chunked(&self) -> String {
    // lctx for the chunk the cursor is currently inside
    let lctx = self.chunk_lctx(self.current_chunk_index());
    let left = (!lctx.is_empty()).then_some(lctx.as_str());
    let right = self.surrounding_context
        .as_ref()
        .and_then(|c| c.right.as_deref());
    self.context_line(left, right)
}

```

This produces aux text like `⚡[あ] 今日は | ctx: lctx: …今日は | model-name`, giving users visibility into the conversion context.

## Incremental Re-chunking with Consistent Context

Karukan optimizes performance by only reconverting the specific portion of text that changed, rather than the entire buffer. When the user edits the middle of a sentence, `ChunkPlan::compute()` in [`chunk.rs`](https://github.com/togatoga/karukan/blob/main/chunk.rs) determines which existing chunks can be reused and which span must be re‑chunked.

Crucially, the **lctx** for each new chunk in the re‑computed span is still built using the same `base_ctx` (from the editor's surrounding text) plus the accumulated converted text of all preceding chunks. This ensures that even during incremental updates, the model receives a consistent view of the document history, maintaining context-aware conversion accuracy without requiring a full buffer re‑send.

## Summary

- **Dual-source context**: Karukan constructs `lctx` by combining editor surrounding text with converted preceding chunks.
- **Truncation safety**: `truncate_context_for_api()` ensures the context never exceeds `max_api_context_len`.
- **Chunk-level processing**: The `lctx_for()` method in [`chunk.rs`](https://github.com/togatoga/karukan/blob/main/chunk.rs) handles the concatenation logic for each conversion unit.
- **Incremental updates**: `ChunkPlan::compute()` enables efficient re‑chunking while preserving contextual consistency.
- **UI transparency**: `display_context_chunked()` exposes the current left context to users via auxiliary text lines.

## Frequently Asked Questions

### What is lctx in Karukan?

**Lctx (left context)** is the contextual string passed to the language model before each conversion chunk. It consists of the editor's left surrounding text concatenated with the converted output of all previous chunks in the current session, truncated to fit within API limits. This allows the model to generate contextually appropriate conversions based on document history.

### How does Karukan handle long documents without exceeding API limits?

Karukan uses **chunked conversion** combined with **context truncation**. The `truncate_context_for_api()` function in [`display.rs`](https://github.com/togatoga/karukan/blob/main/display.rs) limits the left context to `max_api_context_len` before any API call. By processing input in chunks and only sending the most relevant truncated context, the engine maintains flat latency even for very long documents.

### Where is the surrounding text stored in the codebase?

The surrounding text is stored in the `SurroundingContext` struct defined in [`karukan-im/src/core/engine/types.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/types.rs). This struct holds `left` and `right` strings and is populated via the `set_surrounding_text` JSON‑RPC method in [`karukan-im/src/server/protocol.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/server/protocol.rs). The engine accesses this data through `self.surrounding_context` during conversion operations.

### How does Karukan maintain context when editing the middle of a sentence?

When text is edited, `ChunkPlan::compute()` determines which chunks need reconversion. For any new chunks created during this incremental update, the engine still builds the `lctx` using the original `base_ctx` (editor surrounding text) plus the converted text of all preceding chunks. This ensures the model sees the same historical context it would have seen if converting the entire buffer from the start, preserving context-awareness during mid-sentence edits.