# How Karukan's Neural Kana-Kanji Conversion Integrates with llama.cpp for GGUF Inference

> Learn how Karukan integrates with llama.cpp using the llama-cpp-2 Rust bindings for efficient GGUF inference. Discover its three-component architecture and IME API for kanji conversion.

- Repository: [Hitoshi Togasaki/karukan](https://github.com/togatoga/karukan)
- Tags: deep-dive
- Published: 2026-07-03

---

**Karukan delegates neural kana-kanji conversion to llama.cpp through the `llama-cpp-2` Rust bindings, using a three-component architecture that wraps GGUF models, manages global backend initialization via `OnceLock`, and provides a high-level IME API for converting katakana readings to kanji candidates.**

Karukan is an open-source Japanese input method editor (IME) that leverages transformer-based language models for kana-to-kanji conversion. The project integrates with **llama.cpp** to execute optimized CPU inference on GGUF-formatted models, enabling efficient neural conversion without GPU requirements. This integration lives primarily within the `karukan-engine` crate and centers on three core abstractions that bridge the Rust application code with the underlying C++ inference engine.

## Core Architecture Components

The integration is built around three primary components that separate model management, inference execution, and IME-facing API concerns.

### LlamaCppModel

The **`LlamaCppModel`** struct in [`karukan-engine/src/kanji/llamacpp.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/kanji/llamacpp.rs) serves as the low-level wrapper around the llama.cpp backend. It manages a global `LLAMA_BACKEND` initialized via `OnceLock` to ensure the costly `LlamaBackend::init` operation runs only once per process. This component handles:

- GGUF model loading from local file paths
- Tokenization using external HuggingFace tokenizer JSON files
- Decoding with automatic special token stripping (`skip_special_tokens = true`)
- Multiple generation strategies including greedy decoding and beam search

### Backend

The **`Backend`** struct in [`karukan-engine/src/kanji/backend.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/kanji/backend.rs) handles the orchestration layer between model retrieval and instantiation. It downloads GGUF models and their accompanying [`tokenizer.json`](https://github.com/togatoga/karukan/blob/main/tokenizer.json) files from HuggingFace using the `hf_download` module, then constructs an `LlamaCppModel` instance. The `Backend::from_variant_id` method provides a convenient registry-based interface for loading specific model variants (e.g., `"jinen-v1-small-q5"`).

### KanaKanjiConverter

The **`KanaKanjiConverter`** provides the high-level API consumed by the IME engine. Implemented in [`backend.rs`](https://github.com/togatoga/karukan/blob/main/backend.rs), it orchestrates the complete conversion pipeline: building *jinen*-style prompts, converting hiragana input to katakana (as expected by the model), running inference via `LlamaCppModel`, and post-processing outputs to produce clean kanji candidates.

## Inference Pipeline and Data Flow

Understanding the conversion process requires tracing the data flow from user input to kanji output.

### Model Acquisition and Initialization

When `Backend::from_variant_id` is called, the system invokes `hf_download::get_variant_path` and `get_tokenizer_path` to fetch the GGUF model and tokenizer from HuggingFace if not already cached locally. The `LlamaCppModel::from_file` method then loads the model with CPU-only parameters (`with_n_gpu_layers(0)`), creating a `llama_cpp_2::context::LlamaContext` configured with user-specified context window size (`n_ctx`) and thread count (`n_threads`).

### Prompt Construction with Special Tokens

The converter builds a *jinen*-style prompt using private-use-area tokens defined in the kanji module:

- `CONTEXT_TOKEN` (U+E000)
- `INPUT_START_TOKEN` (U+E001)  
- `OUTPUT_START_TOKEN` (U+E002)

The `build_jinen_prompt` function concatenates these tokens with left-hand context and the katakana reading. Since the model expects **katakana** input, `hiragana_to_katakana` converts the user's hiragana reading before embedding it between `INPUT_START_TOKEN` and `OUTPUT_START_TOKEN`.

### External Tokenization Strategy

Critically, `LlamaCppModel::tokenize` uses the external tokenizer loaded from [`tokenizer.json`](https://github.com/togatoga/karukan/blob/main/tokenizer.json) rather than llama.cpp's built-in tokenizer. This design choice prevents mishandling of the private-use-area tokens that define the *jinen* format, ensuring correct token boundaries for the transformer model.

### Generation and Decoding

The system selects inference strategies based on the number of candidates requested:

- **Single candidate**: `LlamaCppModel::generate` uses greedy decoding
- **Multiple candidates**: `LlamaCppModel::generate_beam_search` performs full beam search, or uses depth-1-beam plus greedy variants for faster results

The generation routine produces a `Vec<LlamaToken>` that `LlamaCppModel::decode` converts back to text. The `clean_model_output` function then trims whitespace to yield the final kanji string(s).

## Implementation Examples

### Basic Single-Candidate Conversion

The most common usage pattern loads a default variant and converts hiragana to a single kanji candidate:

```rust
use karukan_engine::kanji::{Backend, KanaKanjiConverter};

fn main() -> anyhow::Result<()> {
    // Load the default "small-q5" variant from the registry
    let backend = Backend::from_variant_id("jinen-v1-small-q5")?;
    let converter = KanaKanjiConverter::new(backend)?;

    // Convert the hiragana reading "かんじ" with no surrounding context
    let candidates = converter.convert("かんじ", "", 1)?;
    println!("Best candidate: {}", candidates[0]);

    Ok(())
}

```

This example demonstrates the `Backend::from_variant_id` → `KanaKanjiConverter::new` → `convert` workflow defined in [`karukan-engine/src/kanji/backend.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/kanji/backend.rs).

### Multi-Candidate Beam Search

To generate multiple conversion candidates for user selection:

```rust
let candidates = converter.convert("ひらがな", "", 5)?;
for (i, cand) in candidates.iter().enumerate() {
    println!("Candidate {}: {}", i + 1, cand);
}

```

When `num_candidates > 1`, `KanaKanjiConverter` automatically switches to `LlamaCppModel::generate_beam_search`, which returns a list of `(token_vec, score)` pairs sorted by cumulative log-probability.

### Direct Low-Level Model Access

For custom inference pipelines, you can interact directly with `LlamaCppModel`:

```rust
use karukan_engine::kanji::llamacpp::LlamaCppModel;

// Load a model file and its tokenizer
let model = LlamaCppModel::from_file("models/jp_small.gguf", "tokenizer.json")?;

// Build a prompt manually (jinen format)
let prompt = format!("{}{}{}{}{}",
    karukan_engine::kanji::CONTEXT_TOKEN,
    "",                      // no left context
    karukan_engine::kanji::INPUT_START_TOKEN,
    "カンジ",                 // katakana reading
    karukan_engine::kanji::OUTPUT_START_TOKEN,
);

// Tokenise, generate 50 new tokens, and decode
let input_tokens = model.tokenize(&prompt)?;
let generated = model.generate(&input_tokens, 50, Some(model.eos_token_id().0))?;
let output = model.decode(&generated[input_tokens.len()..], true)?;
println!("Generated: {}", output);

```

This low-level API exposes the full llama.cpp integration, including custom generation parameters and direct token manipulation.

## Key Source Files

The integration spans several modules in the `karukan-engine` crate:

- **[`karukan-engine/src/kanji/llamacpp.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/kanji/llamacpp.rs)** – Core wrapper around llama.cpp; handles backend initialization, tokenization, decoding, and all generation strategies (greedy, beam search, and depth-1 variants).
- **[`karukan-engine/src/kanji/backend.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/kanji/backend.rs)** – Contains both the `Backend` struct for model lifecycle management and the `KanaKanjiConverter` implementation that the IME calls.
- **[`karukan-engine/src/kanji/model_config.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/kanji/model_config.rs)** – Defines the model registry (`registry()`) used by `Backend::from_variant_id` to resolve variant IDs to HuggingFace repositories.
- **[`karukan-engine/src/kanji/hf_download.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/kanji/hf_download.rs)** – Handles on-demand downloading and caching of GGUF files and external tokenizers from HuggingFace.
- **[`karukan-engine/src/kanji/mod.rs`](https://github.com/togatoga/karukan/blob/main/karukan-engine/src/kanji/mod.rs)** – Public re-exports including `LlamaCppModel` and `NllScorer`.

## Summary

- **Karukan integrates with llama.cpp** through the `llama-cpp-2` Rust bindings, enabling optimized CPU inference on GGUF models without GPU dependencies.
- **Three-component architecture** separates concerns: `LlamaCppModel` manages the inference backend, `Backend` handles model acquisition, and `KanaKanjiConverter` provides the IME-facing API.
- **External tokenization** is required to correctly handle the private-use-area tokens (`CONTEXT_TOKEN`, `INPUT_START_TOKEN`, `OUTPUT_START_TOKEN`) that define the *jinen* prompt format.
- **Global backend initialization** uses `OnceLock` to ensure `LlamaBackend::init` runs only once per process, eliminating redundant initialization overhead.
- **Multiple generation strategies** support both single-candidate greedy decoding and multi-candidate beam search for flexible conversion behavior.

## Frequently Asked Questions

### Why does Karukan use an external tokenizer instead of llama.cpp's built-in tokenizer?

Karukan uses an external HuggingFace tokenizer loaded from [`tokenizer.json`](https://github.com/togatoga/karukan/blob/main/tokenizer.json) because llama.cpp's built-in tokenizer mishandles the private-use-area tokens (U+E000–U+E002) that delimit the *jinen* prompt format. The external tokenizer correctly segments these special tokens, ensuring the model receives the proper input structure for kana-kanji conversion.

### What generation strategies does Karukan support for candidate ranking?

The `LlamaCppModel` implementation supports three generation strategies: **greedy decoding** for single-candidate conversion, **full beam search** for high-quality multi-candidate generation, and **depth-1-beam plus greedy** for faster inference when multiple candidates are needed. The `KanaKanjiConverter` automatically selects the appropriate strategy based on the requested number of candidates.

### How does Karukan handle model downloading and versioning?

The `Backend` struct uses the `hf_download` module to fetch GGUF models and tokenizers from HuggingFace repositories. Models are referenced by variant IDs (e.g., `"jinen-v1-small-q5"`) defined in [`model_config.rs`](https://github.com/togatoga/karukan/blob/main/model_config.rs), which maps to specific HuggingFace repositories. Downloaded files are cached locally, and the system checks for existing files before initiating network requests.

### Can Karukan run on CPU-only systems?

Yes, Karukan is designed specifically for CPU-only inference. The `LlamaCppModel` initializes with `with_n_gpu_layers(0)`, forcing all computation to run on the CPU. This design choice, combined with llama.cpp's optimized GGUF interpreter, enables efficient neural kana-kanji conversion on modest hardware without requiring GPU acceleration.