How Karukan's Neural Kana-Kanji Conversion Integrates with llama.cpp for GGUF Inference

Karukan delegates neural kana-kanji conversion to llama.cpp through the llama-cpp-2 Rust bindings, using a three-component architecture that wraps GGUF models, manages global backend initialization via OnceLock, and provides a high-level IME API for converting katakana readings to kanji candidates.

Karukan is an open-source Japanese input method editor (IME) that leverages transformer-based language models for kana-to-kanji conversion. The project integrates with llama.cpp to execute optimized CPU inference on GGUF-formatted models, enabling efficient neural conversion without GPU requirements. This integration lives primarily within the karukan-engine crate and centers on three core abstractions that bridge the Rust application code with the underlying C++ inference engine.

Core Architecture Components

The integration is built around three primary components that separate model management, inference execution, and IME-facing API concerns.

LlamaCppModel

The LlamaCppModel struct in karukan-engine/src/kanji/llamacpp.rs serves as the low-level wrapper around the llama.cpp backend. It manages a global LLAMA_BACKEND initialized via OnceLock to ensure the costly LlamaBackend::init operation runs only once per process. This component handles:

  • GGUF model loading from local file paths
  • Tokenization using external HuggingFace tokenizer JSON files
  • Decoding with automatic special token stripping (skip_special_tokens = true)
  • Multiple generation strategies including greedy decoding and beam search

Backend

The Backend struct in karukan-engine/src/kanji/backend.rs handles the orchestration layer between model retrieval and instantiation. It downloads GGUF models and their accompanying tokenizer.json files from HuggingFace using the hf_download module, then constructs an LlamaCppModel instance. The Backend::from_variant_id method provides a convenient registry-based interface for loading specific model variants (e.g., "jinen-v1-small-q5").

KanaKanjiConverter

The KanaKanjiConverter provides the high-level API consumed by the IME engine. Implemented in backend.rs, it orchestrates the complete conversion pipeline: building jinen-style prompts, converting hiragana input to katakana (as expected by the model), running inference via LlamaCppModel, and post-processing outputs to produce clean kanji candidates.

Inference Pipeline and Data Flow

Understanding the conversion process requires tracing the data flow from user input to kanji output.

Model Acquisition and Initialization

When Backend::from_variant_id is called, the system invokes hf_download::get_variant_path and get_tokenizer_path to fetch the GGUF model and tokenizer from HuggingFace if not already cached locally. The LlamaCppModel::from_file method then loads the model with CPU-only parameters (with_n_gpu_layers(0)), creating a llama_cpp_2::context::LlamaContext configured with user-specified context window size (n_ctx) and thread count (n_threads).

Prompt Construction with Special Tokens

The converter builds a jinen-style prompt using private-use-area tokens defined in the kanji module:

  • CONTEXT_TOKEN (U+E000)
  • INPUT_START_TOKEN (U+E001)
  • OUTPUT_START_TOKEN (U+E002)

The build_jinen_prompt function concatenates these tokens with left-hand context and the katakana reading. Since the model expects katakana input, hiragana_to_katakana converts the user's hiragana reading before embedding it between INPUT_START_TOKEN and OUTPUT_START_TOKEN.

External Tokenization Strategy

Critically, LlamaCppModel::tokenize uses the external tokenizer loaded from tokenizer.json rather than llama.cpp's built-in tokenizer. This design choice prevents mishandling of the private-use-area tokens that define the jinen format, ensuring correct token boundaries for the transformer model.

Generation and Decoding

The system selects inference strategies based on the number of candidates requested:

  • Single candidate: LlamaCppModel::generate uses greedy decoding
  • Multiple candidates: LlamaCppModel::generate_beam_search performs full beam search, or uses depth-1-beam plus greedy variants for faster results

The generation routine produces a Vec<LlamaToken> that LlamaCppModel::decode converts back to text. The clean_model_output function then trims whitespace to yield the final kanji string(s).

Implementation Examples

Basic Single-Candidate Conversion

The most common usage pattern loads a default variant and converts hiragana to a single kanji candidate:

use karukan_engine::kanji::{Backend, KanaKanjiConverter};

fn main() -> anyhow::Result<()> {
    // Load the default "small-q5" variant from the registry
    let backend = Backend::from_variant_id("jinen-v1-small-q5")?;
    let converter = KanaKanjiConverter::new(backend)?;

    // Convert the hiragana reading "かんじ" with no surrounding context
    let candidates = converter.convert("かんじ", "", 1)?;
    println!("Best candidate: {}", candidates[0]);

    Ok(())
}

This example demonstrates the Backend::from_variant_id → KanaKanjiConverter::new → convert workflow defined in karukan-engine/src/kanji/backend.rs.

To generate multiple conversion candidates for user selection:

let candidates = converter.convert("ひらがな", "", 5)?;
for (i, cand) in candidates.iter().enumerate() {
    println!("Candidate {}: {}", i + 1, cand);
}

When num_candidates > 1, KanaKanjiConverter automatically switches to LlamaCppModel::generate_beam_search, which returns a list of (token_vec, score) pairs sorted by cumulative log-probability.

Direct Low-Level Model Access

For custom inference pipelines, you can interact directly with LlamaCppModel:

use karukan_engine::kanji::llamacpp::LlamaCppModel;

// Load a model file and its tokenizer
let model = LlamaCppModel::from_file("models/jp_small.gguf", "tokenizer.json")?;

// Build a prompt manually (jinen format)
let prompt = format!("{}{}{}{}{}",
    karukan_engine::kanji::CONTEXT_TOKEN,
    "",                      // no left context
    karukan_engine::kanji::INPUT_START_TOKEN,
    "カンジ",                 // katakana reading
    karukan_engine::kanji::OUTPUT_START_TOKEN,
);

// Tokenise, generate 50 new tokens, and decode
let input_tokens = model.tokenize(&prompt)?;
let generated = model.generate(&input_tokens, 50, Some(model.eos_token_id().0))?;
let output = model.decode(&generated[input_tokens.len()..], true)?;
println!("Generated: {}", output);

This low-level API exposes the full llama.cpp integration, including custom generation parameters and direct token manipulation.

Key Source Files

The integration spans several modules in the karukan-engine crate:

Summary

  • Karukan integrates with llama.cpp through the llama-cpp-2 Rust bindings, enabling optimized CPU inference on GGUF models without GPU dependencies.
  • Three-component architecture separates concerns: LlamaCppModel manages the inference backend, Backend handles model acquisition, and KanaKanjiConverter provides the IME-facing API.
  • External tokenization is required to correctly handle the private-use-area tokens (CONTEXT_TOKEN, INPUT_START_TOKEN, OUTPUT_START_TOKEN) that define the jinen prompt format.
  • Global backend initialization uses OnceLock to ensure LlamaBackend::init runs only once per process, eliminating redundant initialization overhead.
  • Multiple generation strategies support both single-candidate greedy decoding and multi-candidate beam search for flexible conversion behavior.

Frequently Asked Questions

Why does Karukan use an external tokenizer instead of llama.cpp's built-in tokenizer?

Karukan uses an external HuggingFace tokenizer loaded from tokenizer.json because llama.cpp's built-in tokenizer mishandles the private-use-area tokens (U+E000–U+E002) that delimit the jinen prompt format. The external tokenizer correctly segments these special tokens, ensuring the model receives the proper input structure for kana-kanji conversion.

What generation strategies does Karukan support for candidate ranking?

The LlamaCppModel implementation supports three generation strategies: greedy decoding for single-candidate conversion, full beam search for high-quality multi-candidate generation, and depth-1-beam plus greedy for faster inference when multiple candidates are needed. The KanaKanjiConverter automatically selects the appropriate strategy based on the requested number of candidates.

How does Karukan handle model downloading and versioning?

The Backend struct uses the hf_download module to fetch GGUF models and tokenizers from HuggingFace repositories. Models are referenced by variant IDs (e.g., "jinen-v1-small-q5") defined in model_config.rs, which maps to specific HuggingFace repositories. Downloaded files are cached locally, and the system checks for existing files before initiating network requests.

Can Karukan run on CPU-only systems?

Yes, Karukan is designed specifically for CPU-only inference. The LlamaCppModel initializes with with_n_gpu_layers(0), forcing all computation to run on the CPU. This design choice, combined with llama.cpp's optimized GGUF interpreter, enables efficient neural kana-kanji conversion on modest hardware without requiring GPU acceleration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →