Understanding the Jinen Format in Karukan: Unicode Tokens for LLM-Based Kanji Conversion

The Jinen format is a specialized plain-text prompt structure that enables Karukan’s kanji-conversion engine to communicate with llama.cpp models using three private-use-area Unicode characters (U+EE00, U+EE01, U+EE02) to delimit context, input reading, and output generation boundaries.

Karukan is an open-source kanji conversion engine that leverages large language models through llama.cpp for accurate Japanese text input. At the core of its inference pipeline lies the Jinen format, a custom prompt protocol developed in the togatoga/karukan repository that uses private-use-area Unicode tokens to ensure precise boundary detection between already-converted context, phonetic reading input, and generated kanji output. This design prevents the model from confusing finalized text with the katakana input requiring conversion.

What Is the Jinen Format?

The Jinen format is implemented as a concatenated string protocol that sandwiches user input between three distinct control tokens. According to the source code in karukan-engine/src/kanji/mod.rs, these tokens are defined as Rust character constants:

pub const CONTEXT_TOKEN: char = '\u{ee02}';
pub const INPUT_START_TOKEN: char = '\u{ee00}';
pub const OUTPUT_START_TOKEN: char = '\u{ee01}';

These Unicode characters occupy the private-use area (PUA), ensuring they never appear in standard Japanese text and remain unambiguous as structural delimiters throughout tokenization and decoding.

The Three Special Unicode Tokens (U+EE00–U+EE02)

Each token serves a specific semantic purpose in prompt construction, creating a linear sequence that guides the model’s attention during inference.

Context Token (U+EE02)

The CONTEXT_TOKEN (\u{EE02}) marks the beginning of the left-hand context—text that has already been converted and accepted by the user. This informs the model about the surrounding linguistic environment while preventing it from modifying or regenerating previously finalized text.

Input Start Token (U+EE00)

The INPUT_START_TOKEN (\u{EE00}) designates the exact boundary where the reading input begins. In Karukan, user hiragana is first converted to katakana via hiragana_to_katakana, then placed immediately after this token. This signals to the model that the following characters represent the phonetic representation requiring kanji conversion.

Output Start Token (U+EE01)

The OUTPUT_START_TOKEN (\u{EE01}) indicates the precise position where the model should begin generating kanji candidates. During inference, the model generates tokens immediately following this marker, and the decoded text is interpreted as the conversion result. The clean_model_output function then strips surrounding whitespace to return clean candidates.

Building the Jinen Prompt in Rust

The helper function build_jinen_prompt in karukan-engine/src/kanji/backend.rs constructs the final prompt string by concatenating the tokens and text segments in the specific order context → input-start → reading → output-start:

pub fn build_jinen_prompt(katakana: &str, context: &str) -> String {
    format!(
        "{}{}{}{}{}",
        CONTEXT_TOKEN, context, INPUT_START_TOKEN, katakana, OUTPUT_START_TOKEN
    )
}

For example, if the user has already typed "今日は" and the current reading is "コンニチハ" (the katakana representation of "こんにちは"), the resulting Jinen prompt becomes:


\u{EE02}今日は\u{EE00}コンニチハ\u{EE01}

The model receives the context token, the existing Japanese text, the input-start token, the katakana reading, and finally the output-start token, creating a clear directive for generation.

How Karukan Uses Jinen for Model Inference

The conversion flow implemented in KanaKanjiConverter::convert (around lines 21‑31 of backend.rs) follows a four-step pipeline:

  1. Convert hiragana to katakana using hiragana_to_katakana to standardize the phonetic input into the katakana script expected by the model.
  2. Build the Jinen prompt via build_jinen_prompt, incorporating optional left context from previously converted text.
  3. Tokenize and infer by feeding the prompt to the LlamaCpp backend defined in karukan-engine/src/kanji/llamacpp.rs. The special Unicode tokens are preserved through tokenization and decoding, ensuring the model can reliably locate boundaries after token-level transformations.
  4. Clean the output using clean_model_output to strip whitespace and return the final kanji candidates to the user interface.

This architecture ensures that the three special tokens survive the entire inference pipeline, preventing context bleed and maintaining strict separation between input, output, and historical context.

Practical Code Examples

Creating a Jinen Prompt Manually

You can construct the prompt manually using the exported constants from the Karukan engine:

use karukan_engine::kanji::{CONTEXT_TOKEN, INPUT_START_TOKEN, OUTPUT_START_TOKEN};

fn jinen_example() {
    let context = "今日は";
    let reading_katakana = "コンニチハ";

    let prompt = format!(
        "{}{}{}{}{}",
        CONTEXT_TOKEN, context,
        INPUT_START_TOKEN, reading_katakana,
        OUTPUT_START_TOKEN
    );

    // Prompt is: "\u{EE02}今日は\u{EE00}コンニチハ\u{EE01}"
    println!("Jinen prompt: {}", prompt);
}

Running Conversion with the Built-in API

For production use, the KanaKanjiConverter abstracts the Jinen format construction:

use karukan_engine::kanji::{Backend, KanaKanjiConverter};

fn convert_example() {
    // Load the default Jinen model variant
    let backend = Backend::from_variant_id("jinen-v1-small-q5")
        .expect("Failed to load model");

    // Create the converter
    let converter = KanaKanjiConverter::new(backend)
        .expect("Failed to create converter");

    // Hiragana reading and empty left context
    let reading = "かんじ";               // user-typed hiragana
    let context = "";                    // no prior converted text

    // Generate one candidate
    let candidates = converter.convert(reading, context, 1)
        .expect("Conversion failed");

    println!("Kanji candidates: {:?}", candidates);
}

Verifying Token Preservation

The test suite in karukan-engine/tests/kanji_conversion_tests.rs verifies that the special tokens survive tokenization:

let prompt = build_jinen_prompt("テスト", "");
let tokens = model.tokenize(&prompt).expect("Tokenisation failed");

let mut found_context = false;
let mut found_input_start = false;
let mut found_output_start = false;

for token in &tokens {
    let display = model.decode_token_for_display(*token);
    if display.contains(CONTEXT_TOKEN) { found_context = true; }
    if display.contains(INPUT_START_TOKEN) { found_input_start = true; }
    if display.contains(OUTPUT_START_TOKEN) { found_output_start = true; }
}
assert!(found_context && found_input_start && found_output_start);

Summary

  • The Jinen format uses three private-use Unicode characters (U+EE00, U+EE01, U+EE02) to structure prompts for LLM-based kanji conversion in the Karukan engine.
  • Context Token (\u{EE02}) precedes already-converted text, Input Start Token (\u{EE00}) marks the katakana reading boundary, and Output Start Token (\u{EE01}) indicates where generation begins.
  • The build_jinen_prompt function in karukan-engine/src/kanji/backend.rs concatenates these elements into a format that survives tokenization and decoding, as implemented in karukan-engine/src/kanji/llamacpp.rs.
  • This architecture ensures clean separation between context, input, and output, enabling accurate incremental kanji conversion through the KanaKanjiConverter::convert method.

Frequently Asked Questions

Why does Karukan use Unicode private-use-area characters instead of plain text markers?

Plain text markers could appear in user input or model output, causing boundary ambiguity. The private-use-area characters U+EE00–U+EE02 are guaranteed not to exist in standard Japanese text, ensuring that the CONTEXT_TOKEN, INPUT_START_TOKEN, and OUTPUT_START_TOKEN remain unambiguous delimiters regardless of the content being processed, even after tokenization.

How does the Jinen format handle tokenization in llama.cpp?

According to the implementation in karukan-engine/src/kanji/llamacpp.rs, the special Unicode tokens are preserved through the tokenization and decoding pipeline. This preservation is essential because it allows the model to reliably locate the OUTPUT_START_TOKEN position after token-to-text conversion, ensuring that generation begins at the correct boundary and that context tokens remain distinguishable from content tokens.

Can I use the Jinen format with models other than those specifically trained for Karukan?

While the format is technically a string convention, effective use requires a model trained to recognize the semantic meaning of U+EE00–U+EE02. The model must understand that text following INPUT_START_TOKEN represents katakana input requiring conversion, and that it should generate kanji immediately after OUTPUT_START_TOKEN. Using untrained models will likely result in improper handling of these control characters or failure to respect the structural boundaries.

Where is the Jinen prompt construction logic located in the source code?

The primary construction logic resides in karukan-engine/src/kanji/backend.rs within the build_jinen_prompt function, while the token constants are defined in karukan-engine/src/kanji/mod.rs. The actual usage during conversion occurs in the KanaKanjiConverter::convert method, which orchestrates the flow from hiragana input through Jinen prompt generation to kanji candidate output, as documented in the togatoga/karukan repository.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →