How the Zenz ML Engine Integrates llama.cpp for AI-Powered Japanese Conversion

The Zenz ML engine integrates llama.cpp by wrapping the native C++ library in a Kotlin ZenzEngine class that loads libzenz.so, forwards model initialization and generation calls to the underlying llama.cpp implementation, and returns Japanese conversion candidates to the Android IME service.

The kazumaproject/japanesekeyboard repository implements an AI-powered Japanese input method editor (IME) that leverages the Zenz ML engine for intelligent text conversion. This engine integrates llama.cpp—the lightweight C++ inference library—to run quantized language models directly on Android devices, enabling offline, low-latency Japanese text generation without cloud dependencies.

Architecture Overview: Zenz as a Native Wrapper

Zenz is architected as a thin Kotlin wrapper around a native shared library. The ZenzEngine class in zenz/src/main/java/com/kazumaproject/zenz/ZenzEngine.kt declares a companion object that loads libzenz.so via System.loadLibrary("zenz"). This native library is compiled from the llama.cpp codebase, specifically tracking the azooKey/llama.cpp upstream repository, and exposes JNI methods that map directly to llama.cpp's model loading, context management, and token generation APIs.

Step-by-Step Integration Flow

Loading the Native Library

The integration begins when the ZenzEngine class initializes. Its init block invokes System.loadLibrary("zenz"), which loads the precompiled shared library containing the llama.cpp inference engine. This happens before any model operations, ensuring the native runtime is available for subsequent JNI calls.

Model Initialization and Configuration

The initModel(modelPath: String) method forwards the absolute path of a GGML-format .bin model file to the native side. On the C++ side, this triggers llama.cpp's model loading sequence: parsing the quantized weights, initializing the KV cache, and preparing the computational graph. If a user-provided model is unavailable, the system falls back to a default model bundled in assets/zenz/, as implemented in the engine provider logic within IMEService.kt.

Runtime Configuration

Before generation begins, setRuntimeConfig(nCtx: Int, nThreads: Int) allows the Kotlin layer to configure the llama.cpp runtime parameters. This sets the context window size (maximum tokens in the KV cache) and the number of CPU threads for inference, enabling performance tuning based on device capabilities.

Generation APIs

The wrapper exposes three native entry points that map directly to llama.cpp generation functions:

  • generate(prompt, maxTokens): Simple text completion without conversational context.
  • generateWithContext(leftContext, input, maxTokens): Generates text conditioned on left-side context (previously typed text).
  • generateWithContextAndConditions(profile, topic, style, preference, leftContext, input, maxTokens): The full-featured API used for Japanese conversion, accepting stylistic parameters alongside context.

These methods return raw UTF-8 strings that the Kotlin layer parses into ZenzCandidate objects.

Candidate Evaluation

The candidateEvaluate(...) method implements the "Zenz-AI → Zenz AI" (Zenzai) evaluation flow. This specialized native call asks llama.cpp to score or rerank a candidate string against the current input context, enabling quality-based candidate filtering before presentation to the user.

Kotlin-Level Orchestration in the IME Service

The IMEService.kt file orchestrates the high-level integration. The performZenzRequest function (lines 82-94) validates that input consists of Hiragana longer than one character, builds the left context from previously typed text, and invokes zenzEngine?.generateWithContextAndConditions(...).

The providesZenzEngine function (lines 13025-13079) handles dependency injection and model selection logic: it checks for custom models in the files directory, falls back to assets/zenz/ for default models, and initializes the engine with appropriate thread counts based on device CPU cores.

Code Examples

Initializing Zenz with a Custom Model

// Called from IMEService.providesZenzEngine(context)
val customFile = File(context.filesDir, "my_zenz_model.bin")
if (customFile.exists()) {
    ZenzEngine.initModel(customFile.absolutePath)
    // Tune llama.cpp runtime parameters
    ZenzEngine.setRuntimeConfig(nCtx = 512, nThreads = 4)
}

Generating Conversion Candidates

suspend fun getZenzCandidates(insert: String): List<ZenzCandidate> =
    withContext(Dispatchers.Default) {
        // Guard: only Hiragana strings longer than 1 char
        if (insert.length <= 1 || !insert.isAllHiraganaWithSymbols()) {
            return@withContext emptyList()
        }

        // Build left-context from previously typed text
        val left = getLeftContext(...)

        // Call native llama.cpp through ZenzEngine
        val raw = ZenzEngine.generateWithContextAndConditions(
            profile = "", topic = "", style = "", preference = "",
            leftContext = left,
            input = insert.hiraganaToKatakana(),
            maxTokens = 32
        )

        // Wrap result for UI presentation
        listOf(
            ZenzCandidate(
                string = raw,
                type = 33.toByte(),
                length = insert.length.toUByte(),
                score = 2000,
                originalString = insert
            )
        )
    }

Evaluating Candidates with Zenzai Flow

suspend fun evaluateZenzCandidate(
    insert: String,
    firstCandidate: String,
    leftContext: String
): ZenzCandidate? = withContext(Dispatchers.Default) {
    val result = ZenzEngine.candidateEvaluate(
        profile = "", topic = "", style = "", preference = "",
        leftContext = leftContext,
        input = insert.hiraganaToKatakana(),
        candidate = firstCandidate
    )
    
    // Parse evaluation result
    if (result.isNotBlank()) {
        ZenzCandidate(
            string = result,
            type = 33.toByte(),
            length = insert.length.toUByte(),
            score = 2000,
            originalString = insert
        )
    } else null
}

Key Source Files and Responsibilities

File Role Link
zenz/src/main/java/com/kazumaproject/zenz/ZenzEngine.kt Kotlin wrapper that loads libzenz.so and declares JNI native methods for llama.cpp integration. https://github.com/kazumaproject/japanesekeyboard/blob/master/zenz/src/main/java/com/kazumaproject/zenz/ZenzEngine.kt
app/src/main/java/com/kazumaproject/markdownhelperkeyboard/ime_service/IMEService.kt (lines 82-94) Orchestrates Zenz calls from the IME, builds left-context, validates Hiragana input, and converts native output to UI candidates. https://github.com/kazumaproject/japanesekeyboard/blob/master/app/src/main/java/com/kazumaproject/markdownhelperkeyboard/ime_service/IMEService.kt#L82-L94
app/src/main/java/com/kazumaproject/markdownhelperkeyboard/ime_service/IMEService.kt (lines 13025-13079) Handles model loading logic, custom vs. default asset fallback, and engine initialization with thread configuration. https://github.com/kazumaproject/japanesekeyboard/blob/master/app/src/main/java/com/kazumaproject/markdownhelperkeyboard/ime_service/IMEService.kt#L13025-L13079
app/src/main/java/com/kazumaproject/markdownhelperkeyboard/setting_activity/ui/opensource/OpenSourceFragment.kt References the upstream azooKey/llama.cpp repository in the open-source attributions. https://github.com/kazumaproject/japanesekeyboard/blob/master/app/src/main/java/com/kazumaproject/markdownhelperkeyboard/setting_activity/ui/opensource/OpenSourceFragment.kt#L64-L66
app/src/main/java/com/kazumaproject/markdownhelperkeyboard/converter/candidate/ZenzCandidate.kt Data class representing a conversion candidate with scoring and metadata returned from the Zenz ML engine. https://github.com/kazumaproject/japanesekeyboard/blob/master/app/src/main/java/com/kazumaproject/markdownhelperkeyboard/converter/candidate/ZenzCandidate.kt

Summary

  • Zenz is a Kotlin-to-native bridge that exposes llama.cpp functionality through libzenz.so, enabling on-device AI inference for Japanese text conversion.
  • The integration follows a wrapper pattern: ZenzEngine.kt declares JNI methods while IMEService.kt orchestrates high-level IME logic, context building, and candidate presentation.
  • Model management supports both custom user models and bundled defaults in assets/zenz/, with runtime configuration for context windows and thread counts via setRuntimeConfig.
  • Generation APIs map directly to llama.cpp functions, supporting simple completion, context-aware generation, and full conditional conversion with profile/topic/style parameters.
  • The Zenzai evaluation flow uses candidateEvaluate to rerank candidates using the underlying ML model, ensuring high-quality conversion suggestions.

Frequently Asked Questions

How does Zenz load the llama.cpp native library on Android?

The ZenzEngine class loads the native library through its init block using System.loadLibrary("zenz"). This call loads libzenz.so, which is the compiled llama.cpp binary packaged within the APK. The library must be loaded before any JNI methods can be invoked, ensuring the C++ runtime is initialized for model inference.

What model format does the Zenz ML engine require?

Zenz requires models in the GGML format (.bin files) generated by llama.cpp. The initModel(modelPath: String) method accepts the absolute path to these binary files, which contain quantized weights compatible with llama.cpp's inference engine. The repository supports both custom user-provided models and a default model bundled in the assets/zenz/ directory as a fallback.

How does the IME service decide when to invoke Zenz for conversion?

The IMEService.kt file contains the performZenzRequest function that validates input before calling the ML engine. It checks that the input string consists of Hiragana characters and exceeds one character in length. If validation passes, the service builds a left-context string from previously typed text and calls generateWithContextAndConditions to retrieve AI-powered conversion candidates, which are then wrapped in ZenzCandidate objects for the suggestion UI.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →