How the Zenz ML Engine Integrates llama.cpp for AI-Powered Japanese Conversion
The Zenz ML engine integrates llama.cpp by wrapping the native C++ library in a Kotlin ZenzEngine class that loads libzenz.so, forwards model initialization and generation calls to the underlying llama.cpp implementation, and returns Japanese conversion candidates to the Android IME service.
The kazumaproject/japanesekeyboard repository implements an AI-powered Japanese input method editor (IME) that leverages the Zenz ML engine for intelligent text conversion. This engine integrates llama.cpp—the lightweight C++ inference library—to run quantized language models directly on Android devices, enabling offline, low-latency Japanese text generation without cloud dependencies.
Architecture Overview: Zenz as a Native Wrapper
Zenz is architected as a thin Kotlin wrapper around a native shared library. The ZenzEngine class in zenz/src/main/java/com/kazumaproject/zenz/ZenzEngine.kt declares a companion object that loads libzenz.so via System.loadLibrary("zenz"). This native library is compiled from the llama.cpp codebase, specifically tracking the azooKey/llama.cpp upstream repository, and exposes JNI methods that map directly to llama.cpp's model loading, context management, and token generation APIs.
Step-by-Step Integration Flow
Loading the Native Library
The integration begins when the ZenzEngine class initializes. Its init block invokes System.loadLibrary("zenz"), which loads the precompiled shared library containing the llama.cpp inference engine. This happens before any model operations, ensuring the native runtime is available for subsequent JNI calls.
Model Initialization and Configuration
The initModel(modelPath: String) method forwards the absolute path of a GGML-format .bin model file to the native side. On the C++ side, this triggers llama.cpp's model loading sequence: parsing the quantized weights, initializing the KV cache, and preparing the computational graph. If a user-provided model is unavailable, the system falls back to a default model bundled in assets/zenz/, as implemented in the engine provider logic within IMEService.kt.
Runtime Configuration
Before generation begins, setRuntimeConfig(nCtx: Int, nThreads: Int) allows the Kotlin layer to configure the llama.cpp runtime parameters. This sets the context window size (maximum tokens in the KV cache) and the number of CPU threads for inference, enabling performance tuning based on device capabilities.
Generation APIs
The wrapper exposes three native entry points that map directly to llama.cpp generation functions:
generate(prompt, maxTokens): Simple text completion without conversational context.generateWithContext(leftContext, input, maxTokens): Generates text conditioned on left-side context (previously typed text).generateWithContextAndConditions(profile, topic, style, preference, leftContext, input, maxTokens): The full-featured API used for Japanese conversion, accepting stylistic parameters alongside context.
These methods return raw UTF-8 strings that the Kotlin layer parses into ZenzCandidate objects.
Candidate Evaluation
The candidateEvaluate(...) method implements the "Zenz-AI → Zenz AI" (Zenzai) evaluation flow. This specialized native call asks llama.cpp to score or rerank a candidate string against the current input context, enabling quality-based candidate filtering before presentation to the user.
Kotlin-Level Orchestration in the IME Service
The IMEService.kt file orchestrates the high-level integration. The performZenzRequest function (lines 82-94) validates that input consists of Hiragana longer than one character, builds the left context from previously typed text, and invokes zenzEngine?.generateWithContextAndConditions(...).
The providesZenzEngine function (lines 13025-13079) handles dependency injection and model selection logic: it checks for custom models in the files directory, falls back to assets/zenz/ for default models, and initializes the engine with appropriate thread counts based on device CPU cores.
Code Examples
Initializing Zenz with a Custom Model
// Called from IMEService.providesZenzEngine(context)
val customFile = File(context.filesDir, "my_zenz_model.bin")
if (customFile.exists()) {
ZenzEngine.initModel(customFile.absolutePath)
// Tune llama.cpp runtime parameters
ZenzEngine.setRuntimeConfig(nCtx = 512, nThreads = 4)
}
Generating Conversion Candidates
suspend fun getZenzCandidates(insert: String): List<ZenzCandidate> =
withContext(Dispatchers.Default) {
// Guard: only Hiragana strings longer than 1 char
if (insert.length <= 1 || !insert.isAllHiraganaWithSymbols()) {
return@withContext emptyList()
}
// Build left-context from previously typed text
val left = getLeftContext(...)
// Call native llama.cpp through ZenzEngine
val raw = ZenzEngine.generateWithContextAndConditions(
profile = "", topic = "", style = "", preference = "",
leftContext = left,
input = insert.hiraganaToKatakana(),
maxTokens = 32
)
// Wrap result for UI presentation
listOf(
ZenzCandidate(
string = raw,
type = 33.toByte(),
length = insert.length.toUByte(),
score = 2000,
originalString = insert
)
)
}
Evaluating Candidates with Zenzai Flow
suspend fun evaluateZenzCandidate(
insert: String,
firstCandidate: String,
leftContext: String
): ZenzCandidate? = withContext(Dispatchers.Default) {
val result = ZenzEngine.candidateEvaluate(
profile = "", topic = "", style = "", preference = "",
leftContext = leftContext,
input = insert.hiraganaToKatakana(),
candidate = firstCandidate
)
// Parse evaluation result
if (result.isNotBlank()) {
ZenzCandidate(
string = result,
type = 33.toByte(),
length = insert.length.toUByte(),
score = 2000,
originalString = insert
)
} else null
}
Key Source Files and Responsibilities
Summary
- Zenz is a Kotlin-to-native bridge that exposes llama.cpp functionality through
libzenz.so, enabling on-device AI inference for Japanese text conversion. - The integration follows a wrapper pattern:
ZenzEngine.ktdeclares JNI methods whileIMEService.ktorchestrates high-level IME logic, context building, and candidate presentation. - Model management supports both custom user models and bundled defaults in
assets/zenz/, with runtime configuration for context windows and thread counts viasetRuntimeConfig. - Generation APIs map directly to llama.cpp functions, supporting simple completion, context-aware generation, and full conditional conversion with profile/topic/style parameters.
- The Zenzai evaluation flow uses
candidateEvaluateto rerank candidates using the underlying ML model, ensuring high-quality conversion suggestions.
Frequently Asked Questions
How does Zenz load the llama.cpp native library on Android?
The ZenzEngine class loads the native library through its init block using System.loadLibrary("zenz"). This call loads libzenz.so, which is the compiled llama.cpp binary packaged within the APK. The library must be loaded before any JNI methods can be invoked, ensuring the C++ runtime is initialized for model inference.
What model format does the Zenz ML engine require?
Zenz requires models in the GGML format (.bin files) generated by llama.cpp. The initModel(modelPath: String) method accepts the absolute path to these binary files, which contain quantized weights compatible with llama.cpp's inference engine. The repository supports both custom user-provided models and a default model bundled in the assets/zenz/ directory as a fallback.
How does the IME service decide when to invoke Zenz for conversion?
The IMEService.kt file contains the performZenzRequest function that validates input before calling the ML engine. It checks that the input string consists of Hiragana characters and exceeds one character in length. If validation passes, the service builds a left-context string from previously typed text and calls generateWithContextAndConditions to retrieve AI-powered conversion candidates, which are then wrapped in ZenzCandidate objects for the suggestion UI.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →