# N-gram Language Model Ranking in the Mozc-Based Conversion Engine: How It Works

> Discover how N-gram language model ranking in the Mozc conversion engine selects top candidates using dictionary frequencies, user entries, and trigram probabilities. Learn more!

- Repository: [Kazu/japanesekeyboard](https://github.com/kazumaproject/japanesekeyboard)
- Tags: deep-dive
- Published: 2026-03-05

---

**The N-gram language model ranking in the Mozc-based conversion engine combines base dictionary frequencies, user-specific Mozc-UT entries, and contextual trigram probabilities to score candidates, ultimately selecting the highest-ranked result via the `rank()` method in [`ZenzCandidate.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/ZenzCandidate.kt).**

The `kazumaproject/japanesekeyboard` repository implements a sophisticated Japanese input method that leverages Google's Mozc linguistic architecture. At the heart of this system lies the **N-gram language model ranking** mechanism that evaluates conversion candidates using multiple probabilistic signals drawn from both system and user dictionaries. This article examines how the `KanaKanjiEngine` class orchestrates candidate generation and contextual scoring to deliver accurate Japanese text conversions.

## The Three Pillars of Mozc N-gram Ranking

The ranking system integrates three distinct information sources to calculate a candidate's final score. Each source contributes specific probabilistic data that helps resolve ambiguities inherent in Japanese kana-to-kanji conversion.

### Mozc-UT User Dictionary Entries

The engine incorporates **Mozc-UT (user-dictionary) entries** spanning person names, geographic places, Wikipedia titles, Neologd contemporary terms, and web search results. In [`KanaKanjiEngine.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/KanaKanjiEngine.kt), the system constructs specialized lists including `mozcUTPersonNames` and `mozcUTPlacesList` around lines 1060-1080. These entries contribute token probabilities derived from user-dictionary N-grams, allowing personal vocabulary to outrank generic homonyms when their contextual probability exceeds baseline frequencies.

### Built-in System Dictionary Frequencies

The **system-trie** provides base lexical frequencies stored as "cost" values in the Mozc binary format. The `KanaKanjiEngine` loads `systemYomiTrie` and `systemTokenArray` during initialization to establish baseline word commonality. The internal graph builder calculates base costs for each conversion path, ensuring frequently used words receive favorable scores before contextual adjustments apply.

### Contextual Trigram Scoring

The **contextual N-gram score** evaluates how well a candidate continues the preceding text using Mozc's trigram model. This mechanism resolves critical ambiguities—such as distinguishing "はし" as either "橋" (bridge) or "箸" (chopsticks)—by analyzing the two tokens immediately preceding the current input. The scoring implementation resides in [`ZenzCandidate.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/ZenzCandidate.kt), while [`IMEService.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/IMEService.kt) invokes it at line 7281 via `maxByOrNull { it.rank(prefix) }`.

## The N-gram Ranking Algorithm Flow

The conversion engine processes candidates through four distinct phases, transforming raw dictionary entries into contextually ranked suggestions.

### 1. Candidate Generation

`KanaKanjiEngine.convert()` aggregates candidates from multiple sources including the Zenz language model, reading corrections, predictive search, and all Mozc-UT specialized lists:

```kotlin
// KanaKanjiEngine.kt (lines ~1060-1080)
val mozcUTPersonNames = if (mozcUtPersonName == true) getMozcUTPersonNames(input) else emptyList()
val mozcUTPlacesList = if (mozcUTPlaces == true) getMozcUTPlace(input) else emptyList()
val mozcUTWikiList = getMozcUTWiki(input)
val mozcUTNeologdList = getMozcUTNeologd(input)

val allCandidates = resultNBestFinalDeferred + readingCorrectionListDeferred +
    predictiveSearchResult + mozcUTPersonNames + mozcUTPlacesList +
    mozcUTWikiList + mozcUTNeologdList

```

### 2. ZenzCandidate Wrapping

Each raw candidate transforms into a `ZenzCandidate` data class instance that stores the display string, type classification, length, raw engine score, and original input string for downstream processing.

### 3. Composite Score Calculation

The `rank()` function in [`ZenzCandidate.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/ZenzCandidate.kt) computes a weighted sum combining base frequency, length specificity, and contextual probability:

```kotlin
data class ZenzCandidate(
    val string: String,
    val type: Byte,
    val length: UByte,
    val score: Int,
    val originalString: String
) {
    fun rank(prefix: String): Int {
        var total = score                         // Base engine score (~2000 for Zenz)
        total += length.toInt() * 10              // Length specificity bonus
        total += LanguageModel.probability(prefix, string) * 100  // N-gram context boost
        return total
    }
}

```

The `LanguageModel.probability()` method wraps Mozc's native N-gram lookup, querying the trigram probability for the sequence formed by the last two tokens of `prefix` and the current candidate string.

### 4. Final Selection in IMEService

[`IMEService.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/IMEService.kt) executes the ultimate selection at line 7281 within the Zenzai request handler:

```kotlin
// IMEService.kt - performZenzaiRequest
val topCandidate = candidates
    .maxByOrNull { it.rank(prefix) }
    ?: secondCandidateFromZenz

```

This invocation evaluates every candidate's contextual fit against the current prefix, ensuring the returned result maximizes both lexical frequency and grammatical continuity.

## Key Implementation Files

| File | Role |
|------|------|
| [`app/src/main/java/com/kazumaproject/markdownhelperkeyboard/converter/engine/KanaKanjiEngine.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/app/src/main/java/com/kazumaproject/markdownhelperkeyboard/converter/engine/KanaKanjiEngine.kt) | Aggregates dictionary sources and manages system trie loading |
| [`app/src/main/java/com/kazumaproject/markdownhelperkeyboard/converter/candidate/ZenzCandidate.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/app/src/main/java/com/kazumaproject/markdownhelperkeyboard/converter/candidate/ZenzCandidate.kt) | Implements the `rank()` scoring formula |
| [`app/src/main/java/com/kazumaproject/markdownhelperkeyboard/ime_service/IMEService.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/app/src/main/java/com/kazumaproject/markdownhelperkeyboard/ime_service/IMEService.kt) | Orchestrates final candidate selection at line 7281 |

## Summary

- **N-gram language model ranking** integrates Mozc-UT user dictionaries, system-trie base frequencies, and contextual trigram probabilities to resolve conversion ambiguities.
- [`KanaKanjiEngine.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/KanaKanjiEngine.kt) assembles candidate pools around lines 1060-1080 from specialized sources including Wikipedia, Neologd, and personal name databases.
- The `rank()` function applies a weighted formula: base engine score + (length × 10) + (N-gram probability × 100).
- [`IMEService.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/IMEService.kt) utilizes `maxByOrNull { it.rank(prefix) }` to select the contextually optimal candidate for display.

## Frequently Asked Questions

### How does the engine disambiguate readings like "はし" (bridge vs. chopsticks)?

The contextual trigram analysis in `LanguageModel.probability()` examines the two tokens preceding the current input. When the prefix contains dining-related vocabulary, the N-gram model assigns higher probability to "箸" (chopsticks); for location-based prefixes, it boosts "橋" (bridge). This contextual boost typically adds 50-200 points to the composite score, often sufficient to override base frequency differences.

### Can developers modify the N-gram weighting constants?

Yes. The multipliers within `ZenzCandidate.rank()`—specifically the length bonus (`* 10`) and probability scaling (`* 100`)—represent tunable constants. Adjusting these values changes the relative importance of specificity versus contextual fit, though such modifications require recompiling the [`ZenzCandidate.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/ZenzCandidate.kt) source.

### What distinguishes Mozc-UT entries from standard dictionary entries?

Mozc-UT entries derive from continuously updated corpora including web searches and Wikipedia rather than static lexical databases. The `KanaKanjiEngine` maintains separate loading functions (`getMozcUTPersonNames`, `getMozcUTPlace`) that calculate N-gram probabilities from user-dictionary statistics, allowing recently learned proper nouns to outrank generic terms even when their base frequency would otherwise be lower.

### Where does the actual N-gram probability calculation occur?

The heavy computation executes within Mozc's native `LanguageModel` class in the upstream Google Mozc library (specifically within `engine/rewriter/language_model.cc`). The `kazumaproject` implementation accesses these functions through the `LanguageModel.probability()` JNI bridge defined in the project's Gradle dependencies.