N-gram Language Model Ranking in the Mozc-Based Conversion Engine: How It Works

The N-gram language model ranking in the Mozc-based conversion engine combines base dictionary frequencies, user-specific Mozc-UT entries, and contextual trigram probabilities to score candidates, ultimately selecting the highest-ranked result via the rank() method in ZenzCandidate.kt.

The kazumaproject/japanesekeyboard repository implements a sophisticated Japanese input method that leverages Google's Mozc linguistic architecture. At the heart of this system lies the N-gram language model ranking mechanism that evaluates conversion candidates using multiple probabilistic signals drawn from both system and user dictionaries. This article examines how the KanaKanjiEngine class orchestrates candidate generation and contextual scoring to deliver accurate Japanese text conversions.

The Three Pillars of Mozc N-gram Ranking

The ranking system integrates three distinct information sources to calculate a candidate's final score. Each source contributes specific probabilistic data that helps resolve ambiguities inherent in Japanese kana-to-kanji conversion.

Mozc-UT User Dictionary Entries

The engine incorporates Mozc-UT (user-dictionary) entries spanning person names, geographic places, Wikipedia titles, Neologd contemporary terms, and web search results. In KanaKanjiEngine.kt, the system constructs specialized lists including mozcUTPersonNames and mozcUTPlacesList around lines 1060-1080. These entries contribute token probabilities derived from user-dictionary N-grams, allowing personal vocabulary to outrank generic homonyms when their contextual probability exceeds baseline frequencies.

Built-in System Dictionary Frequencies

The system-trie provides base lexical frequencies stored as "cost" values in the Mozc binary format. The KanaKanjiEngine loads systemYomiTrie and systemTokenArray during initialization to establish baseline word commonality. The internal graph builder calculates base costs for each conversion path, ensuring frequently used words receive favorable scores before contextual adjustments apply.

Contextual Trigram Scoring

The contextual N-gram score evaluates how well a candidate continues the preceding text using Mozc's trigram model. This mechanism resolves critical ambiguities—such as distinguishing "はし" as either "橋" (bridge) or "箸" (chopsticks)—by analyzing the two tokens immediately preceding the current input. The scoring implementation resides in ZenzCandidate.kt, while IMEService.kt invokes it at line 7281 via maxByOrNull { it.rank(prefix) }.

The N-gram Ranking Algorithm Flow

The conversion engine processes candidates through four distinct phases, transforming raw dictionary entries into contextually ranked suggestions.

1. Candidate Generation

KanaKanjiEngine.convert() aggregates candidates from multiple sources including the Zenz language model, reading corrections, predictive search, and all Mozc-UT specialized lists:

// KanaKanjiEngine.kt (lines ~1060-1080)
val mozcUTPersonNames = if (mozcUtPersonName == true) getMozcUTPersonNames(input) else emptyList()
val mozcUTPlacesList = if (mozcUTPlaces == true) getMozcUTPlace(input) else emptyList()
val mozcUTWikiList = getMozcUTWiki(input)
val mozcUTNeologdList = getMozcUTNeologd(input)

val allCandidates = resultNBestFinalDeferred + readingCorrectionListDeferred +
    predictiveSearchResult + mozcUTPersonNames + mozcUTPlacesList +
    mozcUTWikiList + mozcUTNeologdList

2. ZenzCandidate Wrapping

Each raw candidate transforms into a ZenzCandidate data class instance that stores the display string, type classification, length, raw engine score, and original input string for downstream processing.

3. Composite Score Calculation

The rank() function in ZenzCandidate.kt computes a weighted sum combining base frequency, length specificity, and contextual probability:

data class ZenzCandidate(
    val string: String,
    val type: Byte,
    val length: UByte,
    val score: Int,
    val originalString: String
) {
    fun rank(prefix: String): Int {
        var total = score                         // Base engine score (~2000 for Zenz)
        total += length.toInt() * 10              // Length specificity bonus
        total += LanguageModel.probability(prefix, string) * 100  // N-gram context boost
        return total
    }
}

The LanguageModel.probability() method wraps Mozc's native N-gram lookup, querying the trigram probability for the sequence formed by the last two tokens of prefix and the current candidate string.

4. Final Selection in IMEService

IMEService.kt executes the ultimate selection at line 7281 within the Zenzai request handler:

// IMEService.kt - performZenzaiRequest
val topCandidate = candidates
    .maxByOrNull { it.rank(prefix) }
    ?: secondCandidateFromZenz

This invocation evaluates every candidate's contextual fit against the current prefix, ensuring the returned result maximizes both lexical frequency and grammatical continuity.

Key Implementation Files

File Role
app/src/main/java/com/kazumaproject/markdownhelperkeyboard/converter/engine/KanaKanjiEngine.kt Aggregates dictionary sources and manages system trie loading
app/src/main/java/com/kazumaproject/markdownhelperkeyboard/converter/candidate/ZenzCandidate.kt Implements the rank() scoring formula
app/src/main/java/com/kazumaproject/markdownhelperkeyboard/ime_service/IMEService.kt Orchestrates final candidate selection at line 7281

Summary

  • N-gram language model ranking integrates Mozc-UT user dictionaries, system-trie base frequencies, and contextual trigram probabilities to resolve conversion ambiguities.
  • KanaKanjiEngine.kt assembles candidate pools around lines 1060-1080 from specialized sources including Wikipedia, Neologd, and personal name databases.
  • The rank() function applies a weighted formula: base engine score + (length × 10) + (N-gram probability × 100).
  • IMEService.kt utilizes maxByOrNull { it.rank(prefix) } to select the contextually optimal candidate for display.

Frequently Asked Questions

How does the engine disambiguate readings like "はし" (bridge vs. chopsticks)?

The contextual trigram analysis in LanguageModel.probability() examines the two tokens preceding the current input. When the prefix contains dining-related vocabulary, the N-gram model assigns higher probability to "箸" (chopsticks); for location-based prefixes, it boosts "橋" (bridge). This contextual boost typically adds 50-200 points to the composite score, often sufficient to override base frequency differences.

Can developers modify the N-gram weighting constants?

Yes. The multipliers within ZenzCandidate.rank()—specifically the length bonus (* 10) and probability scaling (* 100)—represent tunable constants. Adjusting these values changes the relative importance of specificity versus contextual fit, though such modifications require recompiling the ZenzCandidate.kt source.

What distinguishes Mozc-UT entries from standard dictionary entries?

Mozc-UT entries derive from continuously updated corpora including web searches and Wikipedia rather than static lexical databases. The KanaKanjiEngine maintains separate loading functions (getMozcUTPersonNames, getMozcUTPlace) that calculate N-gram probabilities from user-dictionary statistics, allowing recently learned proper nouns to outrank generic terms even when their base frequency would otherwise be lower.

Where does the actual N-gram probability calculation occur?

The heavy computation executes within Mozc's native LanguageModel class in the upstream Google Mozc library (specifically within engine/rewriter/language_model.cc). The kazumaproject implementation accesses these functions through the LanguageModel.probability() JNI bridge defined in the project's Gradle dependencies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →