User Dictionary System with Forward-Matching Prediction and Import/Export in JapaneseKeyboard

The JapaneseKeyboard app implements a Room-based user dictionary system that provides real-time forward-matching predictions during typing and supports bulk management through CSV import/export functionality.

The kazumaproject/japanesekeyboard repository delivers a complete user dictionary system designed for Japanese text input. This system enables personalized vocabulary management while integrating seamlessly with the IME service to offer predictive suggestions as users type. Built on Android Room with a clean architecture pattern, the implementation separates data persistence from UI concerns while maintaining high-performance prefix matching for real-time predictions.

Architecture Overview

The user dictionary system follows a three-layer architecture that abstracts database operations through the Repository pattern and exposes reactive data streams via ViewModel.

Entity and Data Access Layer

At the foundation, UserWord.kt defines the Room entity that stores dictionary entries with three core fields: reading (the phonetic input), word (the displayed output), and weight (priority ranking). The corresponding UserWordDao.kt provides the data access interface with methods optimized for forward-matching queries.

Key DAO methods include searchByReadingPrefix for prefix matching using SQL LIKE 'prefix%', commonPrefixSearchInUserDict for IME integration, and standard CRUD operations like insert and deleteAll.

Repository Abstraction

UserDictionaryRepository.kt mediates between the DAO and the rest of the application, ensuring no UI components touch the database directly. It exposes both LiveData streams for UI observation and suspend functions for background operations.

The repository implements searchByReadingPrefix as LiveData (lines 21-23) for reactive UI updates, while commonPrefixSearchInUserDict (lines 33-35) provides a suspend API for the IME service to fetch predictions asynchronously.

ViewModel and UI Layer

UserDictionaryViewModel.kt exposes user-friendly data streams and handles import/export actions. It forwards prediction requests to the repository and posts results to the UI via LiveData, while providing importCsv() and exportCsv() helper methods for file operations.

The UserDictionaryFragment.kt renders the dictionary list, provides a search field for forward-matching, and hosts Import/Export buttons. It observes the ViewModel's LiveData and invokes the ViewModel's import/export functions when users select files through the Storage Access Framework.

Forward-Matching Prediction Implementation

The forward-matching system enables real-time suggestions as users type phonetic readings. When the IME service detects input, it queries the user dictionary for entries where the reading column starts with the current input string.

The mechanism relies on SQL LIKE 'prefix%' queries executed through Room. In IMEService.kt, the prediction flow captures the current reading and requests matching candidates:

// Inside IMEService.kt (prediction flow)
val prefix = currentReading // the characters typed so far
val candidates = userDictionaryRepository.searchByReadingPrefix(prefix)   // LiveData<List<UserWord>>

searchByReadingPrefix delegates to UserWordDao.searchByReadingPrefix, which executes the prefix query. Because the repository exposes this as LiveData, the UI updates instantly when the database changes. For the IME service's direct consumption, commonPrefixSearchInUserDict provides a suspend function that returns results asynchronously without LiveData overhead.

This architecture ensures low-latency predictions even with large user dictionaries, as SQLite indexes optimize prefix queries on the reading column.

Import and Export Functionality

The user dictionary supports bulk management through CSV files, enabling users to backup, restore, or migrate their custom vocabulary across devices.

Exporting to CSV

The export process converts all stored UserWord objects to comma-separated values and writes them to a user-selected location via the Storage Access Framework. The ViewModel handles the background processing:

fun exportCsv(uri: Uri) = viewModelScope.launch {
    val words = repository.allWords.value ?: emptyList()
    val csv = words.joinToString("\n") { "${it.reading},${it.word},${it.weight}" }
    context.contentResolver.openOutputStream(uri)?.use { it.write(csv.toByteArray()) }
}

This method runs in a coroutine to avoid blocking the main thread, ensuring the UI remains responsive during large exports.

Importing from CSV

Import reverses the process by parsing a selected CSV file and inserting entries into the Room database. The fragment launches an ACTION_OPEN_DOCUMENT intent to let users choose files, then delegates to the ViewModel:

fun importCsv(uri: Uri) = viewModelScope.launch {
    val text = context.contentResolver.openInputStream(uri)?.bufferedReader()?.readText()
    val words = text?.lines()
        ?.filter { it.isNotBlank() }
        ?.map {
            val (reading, word, weight) = it.split(',')
            UserWord(reading = reading, word = word, weight = weight.toInt())
        } ?: emptyList()
    repository.insertAll(words)
}

Both operations validate input to prevent crashes from malformed CSV data, and all database writes occur off the UI thread through the repository's suspend functions.

Summary

  • The JapaneseKeyboard app implements a user dictionary system using Android Room with a three-layer architecture (Entity/DAO, Repository, ViewModel).
  • Forward-matching prediction uses SQL LIKE 'prefix%' queries via searchByReadingPrefix in UserWordDao.kt, enabling real-time suggestions as users type.
  • The Repository pattern in UserDictionaryRepository.kt abstracts database access, exposing both LiveData for UI observation and suspend functions for background IME operations.
  • Import and export functionality supports CSV format through UserDictionaryViewModel.kt, allowing users to backup and restore dictionaries via the Storage Access Framework.
  • All database operations run on background threads using Kotlin coroutines, ensuring smooth UI performance even with large custom dictionaries.

Frequently Asked Questions

How does the forward-matching prediction work in the JapaneseKeyboard user dictionary?

The system captures the current phonetic reading as the user types and queries the database for entries where the reading column starts with the input prefix. The UserWordDao.kt executes a SQL LIKE 'prefix%' query through the searchByReadingPrefix method, returning results as LiveData for instant UI updates or as a suspend function for IME service integration.

What file format does the import/export system use?

The user dictionary system uses standard CSV (Comma-Separated Values) format with three columns: reading, word, and weight. During export, UserDictionaryViewModel.kt generates a UTF-8 encoded file where each line represents one dictionary entry. During import, the system parses the CSV, validates the three fields, and inserts valid entries into the Room database via repository.insertAll().

Where is the user dictionary data stored in the codebase?

The dictionary data persists in a SQLite database managed by Android Room, with entity definitions in UserWord.kt and data access methods in UserWordDao.kt. The repository abstraction lives in UserDictionaryRepository.kt, while the UI logic resides in UserDictionaryViewModel.kt and UserDictionaryFragment.kt. The IME service accesses predictions through IMEService.kt in the ime_service package.

Can the user dictionary handle large vocabularies without UI lag?

Yes, the architecture is designed for performance with large dictionaries. All database queries run on background threads using Kotlin coroutines and suspend functions. The UserDictionaryRepository.kt exposes both LiveData streams for reactive UI updates and direct suspend APIs for the IME service. SQLite indexes optimize the prefix queries used in forward-matching, ensuring sub-millisecond response times even with thousands of user-defined entries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →