# User Dictionary System with Forward-Matching Prediction and Import/Export in JapaneseKeyboard

> Explore the JapaneseKeyboard app's user dictionary system with real-time forward-matching predictions and easy CSV import export for efficient word management.

- Repository: [Kazu/japanesekeyboard](https://github.com/kazumaproject/japanesekeyboard)
- Tags: deep-dive
- Published: 2026-03-05

---

**The JapaneseKeyboard app implements a Room-based user dictionary system that provides real-time forward-matching predictions during typing and supports bulk management through CSV import/export functionality.**

The `kazumaproject/japanesekeyboard` repository delivers a complete user dictionary system designed for Japanese text input. This system enables personalized vocabulary management while integrating seamlessly with the IME service to offer predictive suggestions as users type. Built on Android Room with a clean architecture pattern, the implementation separates data persistence from UI concerns while maintaining high-performance prefix matching for real-time predictions.

## Architecture Overview

The user dictionary system follows a three-layer architecture that abstracts database operations through the Repository pattern and exposes reactive data streams via ViewModel.

### Entity and Data Access Layer

At the foundation, [`UserWord.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserWord.kt) defines the Room entity that stores dictionary entries with three core fields: `reading` (the phonetic input), `word` (the displayed output), and `weight` (priority ranking). The corresponding [`UserWordDao.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserWordDao.kt) provides the data access interface with methods optimized for forward-matching queries.

Key DAO methods include `searchByReadingPrefix` for prefix matching using SQL `LIKE 'prefix%'`, `commonPrefixSearchInUserDict` for IME integration, and standard CRUD operations like `insert` and `deleteAll`.

### Repository Abstraction

[`UserDictionaryRepository.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryRepository.kt) mediates between the DAO and the rest of the application, ensuring no UI components touch the database directly. It exposes both `LiveData` streams for UI observation and `suspend` functions for background operations.

The repository implements `searchByReadingPrefix` as LiveData (lines 21-23) for reactive UI updates, while `commonPrefixSearchInUserDict` (lines 33-35) provides a suspend API for the IME service to fetch predictions asynchronously.

### ViewModel and UI Layer

[`UserDictionaryViewModel.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryViewModel.kt) exposes user-friendly data streams and handles import/export actions. It forwards prediction requests to the repository and posts results to the UI via LiveData, while providing `importCsv()` and `exportCsv()` helper methods for file operations.

The [`UserDictionaryFragment.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryFragment.kt) renders the dictionary list, provides a search field for forward-matching, and hosts Import/Export buttons. It observes the ViewModel's LiveData and invokes the ViewModel's import/export functions when users select files through the Storage Access Framework.

## Forward-Matching Prediction Implementation

The forward-matching system enables real-time suggestions as users type phonetic readings. When the IME service detects input, it queries the user dictionary for entries where the reading column starts with the current input string.

The mechanism relies on SQL `LIKE 'prefix%'` queries executed through Room. In [`IMEService.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/IMEService.kt), the prediction flow captures the current reading and requests matching candidates:

```kotlin
// Inside IMEService.kt (prediction flow)
val prefix = currentReading // the characters typed so far
val candidates = userDictionaryRepository.searchByReadingPrefix(prefix)   // LiveData<List<UserWord>>

```

`searchByReadingPrefix` delegates to `UserWordDao.searchByReadingPrefix`, which executes the prefix query. Because the repository exposes this as `LiveData`, the UI updates instantly when the database changes. For the IME service's direct consumption, `commonPrefixSearchInUserDict` provides a suspend function that returns results asynchronously without LiveData overhead.

This architecture ensures low-latency predictions even with large user dictionaries, as SQLite indexes optimize prefix queries on the reading column.

## Import and Export Functionality

The user dictionary supports bulk management through CSV files, enabling users to backup, restore, or migrate their custom vocabulary across devices.

### Exporting to CSV

The export process converts all stored `UserWord` objects to comma-separated values and writes them to a user-selected location via the Storage Access Framework. The ViewModel handles the background processing:

```kotlin
fun exportCsv(uri: Uri) = viewModelScope.launch {
    val words = repository.allWords.value ?: emptyList()
    val csv = words.joinToString("\n") { "${it.reading},${it.word},${it.weight}" }
    context.contentResolver.openOutputStream(uri)?.use { it.write(csv.toByteArray()) }
}

```

This method runs in a coroutine to avoid blocking the main thread, ensuring the UI remains responsive during large exports.

### Importing from CSV

Import reverses the process by parsing a selected CSV file and inserting entries into the Room database. The fragment launches an `ACTION_OPEN_DOCUMENT` intent to let users choose files, then delegates to the ViewModel:

```kotlin
fun importCsv(uri: Uri) = viewModelScope.launch {
    val text = context.contentResolver.openInputStream(uri)?.bufferedReader()?.readText()
    val words = text?.lines()
        ?.filter { it.isNotBlank() }
        ?.map {
            val (reading, word, weight) = it.split(',')
            UserWord(reading = reading, word = word, weight = weight.toInt())
        } ?: emptyList()
    repository.insertAll(words)
}

```

Both operations validate input to prevent crashes from malformed CSV data, and all database writes occur off the UI thread through the repository's suspend functions.

## Summary

- The JapaneseKeyboard app implements a **user dictionary system** using Android Room with a three-layer architecture (Entity/DAO, Repository, ViewModel).
- **Forward-matching prediction** uses SQL `LIKE 'prefix%'` queries via `searchByReadingPrefix` in [`UserWordDao.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserWordDao.kt), enabling real-time suggestions as users type.
- The **Repository pattern** in [`UserDictionaryRepository.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryRepository.kt) abstracts database access, exposing both `LiveData` for UI observation and suspend functions for background IME operations.
- **Import and export** functionality supports CSV format through [`UserDictionaryViewModel.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryViewModel.kt), allowing users to backup and restore dictionaries via the Storage Access Framework.
- All database operations run on background threads using Kotlin coroutines, ensuring smooth UI performance even with large custom dictionaries.

## Frequently Asked Questions

### How does the forward-matching prediction work in the JapaneseKeyboard user dictionary?

The system captures the current phonetic reading as the user types and queries the database for entries where the reading column starts with the input prefix. The [`UserWordDao.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserWordDao.kt) executes a SQL `LIKE 'prefix%'` query through the `searchByReadingPrefix` method, returning results as `LiveData` for instant UI updates or as a suspend function for IME service integration.

### What file format does the import/export system use?

The user dictionary system uses standard CSV (Comma-Separated Values) format with three columns: reading, word, and weight. During export, [`UserDictionaryViewModel.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryViewModel.kt) generates a UTF-8 encoded file where each line represents one dictionary entry. During import, the system parses the CSV, validates the three fields, and inserts valid entries into the Room database via `repository.insertAll()`.

### Where is the user dictionary data stored in the codebase?

The dictionary data persists in a SQLite database managed by Android Room, with entity definitions in [`UserWord.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserWord.kt) and data access methods in [`UserWordDao.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserWordDao.kt). The repository abstraction lives in [`UserDictionaryRepository.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryRepository.kt), while the UI logic resides in [`UserDictionaryViewModel.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryViewModel.kt) and [`UserDictionaryFragment.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryFragment.kt). The IME service accesses predictions through [`IMEService.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/IMEService.kt) in the `ime_service` package.

### Can the user dictionary handle large vocabularies without UI lag?

Yes, the architecture is designed for performance with large dictionaries. All database queries run on background threads using Kotlin coroutines and suspend functions. The [`UserDictionaryRepository.kt`](https://github.com/kazumaproject/japanesekeyboard/blob/main/UserDictionaryRepository.kt) exposes both `LiveData` streams for reactive UI updates and direct suspend APIs for the IME service. SQLite indexes optimize the prefix queries used in forward-matching, ensuring sub-millisecond response times even with thousands of user-defined entries.