How Karukan Manages and Loads User Dictionaries from defaultuserdict.txt
Karukan loads user dictionaries as JSON files from a platform-specific directory, automatically copying the bundled defaultuserdict.txt on first startup, then merges them into a single in-memory dictionary that takes priority over the system dictionary during candidate generation.
Karukan treats every file in the user-dictionary directory as a separate JSON-encoded user dictionary capable of overriding system entries. The Rust-based input method engine handles these dictionaries through a dedicated initialization process in karukan-im/src/core/engine/init.rs, ensuring user customizations persist across sessions while appearing first in conversion candidates.
User Dictionary Storage Location
Karukan determines the storage location through Settings::user_dict_dir() defined in karukan-im/src/config/settings.rs. This method returns platform-specific paths following XDG Base Directory and macOS application support conventions:
- Linux:
$XDG_DATA_HOME/karukan-im/user_dicts(defaults to~/.local/share/karukan-im/user_dicts) - macOS:
$HOME/Library/Application Support/com.karukan.karukan-im/user_dicts
The engine creates this directory automatically using std::fs::create_dir_all if it does not exist during the initialization phase.
The Default Dictionary File
The repository ships a starter dictionary at karukan-im/user_dicts/defaultuserdict.txt. This plain-text file contains JSON-formatted mappings where each line represents a reading-to-candidate pair (for example, てすと,).
During engine startup, the initialize_user_dicts function copies this bundled file into the user's dictionary directory if the directory is missing or empty. This ensures first-time users have a functional dictionary without manual configuration, using std::fs::copy to place the file before proceeding with the loading sequence.
Loading and Merging Process
The initialize_user_dicts function in karukan-im/src/core/engine/init.rs (lines 216-287) orchestrates dictionary loading through a multi-step pipeline:
Directory Initialization and Default Copying
The function first resolves the path via Settings::user_dict_dir(). If the directory is missing, it creates the hierarchy and copies the bundled defaultuserdict.txt into the location. This ensures the engine always has a baseline dictionary before attempting to load user customizations.
File Discovery and JSON Deserialization
The engine scans the directory using fs::read_dir, processing entries with .txt or .json extensions. For each file, it opens a reader and deserializes the content using serde_json::from_reader into UserDictionary structs. Errors during deserialization are logged via warn! macros but do not abort startup, allowing the engine to continue with valid dictionaries while skipping malformed files.
Dictionary Aggregation
All successfully loaded dictionaries merge using UserDictionary::merge, producing a single consolidated dictionary stored in the Engine::user_dict field. This merged structure resides in memory and serves as the authoritative user dictionary for the entire session, consulted before the system dictionary during text conversion.
Search Priority During Conversion
When generating candidates in karukan-im/src/core/engine/conversion.rs (lines 374-382), the engine follows a strict three-tier lookup order:
- Learning cache (recent user inputs)
- Merged user dictionary (from
defaultuserdict.txtand custom files) - System dictionary (built-in vocabulary)
This hierarchy ensures that user-provided entries from defaultuserdict.txt or manually added files appear at the top of candidate lists, overriding system defaults when conflicts occur.
Runtime Dictionary Updates Without Restart
Karukan supports hot-reloading of dictionaries without requiring a process restart. The engine watches the user dictionary directory using the notify crate on Linux and macOS. When file changes are detected, the engine reloads dictionaries on the next key event, making new entries immediately available for conversion.
Practical Usage Examples
Programmatically Adding Entries
use std::fs::{self, OpenOptions};
use std::io::Write;
use karukan_im::config::Settings;
fn add_user_entry(reading: &str, candidate: &str) -> std::io::Result<()> {
let dir = Settings::user_dict_dir()?;
fs::create_dir_all(&dir)?;
let path = dir.join("custom_user_dict.txt");
let mut file = OpenOptions::new()
.create(true)
.append(true)
.open(path)?;
writeln!(file, "\"{}\":\"{}\"", reading, candidate)?;
Ok(())
}
Manual Dictionary Editing
# Create directory if needed and add entry
mkdir -p ~/.local/share/karukan-im/user_dicts
echo '"てすと":"テスト"' >> ~/.local/share/karukan-im/user_dicts/defaultuserdict.txt
The new candidate appears immediately at the next keystroke due to the filesystem watching mechanism implemented in the engine.
Summary
- Storage: Karukan stores user dictionaries in platform-specific directories returned by
Settings::user_dict_dir()inkarukan-im/src/config/settings.rs - Initialization: The bundled
defaultuserdict.txtautomatically copies to the user directory on first run viainitialize_user_dictsinkarukan-im/src/core/engine/init.rs - Format: Dictionary files use JSON format and load via
serde_json::from_reader, with deserialization errors logged but non-fatal to startup - Merging: Multiple dictionaries merge into a single
UserDictionaryusingUserDictionary::merge, consulted before the system dictionary during conversion - Hot-reloading: Filesystem watching via the
notifycrate enables real-time updates without restarting the IME
Frequently Asked Questions
What file format does Karukan use for user dictionaries?
Karukan expects user dictionary files to contain valid JSON where each line maps readings to candidates (for example, "reading":"candidate"). While the files use .txt extensions, they must parse as JSON when loaded via serde_json::from_reader in karukan-im/src/core/engine/init.rs.
Why isn't my custom dictionary entry showing up in candidates?
Check that your file resides in the correct platform-specific directory returned by Settings::user_dict_dir() (typically ~/.local/share/karukan-im/user_dicts on Linux). Ensure the file contains valid JSON—syntax errors will trigger warnings in logs but won't load entries. Also verify the entry doesn't conflict with the learning cache, which takes highest priority in karukan-im/src/core/engine/conversion.rs.
Can I use multiple user dictionary files simultaneously?
Yes. Karukan loads every .txt or .json file in the user dictionary directory and merges them using UserDictionary::merge. You can organize entries into separate files (for example, work_terms.txt and personal_names.txt) and place them all in the directory.
How do I reset to the default dictionary?
Delete the contents of your user dictionary directory (found via Settings::user_dict_dir()). On the next engine startup, initialize_user_dicts will detect the empty directory and recopy the bundled defaultuserdict.txt from karukan-im/user_dicts/ into your user directory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →