# How Karukan Manages and Loads User Dictionaries from defaultuserdict.txt

> Learn how Karukan loads and manages user dictionaries from defaultuserdict.txt. Discover how Karukan merges and prioritizes your custom dictionaries for enhanced candidate generation.

- Repository: [Hitoshi Togasaki/karukan](https://github.com/togatoga/karukan)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Karukan loads user dictionaries as JSON files from a platform-specific directory, automatically copying the bundled [`defaultuserdict.txt`](https://github.com/togatoga/karukan/blob/main/defaultuserdict.txt) on first startup, then merges them into a single in-memory dictionary that takes priority over the system dictionary during candidate generation.**

Karukan treats every file in the user-dictionary directory as a separate JSON-encoded user dictionary capable of overriding system entries. The Rust-based input method engine handles these dictionaries through a dedicated initialization process in [`karukan-im/src/core/engine/init.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/init.rs), ensuring user customizations persist across sessions while appearing first in conversion candidates.

## User Dictionary Storage Location

Karukan determines the storage location through `Settings::user_dict_dir()` defined in [`karukan-im/src/config/settings.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/config/settings.rs). This method returns platform-specific paths following XDG Base Directory and macOS application support conventions:

- **Linux**: `$XDG_DATA_HOME/karukan-im/user_dicts` (defaults to `~/.local/share/karukan-im/user_dicts`)
- **macOS**: `$HOME/Library/Application Support/com.karukan.karukan-im/user_dicts`

The engine creates this directory automatically using `std::fs::create_dir_all` if it does not exist during the initialization phase.

## The Default Dictionary File

The repository ships a starter dictionary at [`karukan-im/user_dicts/defaultuserdict.txt`](https://github.com/togatoga/karukan/blob/main/karukan-im/user_dicts/defaultuserdict.txt). This plain-text file contains JSON-formatted mappings where each line represents a reading-to-candidate pair (for example, `てすと,`).

During engine startup, the `initialize_user_dicts` function copies this bundled file into the user's dictionary directory if the directory is missing or empty. This ensures first-time users have a functional dictionary without manual configuration, using `std::fs::copy` to place the file before proceeding with the loading sequence.

## Loading and Merging Process

The `initialize_user_dicts` function in [`karukan-im/src/core/engine/init.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/init.rs) (lines 216-287) orchestrates dictionary loading through a multi-step pipeline:

### Directory Initialization and Default Copying

The function first resolves the path via `Settings::user_dict_dir()`. If the directory is missing, it creates the hierarchy and copies the bundled [`defaultuserdict.txt`](https://github.com/togatoga/karukan/blob/main/defaultuserdict.txt) into the location. This ensures the engine always has a baseline dictionary before attempting to load user customizations.

### File Discovery and JSON Deserialization

The engine scans the directory using `fs::read_dir`, processing entries with `.txt` or `.json` extensions. For each file, it opens a reader and deserializes the content using `serde_json::from_reader` into `UserDictionary` structs. Errors during deserialization are logged via `warn!` macros but do not abort startup, allowing the engine to continue with valid dictionaries while skipping malformed files.

### Dictionary Aggregation

All successfully loaded dictionaries merge using `UserDictionary::merge`, producing a single consolidated dictionary stored in the `Engine::user_dict` field. This merged structure resides in memory and serves as the authoritative user dictionary for the entire session, consulted before the system dictionary during text conversion.

## Search Priority During Conversion

When generating candidates in [`karukan-im/src/core/engine/conversion.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/conversion.rs) (lines 374-382), the engine follows a strict three-tier lookup order:

1. **Learning cache** (recent user inputs)
2. **Merged user dictionary** (from [`defaultuserdict.txt`](https://github.com/togatoga/karukan/blob/main/defaultuserdict.txt) and custom files)
3. **System dictionary** (built-in vocabulary)

This hierarchy ensures that user-provided entries from [`defaultuserdict.txt`](https://github.com/togatoga/karukan/blob/main/defaultuserdict.txt) or manually added files appear at the top of candidate lists, overriding system defaults when conflicts occur.

## Runtime Dictionary Updates Without Restart

Karukan supports hot-reloading of dictionaries without requiring a process restart. The engine watches the user dictionary directory using the `notify` crate on Linux and macOS. When file changes are detected, the engine reloads dictionaries on the next key event, making new entries immediately available for conversion.

## Practical Usage Examples

### Programmatically Adding Entries

```rust
use std::fs::{self, OpenOptions};
use std::io::Write;
use karukan_im::config::Settings;

fn add_user_entry(reading: &str, candidate: &str) -> std::io::Result<()> {
    let dir = Settings::user_dict_dir()?;
    fs::create_dir_all(&dir)?;                      
    let path = dir.join("custom_user_dict.txt");    
    let mut file = OpenOptions::new()
        .create(true)
        .append(true)
        .open(path)?;
    writeln!(file, "\"{}\":\"{}\"", reading, candidate)?;
    Ok(())
}

```

### Manual Dictionary Editing

```bash

# Create directory if needed and add entry

mkdir -p ~/.local/share/karukan-im/user_dicts
echo '"てすと":"テスト"' >> ~/.local/share/karukan-im/user_dicts/defaultuserdict.txt

```

The new candidate appears immediately at the next keystroke due to the filesystem watching mechanism implemented in the engine.

## Summary

- **Storage**: Karukan stores user dictionaries in platform-specific directories returned by `Settings::user_dict_dir()` in [`karukan-im/src/config/settings.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/config/settings.rs)
- **Initialization**: The bundled [`defaultuserdict.txt`](https://github.com/togatoga/karukan/blob/main/defaultuserdict.txt) automatically copies to the user directory on first run via `initialize_user_dicts` in [`karukan-im/src/core/engine/init.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/init.rs)
- **Format**: Dictionary files use JSON format and load via `serde_json::from_reader`, with deserialization errors logged but non-fatal to startup
- **Merging**: Multiple dictionaries merge into a single `UserDictionary` using `UserDictionary::merge`, consulted before the system dictionary during conversion
- **Hot-reloading**: Filesystem watching via the `notify` crate enables real-time updates without restarting the IME

## Frequently Asked Questions

### What file format does Karukan use for user dictionaries?

Karukan expects user dictionary files to contain valid JSON where each line maps readings to candidates (for example, `"reading":"candidate"`). While the files use `.txt` extensions, they must parse as JSON when loaded via `serde_json::from_reader` in [`karukan-im/src/core/engine/init.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/init.rs).

### Why isn't my custom dictionary entry showing up in candidates?

Check that your file resides in the correct platform-specific directory returned by `Settings::user_dict_dir()` (typically `~/.local/share/karukan-im/user_dicts` on Linux). Ensure the file contains valid JSON—syntax errors will trigger warnings in logs but won't load entries. Also verify the entry doesn't conflict with the learning cache, which takes highest priority in [`karukan-im/src/core/engine/conversion.rs`](https://github.com/togatoga/karukan/blob/main/karukan-im/src/core/engine/conversion.rs).

### Can I use multiple user dictionary files simultaneously?

Yes. Karukan loads every `.txt` or `.json` file in the user dictionary directory and merges them using `UserDictionary::merge`. You can organize entries into separate files (for example, [`work_terms.txt`](https://github.com/togatoga/karukan/blob/main/work_terms.txt) and [`personal_names.txt`](https://github.com/togatoga/karukan/blob/main/personal_names.txt)) and place them all in the directory.

### How do I reset to the default dictionary?

Delete the contents of your user dictionary directory (found via `Settings::user_dict_dir()`). On the next engine startup, `initialize_user_dicts` will detect the empty directory and recopy the bundled [`defaultuserdict.txt`](https://github.com/togatoga/karukan/blob/main/defaultuserdict.txt) from `karukan-im/user_dicts/` into your user directory.