# FlexSearch Charset.LatinBalance vs LatinAdvanced vs LatinExtra: A Technical Comparison

> Explore FlexSearch Charset.LatinBalance LatinAdvanced and LatinExtra presets. Understand their normalization differences for better search accuracy and performance.

- Repository: [Nextapps GmbH/flexsearch](https://github.com/nextapps-de/flexsearch)
- Tags: deep-dive
- Published: 2026-02-23

---

**The three Latin charset presets in FlexSearch differ primarily in their normalization aggressiveness: LatinBalance applies only soundex phonetic mapping, LatinAdvanced adds matcher/replacer rules for digraphs and silent letters, and LatinExtra further compacts tokens by removing non-initial vowels.**

FlexSearch, the high-performance full-text search library by nextapps-de/flexsearch, provides three distinct Latin charset presets that control how text is normalized before indexing. Understanding the differences between **Charset.LatinBalance**, **Charset.LatinAdvanced**, and **Charset.LatinExtra** is crucial for optimizing search relevance and index size in your applications.

## What Are FlexSearch Latin Charset Presets?

Charset presets in FlexSearch define the **encoder** configuration that transforms raw text into normalized tokens. All three Latin presets share a common foundation: a **soundex-based mapper** that collapses similar-sounding consonants (e.g., "ph" → "f", "c" → "k"). However, they diverge in additional processing layers that affect token granularity and fuzzy matching capabilities.

## Technical Differences Between LatinBalance, LatinAdvanced, and LatinExtra

### LatinBalance: Soundex-Only Normalization

**Charset.LatinBalance** provides the most conservative normalization. It applies only the core soundex mapper defined in [`src/charset/latin/balance.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/balance.js) without additional matcher or replacer rules.

This preset performs basic phonetic grouping—converting "ph" to "f" and normalizing consonant clusters—but preserves most of the original token structure. It is the fastest option and produces the largest tokens, making it suitable when you need moderate fuzzy matching without aggressive text compaction.

### LatinAdvanced: Matcher and Replacer Rules

**Charset.LatinAdvanced** extends LatinBalance by adding two critical components defined in [`src/charset/latin/advanced.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/advanced.js): a **matcher** map and **replacer** regex rules.

The **matcher** handles specific character digraphs and common phonetic substitutions:
- `ae → a`
- `oe → o` 
- `sh → s`
- `kh → k`
- `th → t`
- `ph → f`
- `pf → f`

The **replacer** applies regex transformations that remove silent "h" after non-a/e/o letters, collapse duplicate letters, and handle other edge cases that produce cleaner tokens. This results in more aggressive normalization than LatinBalance—"pharmacy" becomes "farmcy" rather than "pharmcy".

### LatinExtra: Aggressive Vowel Compaction

**Charset.LatinExtra** inherits all rules from LatinAdvanced and adds a final **compact** regex that removes non-initial vowels. Defined in [`src/charset/latin/extra.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/extra.js), this preset applies the pattern `/(?!^)[aeo]/g` to strip "a", "e", and "o" characters that appear after the first position in a token.

This produces the most aggressive normalization possible. For example, "pharmacy" collapses to "fmrc", "shampoo" becomes "smpo", and "aeon" reduces to "n". This preset is ideal for high-fuzziness scenarios where you want "movie" and "mv" to match, but it sacrifices token readability and may increase false positives.

## Source Code Implementation

The preset hierarchy is implemented across four key files in the nextapps-de/flexsearch repository:

- **[`src/charset/latin/balance.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/balance.js)** – Defines the base soundex mapper used by all three presets【/cache/repos/github.com/nextapps-de/flexsearch/master/src/charset/latin/balance.js#L1-L48】
- **[`src/charset/latin/advanced.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/advanced.js)** – Adds the matcher map and replacer regex rules【/cache/repos/github.com/nextapps-de/flexsearch/master/src/charset/latin/advanced.js#L4-L33】  
- **[`src/charset/latin/extra.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/extra.js)** – Extends Advanced with the compact vowel-removal regex【/cache/repos/github.com/nextapps-de/flexsearch/master/src/charset/latin/extra.js#L5-L16】
- **[`src/charset.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset.js)** – Exports all three presets for public consumption【/cache/repos/github.com/nextapps-de/flexsearch/master/src/charset.js#L14-L18】

All presets are passed to the FlexSearch index via the `encoder` configuration option.

## Practical Code Examples

The following example demonstrates how each preset transforms identical input strings:

```javascript
import FlexSearch from "flexsearch";

// Helper to encode text using a specific preset
function encode(text, preset) {
  const encoder = new FlexSearch.Encoder(preset);
  return encoder.encode(text);
}

const testWords = ["pharmacy", "pharaoh", "thorough", "shampoo", "aeon", "aeolian"];

// LatinBalance: Soundex only
console.log("LatinBalance:");
testWords.forEach(w => console.log(`${w} → ${encode(w, FlexSearch.Charset.LatinBalance)}`));

// LatinAdvanced: Soundex + matcher/replacer
console.log("\nLatinAdvanced:");
testWords.forEach(w => console.log(`${w} → ${encode(w, FlexSearch.Charset.LatinAdvanced)}`));

// LatinExtra: Advanced + vowel compaction
console.log("\nLatinExtra:");
testWords.forEach(w => console.log(`${w} → ${encode(w, FlexSearch.Charset.LatinExtra)}`));

```

**Typical output:**

```

LatinBalance:
pharmacy → pharmcy
pharaoh → pharoh
thorough → thorough
shampoo → shampoo
aeon → aen
aeolian → aeolian

LatinAdvanced:
pharmacy → farmcy
pharaoh → faroh
thorough → toroug
shampoo → sampo
aeon → aen
aeolian → aelian

LatinExtra:
pharmacy → fmrc
pharaoh → farh
thorough → torug
shampoo → smpo
aeon → n
aeolian → lian

```

Notice how **LatinAdvanced** collapses digraphs like "ph" → "f" and "th" → "t", while **LatinExtra** further strips vowels to create highly compact tokens.

## When to Use Each Preset

Choose your preset based on the trade-off between search fuzziness and precision:

- **Use LatinBalance** when you need fast indexing with basic phonetic normalization. Best for autocomplete fields where exact prefix matching matters but you still want "phone" and "fone" to match.

- **Use LatinAdvanced** when dealing with multilingual Latin text containing digraphs (German "pf", English "ph", Scandinavian "ae") or when you want to normalize silent letters. Ideal for product catalogs and document search.

- **Use LatinExtra** for maximum fuzzy matching where recall is more important than precision. Perfect for typo-tolerant search, short query matching, or mobile input where users skip vowels, but expect higher false-positive rates.

## Summary

- **Charset.LatinBalance** applies only the base soundex mapper from [`src/charset/latin/balance.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/balance.js), providing the fastest normalization with minimal token alteration.

- **Charset.LatinAdvanced** extends Balance with matcher rules (digraph substitution) and replacer regexes (silent letter removal) defined in [`src/charset/latin/advanced.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/advanced.js).

- **Charset.LatinExtra** inherits all Advanced rules and adds aggressive vowel compaction via [`src/charset/latin/extra.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/extra.js), removing non-initial "a", "e", and "o" characters for maximum fuzzy matching.

All three presets are exported from [`src/charset.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset.js) and configured via the `encoder` option when creating a FlexSearch index.

## Frequently Asked Questions

### What is the performance impact of using LatinExtra versus LatinBalance?

**LatinExtra requires more regex operations per token** due to the additional vowel-compaction step, making it slightly slower than LatinBalance during indexing. However, the difference is typically negligible for most applications unless you are processing millions of documents. LatinBalance remains the fastest option because it applies only the soundex mapper without additional string replacements.

### Can I customize the matcher rules in LatinAdvanced?

**Yes, you can create custom encoder configurations** by importing the base components from [`src/charset/latin/advanced.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/charset/latin/advanced.js) and modifying the matcher map or replacer regexes. The FlexSearch encoder system allows you to compose your own preset by combining the soundex mapper with custom replacement rules, though you must ensure your modifications maintain the expected input/output format for the encoder pipeline.

### When should I avoid using LatinExtra?

**Avoid LatinExtra when precision is critical** or when searching short tokens where vowel removal creates ambiguity. For example, "movie" and "move" both reduce to "mv" under LatinExtra, causing false positives in semantic search. Additionally, if your content contains many acronyms or initialisms (like "NASA" or "FBI"), LatinExtra may over-normalize them and reduce search accuracy. Use LatinBalance or LatinAdvanced instead for these scenarios.

### How do I switch between presets in an existing FlexSearch index?

**You must reindex your data when changing charset presets** because FlexSearch encodes tokens at index time using the specified encoder. To switch presets, create a new index instance with the desired `encoder` option (e.g., `encoder: FlexSearch.Charset.LatinAdvanced`), then re-add all documents. There is no runtime migration path for existing indexed tokens because each preset produces fundamentally different token representations.