# FlexSearch Tokenization Modes: strict vs forward vs reverse vs full Trade‑offs

> Explore FlexSearch tokenization modes strict forward reverse and full Understand trade-offs in memory index size and partial matching to optimize your search performance.

- Repository: [Nextapps GmbH/flexsearch](https://github.com/nextapps-de/flexsearch)
- Tags: deep-dive
- Published: 2026-02-23

---

**FlexSearch offers four primary tokenization modes—`strict`, `forward`, `reverse` (bidirectional), and `full`—that trade memory consumption and index size against partial‑matching capabilities, with `strict` using minimal memory but supporting only exact matches, while `full` enables arbitrary substring search at the cost of quadratic memory growth.**

FlexSearch is a high‑performance full‑text search library that allows developers to configure how terms are split into tokens through different tokenization modes. Choosing between `strict`, `forward`, `reverse`, and `full` modes determines whether your index supports prefix matching, suffix matching, or arbitrary substring search, directly impacting memory consumption, index build time, and query performance according to the implementation in [`src/index.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index.js).

## What Are FlexSearch Tokenization Modes?

Tokenization modes determine how a single term is decomposed into index entries. In [`src/index.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index.js) at line 96, the library normalizes the `tokenize` option: `this.tokenize = options.tokenize || "strict"`. This default `strict` mode stores only the complete term, while other modes generate multiple tokens per term to enable partial matching.

The memory factor for each mode indicates how many tokens are generated relative to term length `n`. This factor directly correlates with RAM usage and index size.

## The Four Tokenization Modes Explained

### strict (exact) Mode

The `strict` tokenizer indexes only the complete, unmodified term. It generates exactly **1 token** per term regardless of length, yielding the smallest possible index footprint.

**Key characteristics:**
- **Memory factor:** `1` (minimal)
- **Partial matching:** None—queries must match the full term exactly (unless you lower `minlength` to enable prefix matching via query truncation)
- **Context Search:** ✅ **Only `strict` works with Context Search** (the window‑based relevance scoring). In [`src/index.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index.js) lines 103‑104, the constructor emits a warning when non‑`strict` tokenizers are used with context, and [`src/index/add.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index/add.js) lines 143‑146 enforce this guard internally.

Use `strict` when you need exact lookups, smallest memory footprint, or when utilizing Context Search.

### forward Mode

The `forward` tokenizer generates all prefixes of a term. For a term of length `n`, it creates `n` tokens (e.g., `"flex"` → `"f"`, `"fl"`, `"fle"`, `"flex"`).

**Key characteristics:**
- **Memory factor:** `n` (linear)
- **Partial matching:** **Prefix only**—typing `"fl"` matches `"flexsearch"`, but `"ex"` does not.
- **RTL support:** Supports right‑to‑left languages via the `rtl: true` option.

Use `forward` for autocomplete scenarios where users type the beginning of words. It offers the best balance between partial matching capability and memory efficiency.

### reverse (bidirectional) Mode

The `reverse` tokenizer (also aliased as `bidirectional`) indexes both prefixes and suffixes by storing the term forward and backward. For length `n`, it generates `2n − 1` tokens.

**Key characteristics:**
- **Memory factor:** `2n − 1` (roughly double `forward`)
- **Partial matching:** **Prefix and suffix**—`"sea"` matches `"search"` (suffix) and `"flex"` matches `"flexsearch"` (prefix).
- **Use case:** When users might search from the end of a word (e.g., searching `"arch"` to find `"search"`).

Use `reverse` when you need bidirectional matching but want to avoid the quadratic memory cost of `full` mode.

### full Mode

The `full` tokenizer indexes **every** consecutive substring of a term. For length `n`, it generates `n × (n − 1)` tokens (quadratic growth).

**Key characteristics:**
- **Memory factor:** `n(n‑1)` (quadratic—grows rapidly with term length)
- **Partial matching:** **Any substring**—`"exs"` matches `"flexsearch"` (middle substring).
- **Performance impact:** Slowest index build time and highest memory consumption.

Use `full` only when you need exhaustive substring search (e.g., searching inside long identifiers or DNA sequences) and can afford the memory cost.

## Memory and Performance Trade‑offs

The choice of tokenizer directly impacts three metrics: **memory usage**, **index build time**, and **query flexibility**.

| Mode | Memory Factor | Index Speed | Query Speed | Partial Match | Context Search |
|------|-------------|-------------|-------------|---------------|----------------|
| `strict` | `1` (minimal) | Fastest | Fastest | None | ✅ Yes |
| `forward` | `n` (linear) | Fast | Fast | Prefix | ❌ No |
| `reverse` | `2n‑1` (linear ×2) | Moderate | Moderate | Prefix + Suffix | ❌ No |
| `full` | `n(n‑1)` (quadratic) | Slowest | Slowest | Any substring | ❌ No |

**Key architectural constraints from the source code:**
- In [`src/index.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index.js) line 96, the tokenizer defaults to `"strict"` when not specified.
- Lines 103‑104 of [`src/index.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index.js) enforce that **Context Search** (window‑based relevance scoring) requires `tokenize: "strict"`. Using any other tokenizer with context enabled emits a console warning.
- The internal guard in [`src/index/add.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index/add.js) lines 143‑146 prevents non‑strict tokenizers from being used with context indexes.

## Code Examples: Testing Each Tokenizer

The following examples demonstrate the behavioral differences between tokenization modes using the FlexSearch `Index` class.

```javascript
import { Index } from "flexsearch";

/* 1. strict – exact term only */
const strictIdx = new Index({ tokenize: "strict" });
strictIdx.add(1, "flexsearch");
console.log(strictIdx.search("flex"));        // [] – no prefix match
console.log(strictIdx.search("flexsearch"));  // [1]

/* 2. forward – prefix autocomplete */
const forwardIdx = new Index({ tokenize: "forward" });
forwardIdx.add(1, "flexsearch");
console.log(forwardIdx.search("fl"));         // [1] – matches any prefix
console.log(forwardIdx.search("search"));     // [] – suffix not indexed

/* 3. reverse (bidirectional) – prefix + suffix */
const reverseIdx = new Index({ tokenize: "reverse" });
reverseIdx.add(1, "flexsearch");
console.log(reverseIdx.search("sea"));        // [1] – suffix match works
console.log(reverseIdx.search("flex"));       // [1] – prefix still works

/* 4. full – every substring */
const fullIdx = new Index({ tokenize: "full" });
fullIdx.add(1, "flexsearch");
console.log(fullIdx.search("exs"));           // [1] – inner substring matches
console.log(fullIdx.search("search"));        // [1] – also works

```

All examples run in Node.js or the browser after installing the package (`npm i flexsearch`).

## Summary

- **`strict`** mode offers the smallest memory footprint (1 token per term) and fastest performance, but only supports exact matches. It is the **only mode compatible with Context Search**, as enforced in [`src/index.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index.js) lines 103‑104.

- **`forward`** mode enables prefix matching with linear memory growth (`n` tokens), making it ideal for autocomplete scenarios where users type the beginning of words.

- **`reverse`** (bidirectional) mode doubles the memory cost (`2n‑1` tokens) to index both prefixes and suffixes, supporting searches from either end of a term.

- **`full`** mode generates every possible substring (`n(n‑1)` tokens), enabling exhaustive substring search at the cost of quadratic memory usage and slower index performance.

Choose `strict` for exact lookups and Context Search, `forward` for prefix autocomplete, `reverse` for bidirectional matching, and `full` only when exhaustive substring search justifies the memory overhead.

## Frequently Asked Questions

### What happens if I use Context Search with a non-strict tokenizer?

The library emits a warning in [`src/index.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index.js) lines 103‑104 and internally prevents the combination. Context Search (window‑based relevance scoring) requires whole terms to calculate positional relevance, which only the `strict` tokenizer provides. If you need partial matching alongside context awareness, you must create separate indexes.

### How much memory does each tokenizer use for a 10‑character term?

For a term of length `n = 10`:
- `strict`: 1 token (baseline)
- `forward`: 10 tokens (linear)
- `reverse`: 19 tokens (`2n‑1`)
- `full`: 90 tokens (`n(n‑1)`)

The `full` tokenizer generates nearly 100 times more index entries than `strict` for this term length, explaining its high memory cost.

### Can I switch tokenizers on an existing index?

No. The tokenizer must be defined when the index is instantiated via `new Index({ tokenize: "forward" })` as seen in [`src/index.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/index.js) line 96. Changing the tokenizer requires rebuilding the index from source documents, as the tokenization strategy determines how terms are decomposed and stored internally.

### When should I use reverse instead of full?

Use `reverse` when you need to match prefixes **and** suffixes but not arbitrary substrings in the middle of words. For example, searching for `"arch"` to find `"search"` works with `reverse`, while `"exs"` to find `"flexsearch"` requires `full`. The `reverse` mode uses roughly linear memory (`2n`) compared to the quadratic cost of `full`, making it feasible for large datasets where `full` would exhaust RAM.