FlexSearch Tokenization Modes: strict vs forward vs reverse vs full Trade‑offs

FlexSearch offers four primary tokenization modes—strict, forward, reverse (bidirectional), and full—that trade memory consumption and index size against partial‑matching capabilities, with strict using minimal memory but supporting only exact matches, while full enables arbitrary substring search at the cost of quadratic memory growth.

FlexSearch is a high‑performance full‑text search library that allows developers to configure how terms are split into tokens through different tokenization modes. Choosing between strict, forward, reverse, and full modes determines whether your index supports prefix matching, suffix matching, or arbitrary substring search, directly impacting memory consumption, index build time, and query performance according to the implementation in src/index.js.

What Are FlexSearch Tokenization Modes?

Tokenization modes determine how a single term is decomposed into index entries. In src/index.js at line 96, the library normalizes the tokenize option: this.tokenize = options.tokenize || "strict". This default strict mode stores only the complete term, while other modes generate multiple tokens per term to enable partial matching.

The memory factor for each mode indicates how many tokens are generated relative to term length n. This factor directly correlates with RAM usage and index size.

The Four Tokenization Modes Explained

strict (exact) Mode

The strict tokenizer indexes only the complete, unmodified term. It generates exactly 1 token per term regardless of length, yielding the smallest possible index footprint.

Key characteristics:

  • Memory factor: 1 (minimal)
  • Partial matching: None—queries must match the full term exactly (unless you lower minlength to enable prefix matching via query truncation)
  • Context Search: ✅ Only strict works with Context Search (the window‑based relevance scoring). In src/index.js lines 103‑104, the constructor emits a warning when non‑strict tokenizers are used with context, and src/index/add.js lines 143‑146 enforce this guard internally.

Use strict when you need exact lookups, smallest memory footprint, or when utilizing Context Search.

forward Mode

The forward tokenizer generates all prefixes of a term. For a term of length n, it creates n tokens (e.g., "flex" → "f", "fl", "fle", "flex").

Key characteristics:

  • Memory factor: n (linear)
  • Partial matching: Prefix only—typing "fl" matches "flexsearch", but "ex" does not.
  • RTL support: Supports right‑to‑left languages via the rtl: true option.

Use forward for autocomplete scenarios where users type the beginning of words. It offers the best balance between partial matching capability and memory efficiency.

reverse (bidirectional) Mode

The reverse tokenizer (also aliased as bidirectional) indexes both prefixes and suffixes by storing the term forward and backward. For length n, it generates 2n − 1 tokens.

Key characteristics:

  • Memory factor: 2n − 1 (roughly double forward)
  • Partial matching: Prefix and suffix—"sea" matches "search" (suffix) and "flex" matches "flexsearch" (prefix).
  • Use case: When users might search from the end of a word (e.g., searching "arch" to find "search").

Use reverse when you need bidirectional matching but want to avoid the quadratic memory cost of full mode.

full Mode

The full tokenizer indexes every consecutive substring of a term. For length n, it generates n × (n − 1) tokens (quadratic growth).

Key characteristics:

  • Memory factor: n(n‑1) (quadratic—grows rapidly with term length)
  • Partial matching: Any substring—"exs" matches "flexsearch" (middle substring).
  • Performance impact: Slowest index build time and highest memory consumption.

Use full only when you need exhaustive substring search (e.g., searching inside long identifiers or DNA sequences) and can afford the memory cost.

Memory and Performance Trade‑offs

The choice of tokenizer directly impacts three metrics: memory usage, index build time, and query flexibility.

Mode Memory Factor Index Speed Query Speed Partial Match Context Search
strict 1 (minimal) Fastest Fastest None ✅ Yes
forward n (linear) Fast Fast Prefix ❌ No
reverse 2n‑1 (linear ×2) Moderate Moderate Prefix + Suffix ❌ No
full n(n‑1) (quadratic) Slowest Slowest Any substring ❌ No

Key architectural constraints from the source code:

  • In src/index.js line 96, the tokenizer defaults to "strict" when not specified.
  • Lines 103‑104 of src/index.js enforce that Context Search (window‑based relevance scoring) requires tokenize: "strict". Using any other tokenizer with context enabled emits a console warning.
  • The internal guard in src/index/add.js lines 143‑146 prevents non‑strict tokenizers from being used with context indexes.

Code Examples: Testing Each Tokenizer

The following examples demonstrate the behavioral differences between tokenization modes using the FlexSearch Index class.

import { Index } from "flexsearch";

/* 1. strict – exact term only */
const strictIdx = new Index({ tokenize: "strict" });
strictIdx.add(1, "flexsearch");
console.log(strictIdx.search("flex"));        // [] – no prefix match
console.log(strictIdx.search("flexsearch"));  // [1]

/* 2. forward – prefix autocomplete */
const forwardIdx = new Index({ tokenize: "forward" });
forwardIdx.add(1, "flexsearch");
console.log(forwardIdx.search("fl"));         // [1] – matches any prefix
console.log(forwardIdx.search("search"));     // [] – suffix not indexed

/* 3. reverse (bidirectional) – prefix + suffix */
const reverseIdx = new Index({ tokenize: "reverse" });
reverseIdx.add(1, "flexsearch");
console.log(reverseIdx.search("sea"));        // [1] – suffix match works
console.log(reverseIdx.search("flex"));       // [1] – prefix still works

/* 4. full – every substring */
const fullIdx = new Index({ tokenize: "full" });
fullIdx.add(1, "flexsearch");
console.log(fullIdx.search("exs"));           // [1] – inner substring matches
console.log(fullIdx.search("search"));        // [1] – also works

All examples run in Node.js or the browser after installing the package (npm i flexsearch).

Summary

  • strict mode offers the smallest memory footprint (1 token per term) and fastest performance, but only supports exact matches. It is the only mode compatible with Context Search, as enforced in src/index.js lines 103‑104.

  • forward mode enables prefix matching with linear memory growth (n tokens), making it ideal for autocomplete scenarios where users type the beginning of words.

  • reverse (bidirectional) mode doubles the memory cost (2n‑1 tokens) to index both prefixes and suffixes, supporting searches from either end of a term.

  • full mode generates every possible substring (n(n‑1) tokens), enabling exhaustive substring search at the cost of quadratic memory usage and slower index performance.

Choose strict for exact lookups and Context Search, forward for prefix autocomplete, reverse for bidirectional matching, and full only when exhaustive substring search justifies the memory overhead.

Frequently Asked Questions

What happens if I use Context Search with a non-strict tokenizer?

The library emits a warning in src/index.js lines 103‑104 and internally prevents the combination. Context Search (window‑based relevance scoring) requires whole terms to calculate positional relevance, which only the strict tokenizer provides. If you need partial matching alongside context awareness, you must create separate indexes.

How much memory does each tokenizer use for a 10‑character term?

For a term of length n = 10:

  • strict: 1 token (baseline)
  • forward: 10 tokens (linear)
  • reverse: 19 tokens (2n‑1)
  • full: 90 tokens (n(n‑1))

The full tokenizer generates nearly 100 times more index entries than strict for this term length, explaining its high memory cost.

Can I switch tokenizers on an existing index?

No. The tokenizer must be defined when the index is instantiated via new Index({ tokenize: "forward" }) as seen in src/index.js line 96. Changing the tokenizer requires rebuilding the index from source documents, as the tokenization strategy determines how terms are decomposed and stored internally.

When should I use reverse instead of full?

Use reverse when you need to match prefixes and suffixes but not arbitrary substrings in the middle of words. For example, searching for "arch" to find "search" works with reverse, while "exs" to find "flexsearch" requires full. The reverse mode uses roughly linear memory (2n) compared to the quadratic cost of full, making it feasible for large datasets where full would exhaust RAM.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →