How to Use Custom Encoders with Stemmers and Filters for Language-Specific Search in FlexSearch

FlexSearch processes text through an Encoder class that applies a configurable pipeline of stemmers, filters, mappers, and replacers, enabling language-specific search by passing custom EncoderOptions to new Encoder() or extending built-in presets from flexsearch/lang/*.

FlexSearch is a high-performance full-text search library that handles tokenization, normalization, and fuzzy matching through its Encoder class. To implement language-specific search with custom text processing, you configure the encoder pipeline with stemmers for word reduction, filters for stop-word removal, and character mappers for normalization. This guide demonstrates how to build custom encoders using the FlexSearch source code architecture found in nextapps-de/flexsearch.

Understanding the FlexSearch Encoder Pipeline

The Encoder class defined in src/encoder.js merges every supplied option into an internal processing pipeline. When an Index is instantiated, FlexSearch builds an encoder instance from the options.encoder (or options.encode) field in src/index.js (lines 76-89):

const encoder = SUPPORT_CHARSET && is_string(options.encoder)
    ? Charset[options.encoder]                     // built‑in charset preset
    : options.encode || options.encoder || (
        SUPPORT_ENCODER ? {} : fallback_encoder
      );
this.encoder = encoder.encode
    ? encoder
    : typeof encoder === "object"
        ? new Encoder(/** @type {EncoderOptions} */ (encoder))
        : { encode: encoder };

The encoder pipeline processes text through sequential steps: normalize → prepare → split → mapper → matcher → stemmer → filter → finalize. Each step corresponds to an option in EncoderOptions defined in src/type.js (lines 14-38):

Step Option Purpose
Stemming stemmer Map<string,string> that replaces term endings (e.g., ly → '').
Filtering filter Set<string> or function that removes stop-words or applies custom logic.
Character Mapping mapper Map<string,string> that maps single characters (e.g., é → e).
String Matching matcher Map<string,string> that replaces whole strings before stemming.
Regex Replacement replacer Array<[RegExp, string]> for pattern-based replacements after mapping.
Preparation prepare function(str) for custom pre-split processing (e.g., HTML entity replacement).
Finalization finalize function(array) for post-split processing (e.g., dropping short terms).

Creating a Basic Custom Encoder with Stemmers and Filters

To implement language-specific processing, instantiate Encoder with a configuration object containing your custom stemmers and filters. The Encoder.prototype.assign method in src/encoder.js (lines 93-124) merges these options into the pipeline.

import FlexSearch from "flexsearch";
import { Encoder } from "flexsearch";

// Create a custom encoder with stemming and stop-word filtering
const myEncoder = new Encoder({
  // Replace English adverb suffixes
  stemmer: new Map([["ly", ""]]),

  // Filter out common stop-words
  filter: new Set(["and", "the", "of"]),

  // Optional pre-process step (e.g., replace ampersand)
  prepare: str => str.replace(/&/g, " and "),

  // Optional post-process step (drop very short terms)
  finalize: terms => terms.filter(t => t.length > 2)
});

// Attach encoder to an index
const idx = new FlexSearch.Index({
  encoder: myEncoder,
  tokenize: "strict",   // Required for context search
  resolution: 9
});

The addStemmer and addFilter helper methods (defined in src/encoder.js, lines 74-88 and 83-92) allow you to extend an existing encoder instance without reconstructing the configuration object.

Extending Built-in Language Presets

FlexSearch provides language-specific presets in src/lang/en.js (and similar files for German, French, etc.) that export pre-filled EncoderOptions objects. You can extend these presets by passing them to the Encoder constructor along with override options. The constructor merges multiple option objects sequentially (see src/encoder.js, lines 80-84).

import FlexSearch from "flexsearch";
import { Encoder } from "flexsearch";
import EnglishPreset from "flexsearch/lang/en";

// Start from the English preset and modify specific options
const encoder = new Encoder(
  EnglishPreset,               // Built-in stemmer, filter, mapper, etc.
  { filter: false }            // Disable the preset's default stop-word filter
);

// Add domain-specific terms to the filter
encoder.addFilter("example");

// Create the index with the customized encoder
const idx = new FlexSearch.Index({
  encoder,
  tokenize: "strict"
});

This approach ensures you inherit the optimized stemming rules and character mappings from the language preset while customizing the filtering logic for your specific domain.

Building CJK-Aware Encoders with Character Mapping

For languages that require character-level processing, such as Chinese, Japanese, and Korean (CJK), FlexSearch provides charset presets in src/charset.js. You can combine these with custom mappers and replacers to handle full-width punctuation and ideographic spaces.

import FlexSearch from "flexsearch";
import { Encoder, Charset } from "flexsearch";

// Start from the CJK charset (character-by-character split)
const encoder = new Encoder(Charset.CJK)
  .addMapper(",", ",")            // Map full-width comma to ASCII comma
  .addReplacer(/[\u3000]/g, " ")  // Replace IDEOGRAPHIC SPACE with ordinary space
  .addFilter(term => term.length > 1); // Keep only multi-character terms

const idx = new FlexSearch.Index({
  encoder,
  tokenize: "strict"
});

The addMapper and addReplacer methods (implemented in src/encoder.js, lines 83-106 and 115-124) allow you to inject character-level transformations into the pipeline after the initial split but before stemming and filtering.

How the Encoder Integrates with Index Operations

The encoder is attached to the index instance during construction in src/index.js (lines 76-89). All subsequent search and add operations internally call this.encoder.encode(...) to ensure consistent text processing. This synchronization is essential for language-aware search, as it guarantees that query terms undergo the same stemming, filtering, and normalization as indexed content.

According to the source code in src/index/search.js (line 85) and src/index/add.js (line 40), the encoder's encode method is invoked immediately before indexing or matching, ensuring the pipeline is applied uniformly across both operations.

Summary

  • FlexSearch encoders process text through a configurable pipeline defined in src/encoder.js, handling normalization, stemming, filtering, and finalization.
  • Custom encoders are created by instantiating new Encoder() with an EncoderOptions object containing stemmer (Map), filter (Set/function), mapper (Map), and other pipeline steps.
  • Language presets in src/lang/en.js (and similar) provide pre-configured encoders for specific languages that you can extend or override.
  • Charset presets like Charset.CJK in src/charset.js enable character-level processing for CJK languages, combinable with custom mappers and replacers.
  • Integration occurs in src/index.js, where the encoder is attached to the index and automatically invoked during add() and search() operations to ensure consistent language processing.

Frequently Asked Questions

How do I disable the default stop-word filter in a language preset?

Pass { filter: false } as a second argument to the Encoder constructor when extending a preset. According to src/encoder.js (lines 80-84), the constructor merges multiple option objects sequentially, allowing you to override specific properties while inheriting the preset's stemmer and mapper configurations.

Can I use a function instead of a Set for the filter option?

Yes, the filter option accepts either a Set<string> for exact string matching or a function(term: string): boolean for custom logic. When a function is provided, the encoder invokes it for each term during the filter stage of the pipeline defined in src/encoder.js.

What is the difference between mapper and matcher in FlexSearch encoders?

The mapper option processes single characters (e.g., mapping é to e), while the matcher replaces whole strings before stemming occurs (e.g., mapping "New York" to "newyork"). According to src/type.js, both accept Map<string,string> but operate at different pipeline stages.

How do I ensure my custom encoder uses the same processing for indexing and searching?

Attach your encoder instance to the Index constructor via the encoder option. As implemented in src/index.js (lines 76-89), the index stores the encoder reference and automatically invokes this.encoder.encode() during both add() operations (src/index/add.js, line 40) and search() operations (src/index/search.js, line 85), ensuring identical text processing for both stages.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →