Common Pitfalls When Indexing Nested JSON Documents with the FlexSearch Document API
Most indexing errors in FlexSearch occur when developers use dot notation instead of colons for nested paths, omit the [] marker for arrays, or fail to remove documents before re-indexing fields that no longer exist.
The FlexSearch Document API provides a powerful way to index arbitrarily structured JSON objects by mapping properties to searchable fields. However, the internal tree-parsing mechanism—implemented in src/document.js—relies on strict descriptor syntax that differs from standard JavaScript object notation. Understanding how the parse_tree function processes colons and array markers is essential to avoid silent indexing failures when working with nested data structures.
How FlexSearch Processes Nested Documents
Under the hood, a Document instance is a thin wrapper that creates one or more ordinary Index instances for each field you declare. The wrapper parses the document descriptor into a tree that drives recursive extraction of values.
Descriptor to Tree Conversion
When you instantiate a Document, the constructor calls parse_tree (lines 387‑410 in src/document.js) to transform your field paths into an internal tree structure. The function splits each field name on the colon character (:), treating each segment as a level in the tree. If a segment ends with ] (for example, tags[]), the corresponding entry in a marker array is set to true (lines 389‑401). This marker tells the indexer that the value is an array whose items must be indexed separately.
Recursive Value Extraction
The add_index function in src/document/add.js (lines 24‑62) walks the JSON object recursively according to the tree structure. When it reaches the final segment of a path, the value is handed to the underlying Index via index.add. If the marker for that segment is true and the value is an array, each element is indexed independently (lines 34‑41).
Deep Array Handling
When an array appears before the leaf field—for example, authors[].name—the recursion does not advance the tree position. Instead, the array is iterated, and the same tree level is processed for each item (lines 50‑55 in src/document/add.js). This allows indexing of values inside nested arrays without consuming the tree segments prematurely.
Common Pitfalls and How to Avoid Them
Using Dot Notation Instead of Colon Separators
The most frequent mistake is using JavaScript dot notation (author.name) instead of the required colon syntax (author:name). The parse_tree function only splits paths on :, so a dot remains part of the key name, causing the field to be silently skipped during indexing.
// Wrong: uses dot notation, field is never indexed
index.add({author:{name:'Bob'}}, {
document:{id:'id', index:'author.name'}
});
// Correct: uses colon separator
index.add({author:{name:'Bob'}}, {
document:{id:'id', index:'author:name'}
});
Omitting the Array Marker Brackets
Without the [] suffix, FlexSearch treats an array as a single string value (effectively calling tags.join(' ')), which prevents individual element matching and highlighting.
// Wrong: indexes entire array as one token
index.add(doc, {document:{id:'id', index:'tags'}});
// Correct: each array item indexed separately
index.add(doc, {document:{id:'id', index:'tags[]'}});
Incorrect Syntax for Deeply Nested Arrays
When paths contain multiple arrays—such as authors[].books[].title—you must append [] to every array segment. Omitting a marker causes the recursion to stop early, indexing only the first array level.
// Correct descriptor for deep arrays
const idx = new Document({
document:{
id:'id',
index:['authors[].books[].title']
}
});
Custom Extractors Returning Undefined
When using a custom function to transform data, returning undefined or omitting a return value causes add_index to skip the field entirely (see the if (!tree) continue check at lines 75‑78 in src/document/add.js).
// Wrong: may return undefined
custom: doc => doc.title
// Correct: always return a value or empty array
custom: doc => doc.title || null
// For arrays:
custom: doc => (doc.tags || "").split(",")
Confusing Index Fields with Store Fields
The storetree is built only when the store option is truthy (lines 84‑88 in src/document.js). If you attempt to enrich search results with data from a field that is only indexed but not stored, the enrichment will fail silently.
// Enable storing for enrichment
const idx = new Document({
document:{
id:'id',
index:['content'],
store:['content', 'title'] // explicitly store these
}
});
Mismatched ID Path Syntax
The id field uses the same colon-based path parsing as index fields (lines 68‑69 in src/document.js). Using dot notation or an incorrect path causes the document to be registered under an undefined key, making it unsearchable.
// Wrong: dot notation
document:{ id:'meta.id' }
// Correct: colon separator
document:{ id:'meta:id' }
Performance Issues Without Fast Update
When indexing high-cardinality tag fields, the internal register (this.reg) grows rapidly. Without fastupdate:true, each addition triggers a full scan of the register, causing severe performance degradation during bulk updates.
const idx = new Document({
fastupdate: true, // essential for heavy tag usage
document:{
id:'id',
tag:['categories[]']
}
});
Stale Tokens When Re-indexing Documents
The removal path in add_index is currently commented out (lines 62‑66 in src/document/add.js). Consequently, persistent indexes never automatically delete terms when a source field disappears. Re-adding a document without first removing it leaves stale tokens in the index.
// Correct re-indexing pattern
idx.remove(doc.id);
idx.add(doc);
Practical Implementation Examples
Indexing a Moderately Nested Document
This example demonstrates proper colon syntax, array markers, and storage configuration:
import Document from "flexsearch";
const idx = new Document({
fastupdate: true,
document: {
id: "meta:id",
index: [
"title",
"author:name",
"author:tags[]",
"categories[]",
"meta:stats:views"
],
store: ["title", "author:name"],
tag: ["author:tags[]", "categories[]"]
}
});
const doc = {
meta: { id: 42, stats: { views: 123 } },
title: "FlexSearch Deep Dive",
author: { name: "Alice", tags: ["javascript", "search"] },
categories: ["library", "frontend"]
};
idx.add(doc);
const results = idx.search("javascript");
console.log(results);
Using a Custom Extractor for Non-Standard Shapes
When data arrives in formats like comma-separated strings, use a custom function to normalize values:
const idx = new Document({
document: {
id: "id",
index: [
{
field: "tags",
custom: doc => (doc.tags || "").split(",")
}
]
}
});
idx.add({ id: 1, tags: "node,js,backend" });
console.log(idx.search("js")); // → [{ id: 1 }]
Re-indexing a Document with Removed Fields
To prevent stale search results when fields are deleted from your data model:
// First version contains a summary field
idx.add({ id: 5, title: "Intro", summary: "short text" });
// Later version removes the summary field
idx.remove(5);
idx.add({ id: 5, title: "Intro" }); // Old summary tokens are now cleared
Summary
- Always use colons (
:) to separate nested path segments; dot notation causes silent indexing failures. - Append
[]to every array segment in your descriptor to ensure individual elements are indexed separately. - Enable
fastupdate: truewhen working with high-cardinality tag fields to maintain performance during bulk operations. - Explicitly declare
storefor any fields you intend to use for result enrichment. - Call
remove(id)before re-adding documents when fields have been deleted to eliminate stale tokens. - Ensure custom extractor functions always return a value or empty array, never
undefined.
Frequently Asked Questions
Why does FlexSearch use colons instead of dots for nested paths?
FlexSearch uses colons as path separators because the parse_tree function in src/document.js (lines 387‑410) explicitly splits field names on the colon character to build the internal tree structure. Dots are treated as literal characters and become part of the key name, causing the indexer to look for a property literally named "author.name" rather than the name property inside an author object.
What happens if I forget to add [] to an array field in the descriptor?
Without the [] marker, FlexSearch treats the entire array as a single string value (effectively joining elements with spaces) and indexes it as one token. This prevents individual element matching—searching for a specific tag will fail because the index contains the concatenated string rather than discrete tokens. The [] suffix sets the internal marker flag (lines 389‑401 in src/document.js) that triggers per-element indexing in add_index (lines 34‑41 in src/document/add.js).
How do I update a document when nested fields have been deleted?
Because the removal path in add_index is currently commented out (lines 62‑66 in src/document/add.js), FlexSearch does not automatically prune terms when a source field disappears. To prevent stale tokens from remaining in the index, you must explicitly call index.remove(id) before calling index.add(doc) with the updated data. This ensures the old document registration is cleared before the new version is indexed.
Can I use custom functions to transform data before indexing?
Yes, the Document API supports a custom property in field descriptors that accepts a function receiving the document object and returning the value to index. However, your function must always return a defined value—returning undefined causes add_index to skip the field entirely (see the if (!tree) continue check at lines 75‑78 in src/document/add.js). For array fields, return an array; for single values, return the primitive or null to explicitly index nothing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →