# Common Pitfalls When Indexing Nested JSON Documents with the FlexSearch Document API

> Avoid common pitfalls when indexing nested JSON FlexSearch documents. Learn to use colons for paths, include array markers, and manage document updates effectively.

- Repository: [Nextapps GmbH/flexsearch](https://github.com/nextapps-de/flexsearch)
- Tags: best-practices
- Published: 2026-02-23

---

**Most indexing errors in FlexSearch occur when developers use dot notation instead of colons for nested paths, omit the `[]` marker for arrays, or fail to remove documents before re-indexing fields that no longer exist.**

The FlexSearch Document API provides a powerful way to index arbitrarily structured JSON objects by mapping properties to searchable fields. However, the internal tree-parsing mechanism—implemented in [`src/document.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document.js)—relies on strict descriptor syntax that differs from standard JavaScript object notation. Understanding how the `parse_tree` function processes colons and array markers is essential to avoid silent indexing failures when working with nested data structures.

## How FlexSearch Processes Nested Documents

Under the hood, a `Document` instance is a thin wrapper that creates one or more ordinary `Index` instances for each field you declare. The wrapper parses the document descriptor into a tree that drives recursive extraction of values.

### Descriptor to Tree Conversion

When you instantiate a `Document`, the constructor calls `parse_tree` (lines 387‑410 in [`src/document.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document.js)) to transform your field paths into an internal tree structure. The function splits each field name on the colon character (`:`), treating each segment as a level in the tree. If a segment ends with `]` (for example, `tags[]`), the corresponding entry in a **marker** array is set to `true` (lines 389‑401). This marker tells the indexer that the value is an array whose items must be indexed separately.

### Recursive Value Extraction

The `add_index` function in [`src/document/add.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document/add.js) (lines 24‑62) walks the JSON object recursively according to the tree structure. When it reaches the final segment of a path, the value is handed to the underlying `Index` via `index.add`. If the marker for that segment is `true` and the value is an array, each element is indexed independently (lines 34‑41).

### Deep Array Handling

When an array appears before the leaf field—for example, `authors[].name`—the recursion **does not advance the tree position**. Instead, the array is iterated, and the same tree level is processed for each item (lines 50‑55 in [`src/document/add.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document/add.js)). This allows indexing of values inside nested arrays without consuming the tree segments prematurely.

## Common Pitfalls and How to Avoid Them

### Using Dot Notation Instead of Colon Separators

The most frequent mistake is using JavaScript dot notation (`author.name`) instead of the required colon syntax (`author:name`). The `parse_tree` function only splits paths on `:`, so a dot remains part of the key name, causing the field to be silently skipped during indexing.

```javascript
// Wrong: uses dot notation, field is never indexed
index.add({author:{name:'Bob'}}, {
  document:{id:'id', index:'author.name'}
});

// Correct: uses colon separator
index.add({author:{name:'Bob'}}, {
  document:{id:'id', index:'author:name'}
});

```

### Omitting the Array Marker Brackets

Without the `[]` suffix, FlexSearch treats an array as a single string value (effectively calling `tags.join(' ')`), which prevents individual element matching and highlighting.

```javascript
// Wrong: indexes entire array as one token
index.add(doc, {document:{id:'id', index:'tags'}});

// Correct: each array item indexed separately
index.add(doc, {document:{id:'id', index:'tags[]'}});

```

### Incorrect Syntax for Deeply Nested Arrays

When paths contain multiple arrays—such as `authors[].books[].title`—you must append `[]` to **every** array segment. Omitting a marker causes the recursion to stop early, indexing only the first array level.

```javascript
// Correct descriptor for deep arrays
const idx = new Document({
  document:{
    id:'id',
    index:['authors[].books[].title']
  }
});

```

### Custom Extractors Returning Undefined

When using a `custom` function to transform data, returning `undefined` or omitting a return value causes `add_index` to skip the field entirely (see the `if (!tree) continue` check at lines 75‑78 in [`src/document/add.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document/add.js)).

```javascript
// Wrong: may return undefined
custom: doc => doc.title

// Correct: always return a value or empty array
custom: doc => doc.title || null
// For arrays:
custom: doc => (doc.tags || "").split(",")

```

### Confusing Index Fields with Store Fields

The `storetree` is built only when the `store` option is truthy (lines 84‑88 in [`src/document.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document.js)). If you attempt to enrich search results with data from a field that is only indexed but not stored, the enrichment will fail silently.

```javascript
// Enable storing for enrichment
const idx = new Document({
  document:{
    id:'id',
    index:['content'],
    store:['content', 'title']  // explicitly store these
  }
});

```

### Mismatched ID Path Syntax

The `id` field uses the same colon-based path parsing as index fields (lines 68‑69 in [`src/document.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document.js)). Using dot notation or an incorrect path causes the document to be registered under an undefined key, making it unsearchable.

```javascript
// Wrong: dot notation
document:{ id:'meta.id' }

// Correct: colon separator
document:{ id:'meta:id' }

```

### Performance Issues Without Fast Update

When indexing high-cardinality tag fields, the internal register (`this.reg`) grows rapidly. Without `fastupdate:true`, each addition triggers a full scan of the register, causing severe performance degradation during bulk updates.

```javascript
const idx = new Document({
  fastupdate: true,  // essential for heavy tag usage
  document:{
    id:'id',
    tag:['categories[]']
  }
});

```

### Stale Tokens When Re-indexing Documents

The removal path in `add_index` is currently commented out (lines 62‑66 in [`src/document/add.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document/add.js)). Consequently, persistent indexes never automatically delete terms when a source field disappears. Re-adding a document without first removing it leaves stale tokens in the index.

```javascript
// Correct re-indexing pattern
idx.remove(doc.id);
idx.add(doc);

```

## Practical Implementation Examples

### Indexing a Moderately Nested Document

This example demonstrates proper colon syntax, array markers, and storage configuration:

```javascript
import Document from "flexsearch";

const idx = new Document({
  fastupdate: true,
  document: {
    id: "meta:id",
    index: [
      "title",
      "author:name",
      "author:tags[]",
      "categories[]",
      "meta:stats:views"
    ],
    store: ["title", "author:name"],
    tag: ["author:tags[]", "categories[]"]
  }
});

const doc = {
  meta: { id: 42, stats: { views: 123 } },
  title: "FlexSearch Deep Dive",
  author: { name: "Alice", tags: ["javascript", "search"] },
  categories: ["library", "frontend"]
};

idx.add(doc);

const results = idx.search("javascript");
console.log(results);

```

### Using a Custom Extractor for Non-Standard Shapes

When data arrives in formats like comma-separated strings, use a custom function to normalize values:

```javascript
const idx = new Document({
  document: {
    id: "id",
    index: [
      { 
        field: "tags", 
        custom: doc => (doc.tags || "").split(",") 
      }
    ]
  }
});

idx.add({ id: 1, tags: "node,js,backend" });
console.log(idx.search("js")); // → [{ id: 1 }]

```

### Re-indexing a Document with Removed Fields

To prevent stale search results when fields are deleted from your data model:

```javascript
// First version contains a summary field
idx.add({ id: 5, title: "Intro", summary: "short text" });

// Later version removes the summary field
idx.remove(5);
idx.add({ id: 5, title: "Intro" }); // Old summary tokens are now cleared

```

## Summary

- **Always use colons (`:`)** to separate nested path segments; dot notation causes silent indexing failures.
- **Append `[]`** to every array segment in your descriptor to ensure individual elements are indexed separately.
- **Enable `fastupdate: true`** when working with high-cardinality tag fields to maintain performance during bulk operations.
- **Explicitly declare `store`** for any fields you intend to use for result enrichment.
- **Call `remove(id)`** before re-adding documents when fields have been deleted to eliminate stale tokens.
- **Ensure custom extractor functions** always return a value or empty array, never `undefined`.

## Frequently Asked Questions

### Why does FlexSearch use colons instead of dots for nested paths?

FlexSearch uses colons as path separators because the `parse_tree` function in [`src/document.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document.js) (lines 387‑410) explicitly splits field names on the colon character to build the internal tree structure. Dots are treated as literal characters and become part of the key name, causing the indexer to look for a property literally named `"author.name"` rather than the `name` property inside an `author` object.

### What happens if I forget to add `[]` to an array field in the descriptor?

Without the `[]` marker, FlexSearch treats the entire array as a single string value (effectively joining elements with spaces) and indexes it as one token. This prevents individual element matching—searching for a specific tag will fail because the index contains the concatenated string rather than discrete tokens. The `[]` suffix sets the internal marker flag (lines 389‑401 in [`src/document.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document.js)) that triggers per-element indexing in `add_index` (lines 34‑41 in [`src/document/add.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document/add.js)).

### How do I update a document when nested fields have been deleted?

Because the removal path in `add_index` is currently commented out (lines 62‑66 in [`src/document/add.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document/add.js)), FlexSearch does not automatically prune terms when a source field disappears. To prevent stale tokens from remaining in the index, you must explicitly call `index.remove(id)` before calling `index.add(doc)` with the updated data. This ensures the old document registration is cleared before the new version is indexed.

### Can I use custom functions to transform data before indexing?

Yes, the Document API supports a `custom` property in field descriptors that accepts a function receiving the document object and returning the value to index. However, your function must always return a defined value—returning `undefined` causes `add_index` to skip the field entirely (see the `if (!tree) continue` check at lines 75‑78 in [`src/document/add.js`](https://github.com/nextapps-de/flexsearch/blob/main/src/document/add.js)). For array fields, return an array; for single values, return the primitive or `null` to explicitly index nothing.