How to Use Keystore to Handle Large In-Memory Indexes in FlexSearch
Enable the keystore option in FlexSearch to expand the default 16 million key limit to billions or trillions of distinct terms without allocating extra memory per key.
FlexSearch is a high-performance full-text search library that stores term and partial keys in an in-memory keystore. By default, this keystore supports up to 2²⁴ (approximately 16 million) distinct keys. When building indexes for massive datasets, you must configure the keystore to handle large in-memory indexes in FlexSearch by extending the address space.
Understanding the FlexSearch Keystore Architecture
The keystore is the internal structure that maps distinct terms and partials to their posting lists. Understanding its limitations and extension mechanism is critical for scaling.
Default Key Space Limitations
By default, the keystore can hold up to 2²⁴ distinct terms/partials (approximately 16 million). This limit is hardcoded in the address space allocation within src/keystore.js. When you attempt to index more unique terms than this limit allows, the index will fail to store additional keys.
How Address Space Extension Works
The keystore configuration option accepts a value between 8 and 64, representing the number of extra address bits to allocate. The total key space becomes 2^(24+bits):
keystore: 8→ 2³² keys (≈ 4.3 billion)keystore: 16→ 2⁴⁰ keys (≈ 1 trillion)keystore: 64→ 2⁸⁸ keys (theoretical upper bound)
According to the documentation in doc/keystore.md, enabling a larger keystore does not allocate additional RAM per key; it only expands the addressable range.
Configuring Keystore for Large-Scale Indexes
Proper configuration ensures you can handle massive datasets without memory waste or performance degradation.
Calculating Required Address Bits
To determine your keystore value:
- Estimate your maximum distinct term count (including partials if using contextual indexes).
- Calculate required bits:
bits = ceil(log2(estimated_terms) - 24). - Round up to the nearest valid option (8, 16, 32, or 64).
For most large-scale applications, keystore: 8 (4.3 billion keys) provides sufficient headroom. Only extreme use cases require keystore: 16 or higher.
Recommended Configuration Settings
When enabling extended keystore support, combine it with performance optimizations:
import FlexSearch from "flexsearch";
const index = new FlexSearch.Index({
keystore: 16, // 2^40 possible keys
fastupdate: true, // Reduces overhead during bulk inserts
context: true // Enables phrase searching
});
The fastupdate option is particularly important when handling large in-memory indexes, as it minimizes the cost of rebuilding posting lists after each insertion.
Handling Massive Posting Lists
Beyond key count limits, individual terms may accumulate billions of document references. FlexSearch handles this through automatic scaling mechanisms.
Automatic Proxy-Based Scaling
When a term’s posting list grows beyond 2³¹ entries, FlexSearch automatically switches to a Proxy-based storage system. This is implemented in src/keystore.js and verified in test/keystore.js.
The test suite artificially inflates posting lists to exceed the 2³¹ threshold:
// From test/keystore.js
const index = new Index({
fastupdate: true,
keystore: 16,
context: true
});
// Simulate massive posting list
let foo = index.map.get("fo");
foo[0].length = 2**31 - 10;
// Continue adding beyond 2^31
for (let i = foo[0].length; i <= 2**31 + 9; i++) {
index.add(i, "foo bar");
}
This automatic scaling ensures you can exceed 2 billion documents per term without manual intervention or array overflow errors.
Memory Management with cleanup()
After massive deletions or bulk updates, obsolete posting entries may remain in memory. The cleanup() method, available on index instances, prunes these obsolete postings and reduces the map size to the minimum required.
Call cleanup() after bulk removal operations:
// After deleting millions of documents
index.cleanup(); // Reclaims memory from empty term buckets
Code Examples
Scaling to Hundreds of Millions of Terms
This example demonstrates configuring FlexSearch for extreme scale with 64-bit addressing:
import FlexSearch from "flexsearch";
// 64-bit keystore → 2^88 possible keys (theoretical upper bound)
const index = new FlexSearch.Index({
keystore: 64,
fastupdate: true,
context: true,
});
// Simulate indexing massive dataset
for (let i = 0; i < 10_000_000; i++) {
index.add(i, `document number ${i} with many unique terms ${Math.random()}`);
}
// Compact after bulk load
index.cleanup();
Key implementation details:
- The
keystorevalue represents extra address bits, not byte size. fastupdatemaintains insertion performance at scale.cleanup()frees unused term slots after bulk operations.
Testing Extended Limits
Verify your keystore configuration handles posting lists exceeding 2³¹ entries:
import { Index } from "flexsearch";
const index = new Index({
fastupdate: true,
keystore: 16, // 2^40 address space
context: true
});
// Seed initial documents
index.add(0, "foo bar");
index.add(1, "foo bar");
index.add(2, "foo bar");
// Access internal posting lists (advanced usage)
let foo = index.map.get("fo");
let bar = index.ctx.get("fo").get("bar");
// Artificially inflate to near 2^31 limit
foo[0].length = 2**31 - 10;
bar[0].length = 2**31 - 10;
// Push beyond 2^31 to trigger Proxy-based storage
for (let i = foo[0].length; i <= 2**31 + 9; i++) {
index.add(i, "foo bar");
}
console.log("Successfully stored >2^31 postings per term");
This test confirms that FlexSearch automatically switches to Proxy-based arrays when posting lists exceed 2³¹ entries, as implemented in src/keystore.js and validated in test/keystore.js.
Summary
- Default limits: FlexSearch supports 2²⁴ (≈16M) distinct terms by default, defined in the keystore address space.
- Address extension: Add extra bits via
keystore: 8throughkeystore: 64to support billions or trillions of keys without per-key memory overhead. - Posting list scaling: Automatic Proxy-based storage activates when individual terms exceed 2³¹ document references.
- Performance optimization: Combine
keystorewithfastupdate: truefor bulk inserts and callcleanup()after deletions to reclaim memory. - Persistent storage incompatibility: Do not enable
keystoreon SQLite, Redis, or other persistent backends, as they ignore the limit and may buffer excessively.
Frequently Asked Questions
What is the default key limit in FlexSearch?
By default, FlexSearch stores up to 2²⁴ distinct terms or partials (approximately 16.7 million) in the in-memory keystore. This limit is hardcoded in the address space allocation within src/keystore.js and documented in doc/keystore.md. Once you exceed this threshold, the index cannot store additional unique keys unless you enable the extended keystore.
Does enabling a larger keystore consume more RAM?
No. Enabling a larger keystore via the keystore option does not allocate additional RAM per key. According to the documentation in doc/keystore.md, the option only expands the addressable range by adding extra bits to the key space (2^(24+bits)). The memory overhead is negligible because the library does not pre-allocate space for the theoretical maximum; it merely allows the index to address more distinct terms as they are inserted.
Can I use keystore with persistent storage backends?
No. The keystore option is designed exclusively for in-memory indexes. Persistent storage backends such as SQLite, Redis, and other custom storages ignore the keystore limit because they manage their own key spaces. Enabling keystore on a persistent index is unnecessary and may cause the internal buffer to stress before you call index.commit(), as noted in doc/keystore.md. Use keystore only when keeping the entire index in memory.
How do I know if I need to increase the keystore size?
You need to increase the keystore size if you anticipate indexing more than 16 million distinct terms or partials. Calculate your requirements by estimating the cardinality of your vocabulary: if you expect billions of unique terms (common in large-scale log analysis or genomic data), set keystore: 8 (4.3 billion keys) or higher. Monitor for insertion failures or memory errors during bulk indexing; if the index stops accepting new unique terms despite available system memory, you have likely hit the 2²⁴ default limit and must restart with a higher keystore value.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →