Intent-Driven Filtering vs Auto-Indexing in ctx_execute: Managing Large Command Outputs
ctx_execute employs intent-driven filtering to return BM25-ranked excerpts for outputs over 5 KB when a search intent is provided, whereas auto-indexing stores the entire output in SQLite FTS5 without returning text when outputs exceed 100 KB.
The ctx_execute tool in the mksglu/context-mode repository handles massive command outputs through two complementary optimization strategies. While both mechanisms leverage a persistent SQLite FTS5 store, they differ fundamentally in activation thresholds, retrieval logic, and what the LLM ultimately receives. Understanding these distinctions ensures you can predict how large execution results are processed and when to apply manual search queries.
Activation Thresholds and Conditions
The execution pipeline checks output size against two distinct constants defined in the source code, triggering different behaviors based on whether the user has supplied a search intent.
Intent-Driven Filtering (5 KB Threshold)
When a user provides an intent and the command output exceeds approximately 5 KB (INTENT_SEARCH_THRESHOLD = 5_000 bytes), the system activates intent-driven filtering. Rather than truncating or omitting the data, the tool indexes the full output and runs a BM25-style search against the supplied intent query. This occurs in src/server.ts at lines 690-697 for standard output and lines 740-748 for error output.
The function returns only the matching section titles and brief previews, formatted as a concise list of relevant excerpts.
Auto-Indexing (100 KB Threshold)
Auto-indexing triggers when output exceeds 100 KB (LARGE_OUTPUT_THRESHOLD = 102_400 bytes), regardless of whether an intent was provided. Implemented at lines 779-782 in src/server.ts, this mechanism calls indexStdout to persist the entire output blob to the FTS5 knowledge base. Instead of returning any raw text, the tool responds with a short pointer message indicating the data is now searchable.
Implementation Details in src/server.ts
The logic branches are clearly separated in the core server implementation, with each strategy calling distinct indexing functions.
Intent-Driven Search Logic
When the 5 KB threshold is met and intent is present, the code indexes the output and executes intentSearch:
// Lines 690-697: Intent-driven filtering for stdout ≥ 5 KB
if (intent && intent.trim().length > 0 && Buffer.byteLength(stdout) > INTENT_SEARCH_THRESHOLD) {
trackIndexed(Buffer.byteLength(stdout));
return trackResponse("ctx_execute", {
content: [{ type: "text", text: intentSearch(stdout, intent, `execute:${language}`) }],
});
}
The same pattern applies to error output at lines 740-748, using the label execute:${language}:error to distinguish stderr content.
Auto-Indexing Logic
For massive outputs exceeding 100 KB, the tool bypasses content return entirely:
// Lines 779-782: Auto-indexing for stdout ≥ 100 KB
if (Buffer.byteLength(stdout) > LARGE_OUTPUT_THRESHOLD) {
// Returns a short "Indexed ... sections ..." message.
return trackResponse("ctx_execute", indexStdout(stdout, `execute:${language}`));
}
Key Differences in Behavior
| Characteristic | Intent-Driven Filtering | Auto-Indexing |
|---|---|---|
| Activation | Output > 5 KB and intent provided | Output > 100 KB (intent optional) |
| Search Query | Uses user-supplied intent | No immediate query; full text stored |
| LLM Response | Filtered excerpts with section titles | Pointer message only; no raw text |
| Underlying Call | intentSearch() |
indexStdout() |
| Source Location | src/server.ts:690-697 and :740-748 |
src/server.ts:779-782 |
The Persistent Storage Layer
Both mechanisms rely on the same underlying infrastructure defined in src/store.ts. The getStore() function provides access to an SQLite FTS5 database, ensuring that indexed content persists across tool invocations. When auto-indexing stores a 100 KB blob, or when intent-driven filtering indexes 5 KB of output, the data remains available for later retrieval via search(queries: [...]) calls even if the initial response omitted the full content.
Outputs that fall between these thresholds—larger than typical but under 100 KB without an intent—are handled by standard truncation utilities (indirectly referencing src/truncate.ts) rather than the indexing pipeline.
Summary
- Intent-driven filtering activates at the 5 KB threshold when a search intent is present, returning only BM25-ranked excerpts from
src/server.tslines 690-697. - Auto-indexing activates at the 100 KB threshold via lines 779-782, persisting the entire output to SQLite FTS5 without returning text to the LLM.
- Both strategies use the
getStore()abstraction fromsrc/store.tsfor durable storage. - Error outputs support the same intent-driven filtering logic at lines 740-748, using the
:errorlabel suffix.
Frequently Asked Questions
What happens if output is 50 KB but no intent is provided?
The output does not trigger intent-driven filtering (no intent present) and remains below the 100 KB auto-indexing threshold. According to the codebase structure, such cases fall back to standard truncation utilities rather than FTS5 indexing.
How do I retrieve content after auto-indexing?
Once auto-indexing stores the output, you can retrieve specific sections by calling the search functionality with targeted queries. The indexed data remains accessible in the SQLite FTS5 store via search(queries: [...]) with appropriate query strings.
Does intent-driven filtering work for stderr?
Yes. The same logic applied to stdout at lines 690-697 is duplicated for error output at lines 740-748 in src/server.ts. When indexing errors, the system appends :error to the language label (e.g., execute:python:error).
What database powers the indexing in ctx_execute?
Both mechanisms utilize an SQLite FTS5 database accessed through the getStore() wrapper in src/store.ts. This ensures full-text search capabilities and persistent storage for all indexed command outputs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →