Intent-Driven Filtering vs Auto-Indexing in ctx_execute: Managing Large Command Outputs

ctx_execute employs intent-driven filtering to return BM25-ranked excerpts for outputs over 5 KB when a search intent is provided, whereas auto-indexing stores the entire output in SQLite FTS5 without returning text when outputs exceed 100 KB.

The ctx_execute tool in the mksglu/context-mode repository handles massive command outputs through two complementary optimization strategies. While both mechanisms leverage a persistent SQLite FTS5 store, they differ fundamentally in activation thresholds, retrieval logic, and what the LLM ultimately receives. Understanding these distinctions ensures you can predict how large execution results are processed and when to apply manual search queries.

Activation Thresholds and Conditions

The execution pipeline checks output size against two distinct constants defined in the source code, triggering different behaviors based on whether the user has supplied a search intent.

Intent-Driven Filtering (5 KB Threshold)

When a user provides an intent and the command output exceeds approximately 5 KB (INTENT_SEARCH_THRESHOLD = 5_000 bytes), the system activates intent-driven filtering. Rather than truncating or omitting the data, the tool indexes the full output and runs a BM25-style search against the supplied intent query. This occurs in src/server.ts at lines 690-697 for standard output and lines 740-748 for error output.

The function returns only the matching section titles and brief previews, formatted as a concise list of relevant excerpts.

Auto-Indexing (100 KB Threshold)

Auto-indexing triggers when output exceeds 100 KB (LARGE_OUTPUT_THRESHOLD = 102_400 bytes), regardless of whether an intent was provided. Implemented at lines 779-782 in src/server.ts, this mechanism calls indexStdout to persist the entire output blob to the FTS5 knowledge base. Instead of returning any raw text, the tool responds with a short pointer message indicating the data is now searchable.

Implementation Details in src/server.ts

The logic branches are clearly separated in the core server implementation, with each strategy calling distinct indexing functions.

Intent-Driven Search Logic

When the 5 KB threshold is met and intent is present, the code indexes the output and executes intentSearch:

// Lines 690-697: Intent-driven filtering for stdout ≥ 5 KB
if (intent && intent.trim().length > 0 && Buffer.byteLength(stdout) > INTENT_SEARCH_THRESHOLD) {
  trackIndexed(Buffer.byteLength(stdout));
  return trackResponse("ctx_execute", {
    content: [{ type: "text", text: intentSearch(stdout, intent, `execute:${language}`) }],
  });
}

The same pattern applies to error output at lines 740-748, using the label execute:${language}:error to distinguish stderr content.

Auto-Indexing Logic

For massive outputs exceeding 100 KB, the tool bypasses content return entirely:

// Lines 779-782: Auto-indexing for stdout ≥ 100 KB
if (Buffer.byteLength(stdout) > LARGE_OUTPUT_THRESHOLD) {
  // Returns a short "Indexed ... sections ..." message.
  return trackResponse("ctx_execute", indexStdout(stdout, `execute:${language}`));
}

Key Differences in Behavior

Characteristic Intent-Driven Filtering Auto-Indexing
Activation Output > 5 KB and intent provided Output > 100 KB (intent optional)
Search Query Uses user-supplied intent No immediate query; full text stored
LLM Response Filtered excerpts with section titles Pointer message only; no raw text
Underlying Call intentSearch() indexStdout()
Source Location src/server.ts:690-697 and :740-748 src/server.ts:779-782

The Persistent Storage Layer

Both mechanisms rely on the same underlying infrastructure defined in src/store.ts. The getStore() function provides access to an SQLite FTS5 database, ensuring that indexed content persists across tool invocations. When auto-indexing stores a 100 KB blob, or when intent-driven filtering indexes 5 KB of output, the data remains available for later retrieval via search(queries: [...]) calls even if the initial response omitted the full content.

Outputs that fall between these thresholds—larger than typical but under 100 KB without an intent—are handled by standard truncation utilities (indirectly referencing src/truncate.ts) rather than the indexing pipeline.

Summary

  • Intent-driven filtering activates at the 5 KB threshold when a search intent is present, returning only BM25-ranked excerpts from src/server.ts lines 690-697.
  • Auto-indexing activates at the 100 KB threshold via lines 779-782, persisting the entire output to SQLite FTS5 without returning text to the LLM.
  • Both strategies use the getStore() abstraction from src/store.ts for durable storage.
  • Error outputs support the same intent-driven filtering logic at lines 740-748, using the :error label suffix.

Frequently Asked Questions

What happens if output is 50 KB but no intent is provided?

The output does not trigger intent-driven filtering (no intent present) and remains below the 100 KB auto-indexing threshold. According to the codebase structure, such cases fall back to standard truncation utilities rather than FTS5 indexing.

How do I retrieve content after auto-indexing?

Once auto-indexing stores the output, you can retrieve specific sections by calling the search functionality with targeted queries. The indexed data remains accessible in the SQLite FTS5 store via search(queries: [...]) with appropriate query strings.

Does intent-driven filtering work for stderr?

Yes. The same logic applied to stdout at lines 690-697 is duplicated for error output at lines 740-748 in src/server.ts. When indexing errors, the system appends :error to the language label (e.g., execute:python:error).

What database powers the indexing in ctx_execute?

Both mechanisms utilize an SQLite FTS5 database accessed through the getStore() wrapper in src/store.ts. This ensures full-text search capabilities and persistent storage for all indexed command outputs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →