Think in Code Paradigm in context-mode: How It Reduces LLM Context Consumption
The "Think in Code" paradigm requires LLMs to generate executable scripts that run in a sandbox environment, returning only concise results rather than ingesting raw data into the context window, reducing token usage by up to 99%.
The Think in Code paradigm is a mandatory design rule enforced throughout mksglu/context-mode that fundamentally changes how language models interact with data. Instead of prompting the LLM to read and analyze large file contents directly—which rapidly consumes context window capacity—the model must write short programs that execute locally and console.log() only the essential findings.
What Is the Think in Code Paradigm?
According to the repository's README.md at line 40, Think in Code is not optional guidance but a strict architectural constraint. The paradigm treats the LLM as a code generator rather than a data processor. When the model needs to perform analysis—such as counting functions across 50 files or processing large API responses—it must write a script that runs inside the sandbox environment via tools like ctx_execute.
The web interface documentation in web/index.html (lines 255-262) reinforces this rule, presenting it as the primary interaction pattern for users. Meanwhile, the routing hooks in hooks/core/routing.mjs (lines 235-270) enforce compliance by intercepting unsafe commands like curl, wget, or raw HTTP calls, and redirecting them to sandbox tools with explicit instructions to "Think in Code."
How Think in Code Reduces Context Consumption
The context savings follow three distinct phases, moving data processing out of the LLM's context window and into the sandbox:
Data Ingestion: Without Think in Code, the LLM pulls every file into the prompt—often exceeding 50 KB of raw source code. With Think in Code, the model sends only a tiny script (approximately 200 bytes) to ctx_execute, achieving roughly 99% reduction in immediate context usage.
Processing: When the LLM processes data internally, it must maintain intermediate calculations and partial results in the conversation history, consuming additional tokens. The sandbox performs all heavy lifting, returning only the final value—such as "27 functions"—reducing the operation from thousands of tokens to a few dozen bytes.
Re-use: Re-running analysis without Think in Code duplicates the entire file payload across multiple turns. With Think in Code, the same script can be reused, and sandbox output caching eliminates repeated token costs for identical operations.
The analytics module in src/session/analytics.ts (line 191) tracks these metrics specifically, monitoring the file-to-output ratio to demonstrate actual context savings in production usage.
Practical Implementation Examples
Analyzing Codebases Without Loading Source Files
Rather than reading 50 JavaScript files into the context to count functions, the LLM generates a script executed via ctx_execute:
// Built-in sandbox tool: ctx_execute (JavaScript runtime)
await ctx_execute(
"javascript",
`
const fs = require('fs');
const path = require('path');
// Recursively collect *.js files
const walk = (dir) => {
return fs.readdirSync(dir).flatMap(entry => {
const full = path.join(dir, entry);
return fs.statSync(full).isDirectory()
? walk(full)
: full.endsWith('.js') ? [full] : [];
});
};
const files = walk(process.cwd());
let total = 0;
const fnRegex = /function\\s+\\w+|\\w+\\s*[:=]\\s*\\(.*\\)\\s*=>/g;
for (const f of files) {
const src = fs.readFileSync(f, 'utf8');
const matches = src.match(fnRegex);
total += matches ? matches.length : 0;
}
console.log('Total JS functions:', total);
`
);
The script occupies approximately 1 KB, while the interaction consumes only a few dozen tokens. The LLM sees only Total JS functions: 42 rather than the entire codebase.
Fetching and Indexing External Content
For external data retrieval, the paradigm uses specialized tools that minimize context exposure:
// Fetch and convert HTML to markdown, then index
await ctx_fetch_and_index(
"https://example.com/docs",
{ source: "html" }
);
// Search indexed content with multiple queries in one call
await ctx_search(
["authentication flow", "API rate limit"]
);
The LLM never receives the raw HTML (typically 50+ KB). Only the concise search results enter the context window, preserving tokens for decision-making and high-level guidance.
Enforcement and Configuration
The Think in Code rule extends beyond the core engine into platform-specific configurations. The routing hooks in hooks/core/routing.mjs (lines 235-270) actively intercept dangerous commands—such as direct HTTP requests or build-tool output—and redirect them to appropriate sandbox tools, ensuring the paradigm cannot be accidentally bypassed.
Platform integrations declare this rule explicitly. For example, configs/zed/AGENTS.md (line 5) includes the Think in Code requirement in the system prompt, making the constraint visible to the model at the start of every session regardless of the client interface.
Summary
- Think in Code is a mandatory architectural rule in mksglu/context-mode that requires LLMs to generate sandbox scripts rather than process raw data directly.
- Context consumption drops by approximately 99% when data ingestion moves from the prompt (50+ KB) to sandbox scripts (~200 bytes).
- The
ctx_execute,ctx_fetch_and_index, andctx_searchtools enable complex analysis while returning only final results to the LLM. - Routing hooks in
hooks/core/routing.mjsenforce the paradigm by intercepting unsafe commands and redirecting them to sandbox tools. - Configuration files like
configs/zed/AGENTS.mdensure the rule persists across different editor integrations.
Frequently Asked Questions
What makes Think in Code mandatory in context-mode?
The enforcement occurs at the routing layer. The file hooks/core/routing.mjs (lines 235-270) intercepts commands like curl or wget and redirects them to sandbox tools with explicit "Think in Code" instructions. Additionally, platform-specific configurations such as configs/zed/AGENTS.md embed this requirement directly into system prompts, ensuring the model cannot bypass the paradigm regardless of user input.
How does the sandbox execute the generated code?
The sandbox runs within a controlled environment via the ctx_execute tool, which accepts a language identifier (such as "javascript") and the code string. The execution occurs outside the LLM's context window, with only the console.log() output returned to the conversation. This isolation prevents large data structures from ever entering the token stream while still allowing complex file system operations and computations.
Can Think in Code be disabled or bypassed?
No. The architecture treats this as a core constraint rather than a configurable preference. The routing hooks actively block direct data ingestion methods, redirecting them to sanctioned sandbox tools. Attempts to read large files directly or make raw HTTP requests are intercepted before reaching execution, ensuring context conservation remains enforced across all interactions.
What types of analysis work best with this paradigm?
Operations that summarize large datasets—such as counting symbols across directories, aggregating log files, scraping web content, or processing API responses—provide the highest context savings. Any analysis where the input data volume exceeds the output volume by orders of magnitude benefits from Think in Code, as the LLM handles only the condensed result rather than the raw source material.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →