Configuring Thinking Budget in Earendil Pi: Complete Technical Guide
TLDR: Earendil Pi exposes a five-level thinking budget (off, low, medium, high, max) that controls internal model deliberation tokens, configurable via ~/.pi/agent/settings.json, the PI_THINKING_LEVEL environment variable, or runtime slash commands.
Earendil Pi (also known as Instagit Pi) is a modular, TypeScript-first command-line interface for executing tool-augmented AI workflows from the terminal. When configuring thinking budget in earendil pi, you manage the thinking level—a soft limit on reasoning tokens that affects how extensively the model deliberates before generating visible output. This article details the configuration architecture, runtime controls, and source implementation based on the @earendil-works/pi monorepo.
Understanding the Thinking Budget Architecture
The thinking budget implementation spans four primary packages in the earendil-works/pi repository:
- @earendil-works/pi-coding-agent: Contains the high-level CLI, slash-command parser, and session management
- @earendil-works/pi-agent-core: Handles the core agent loop and streaming events
- @earendil-works/pi-ai: Provides provider-agnostic model definitions and token counting
- @earendil-works/pi-tui: Renders thinking deltas in the terminal UI
The data flow follows a clear pipeline. First, SessionManager (packages/coding-agent/src/core/session-manager.ts) stores the thinkingLevel property in session state. Next, model-resolver.ts reads this value and, if the provider supports thinking capabilities, injects it into the request payload. Finally, TuiRenderer (packages/tui/src/TuiRenderer.ts) displays incoming thinking tokens with distinct theme colors while system-prompt.ts incorporates the current level into the system prompt so the model can adapt its behavior.
Configuring Thinking Levels via Settings File
Set your default thinking budget by editing the user configuration file at ~/.pi/agent/settings.json. The thinkingLevel key accepts one of five string values:
{
"thinkingLevel": "medium"
}
Available levels and their approximate token budgets are:
| Level | Approximate Token Budget | Visual Cue |
|---|---|---|
off |
0 tokens (no thinking output) | Gray text |
low |
~50 tokens | Light cyan |
medium |
~150 tokens | Cyan |
high |
~300 tokens | Bright cyan |
max |
Unlimited (until model stops) | Bold cyan |
Pi applies this default on startup. The configuration is read by utility functions in packages/coding-agent/src/utils/config.ts:
export function getThinkingLevel(session: Session): ThinkingLevel {
return session.thinkingLevel ?? "off";
}
Runtime Configuration Methods
You can override the configured thinking budget during an active session without restarting the application.
Keyboard Shortcut and Slash Commands
Press Shift + Tab to cycle through thinking levels, or use the explicit slash command:
/thinking high
This command updates the session state immediately. As implemented in packages/coding-agent/src/core/model-resolver.ts, the change propagates to the next model call:
if (session.thinkingLevel && providerSupportsThinking) {
request.thinking = session.thinkingLevel;
}
Environment Variables for Scripting
When automating Pi in shell scripts, set the PI_THINKING_LEVEL variable before invoking the CLI (parsed by cli/args.ts):
PI_THINKING_LEVEL=high pi -p "Explain quantum computing"
This approach is useful for CI/CD pipelines or batch processing where you want consistent reasoning depth across multiple invocations.
How Thinking Tokens Are Processed
The thinking budget acts as a soft limit rather than a hard token ceiling. Pi leverages the provider's native "thinking" features (such as Anthropic's reasoning parameters) combined with client-side streaming buffers.
When the model generates thinking content, TuiRenderer receives three event types:
thinking_start— Begins the reasoning blockthinking_delta— Streams incremental reasoning tokensthinking_end— Closes the reasoning block
The renderer applies theme colors based on the current level. Visual cues include gray text for off, light cyan for low, cyan for medium, bright cyan for high, and bold cyan for max. These definitions reside in packages/tui/theme/theme.ts.
Meanwhile, packages/coding-agent/src/core/system-prompt.ts injects a description of the current thinking level into the system prompt, ensuring the model understands its constraints.
Combining Budgets with Output Limits
For workflows requiring strict resource constraints, combine the thinking level with the max_output_tokens provider option (defined in model-config.ts). This creates a hard ceiling on total output while preserving the allocated reasoning budget:
const config = {
thinkingLevel: "high",
max_output_tokens: 4000 // Hard limit including thinking + response
};
This configuration ensures the model thinks within the ~300 token budget but stops generation entirely at 4000 tokens.
Summary
- The thinking budget in earendil pi uses five discrete levels (
offthroughmax) mapped to approximate token allowances - Default configuration lives in
~/.pi/agent/settings.jsonunder thethinkingLevelkey - Runtime adjustments work via Shift + Tab,
/thinkingcommands, or thePI_THINKING_LEVELenvironment variable SessionManagerstores state,model-resolver.tsapplies it to provider requests, andTuiRendererdisplays thinking streams with theme-specific colors- The system prompt builder automatically informs the model of its current thinking constraints
Frequently Asked Questions
What is the difference between thinking level and max_output_tokens?
The thinking level controls internal deliberation tokens before the model produces visible output, while max_output_tokens sets a hard ceiling on the total response length including both thinking and final output. According to the source code in packages/coding-agent/src/core/model-resolver.ts, the thinking level is passed to the provider's thinking parameter, whereas max_output_tokens caps the entire generation. Use them together when you need reasoning capability with strict overall limits.
How do I disable thinking output completely in earendil pi?
Set the thinking level to off either in your ~/.pi/agent/settings.json configuration file or by typing /thinking off during a session. When set to off, the getThinkingLevel function in packages/coding-agent/src/utils/config.ts returns zero tokens, and TuiRenderer displays thinking blocks in gray text (or suppresses them entirely depending on the theme configuration).
Can I programmatically check the current thinking budget?
Yes. Import SessionManager from @earendil-works/pi-coding-agent and access the thinkingLevel property:
import { SessionManager } from "@earendil-works/pi-coding-agent";
const session = SessionManager.current();
console.log(session.thinkingLevel); // "medium", "high", etc.
This reads the live session state maintained in packages/coding-agent/src/core/session-manager.ts, reflecting any runtime changes made via slash commands or keyboard shortcuts.
Why does the thinking budget use approximate token counts?
The budget is a soft limit (as noted in the packages/coding-agent/src/core/model-resolver.ts implementation) because Pi delegates token enforcement to the underlying LLM provider's native reasoning features. The client buffers and displays thinking deltas but relies on the provider to halt reasoning generation. The approximate values (~50, ~150, ~300 tokens) represent heuristic targets that align with provider-specific implementations of reasoning depth controls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →