# Configuring Thinking Budget in Earendil Pi: Complete Technical Guide

> Learn to configure thinking budget in Earendil Pi. This guide details using settings.json, environment variables, and slash commands for optimal model deliberation. Fine-tune your AI's thought process effortlessly.

- Repository: [Earendil Works/pi](https://github.com/earendil-works/pi)
- Tags: how-to-guide
- Published: 2026-05-25

---

**TLDR:** Earendil Pi exposes a five-level thinking budget (`off`, `low`, `medium`, `high`, `max`) that controls internal model deliberation tokens, configurable via `~/.pi/agent/settings.json`, the `PI_THINKING_LEVEL` environment variable, or runtime slash commands.

Earendil Pi (also known as Instagit Pi) is a modular, TypeScript-first command-line interface for executing tool-augmented AI workflows from the terminal. When configuring thinking budget in earendil pi, you manage the **thinking level**—a soft limit on reasoning tokens that affects how extensively the model deliberates before generating visible output. This article details the configuration architecture, runtime controls, and source implementation based on the `@earendil-works/pi` monorepo.

## Understanding the Thinking Budget Architecture

The thinking budget implementation spans four primary packages in the `earendil-works/pi` repository:

- **@earendil-works/pi-coding-agent**: Contains the high-level CLI, slash-command parser, and session management
- **@earendil-works/pi-agent-core**: Handles the core agent loop and streaming events  
- **@earendil-works/pi-ai**: Provides provider-agnostic model definitions and token counting
- **@earendil-works/pi-tui**: Renders thinking deltas in the terminal UI

The data flow follows a clear pipeline. First, `SessionManager` ([`packages/coding-agent/src/core/session-manager.ts`](https://github.com/earendil-works/pi/blob/main/packages/coding-agent/src/core/session-manager.ts)) stores the `thinkingLevel` property in session state. Next, [`model-resolver.ts`](https://github.com/earendil-works/pi/blob/main/model-resolver.ts) reads this value and, if the provider supports thinking capabilities, injects it into the request payload. Finally, `TuiRenderer` ([`packages/tui/src/TuiRenderer.ts`](https://github.com/earendil-works/pi/blob/main/packages/tui/src/TuiRenderer.ts)) displays incoming thinking tokens with distinct theme colors while [`system-prompt.ts`](https://github.com/earendil-works/pi/blob/main/system-prompt.ts) incorporates the current level into the system prompt so the model can adapt its behavior.

## Configuring Thinking Levels via Settings File

Set your default thinking budget by editing the user configuration file at `~/.pi/agent/settings.json`. The `thinkingLevel` key accepts one of five string values:

```json
{
  "thinkingLevel": "medium"
}

```

Available levels and their approximate token budgets are:

| Level | Approximate Token Budget | Visual Cue |
|-------|-------------------------|------------|
| `off` | 0 tokens (no thinking output) | Gray text |
| `low` | ~50 tokens | Light cyan |
| `medium` | ~150 tokens | Cyan |
| `high` | ~300 tokens | Bright cyan |
| `max` | Unlimited (until model stops) | Bold cyan |

Pi applies this default on startup. The configuration is read by utility functions in [`packages/coding-agent/src/utils/config.ts`](https://github.com/earendil-works/pi/blob/main/packages/coding-agent/src/utils/config.ts):

```typescript
export function getThinkingLevel(session: Session): ThinkingLevel {
    return session.thinkingLevel ?? "off";
}

```

## Runtime Configuration Methods

You can override the configured thinking budget during an active session without restarting the application.

### Keyboard Shortcut and Slash Commands

Press **Shift + Tab** to cycle through thinking levels, or use the explicit slash command:

```bash
/thinking high

```

This command updates the session state immediately. As implemented in [`packages/coding-agent/src/core/model-resolver.ts`](https://github.com/earendil-works/pi/blob/main/packages/coding-agent/src/core/model-resolver.ts), the change propagates to the next model call:

```typescript
if (session.thinkingLevel && providerSupportsThinking) {
    request.thinking = session.thinkingLevel;
}

```

### Environment Variables for Scripting

When automating Pi in shell scripts, set the `PI_THINKING_LEVEL` variable before invoking the CLI (parsed by [`cli/args.ts`](https://github.com/earendil-works/pi/blob/main/cli/args.ts)):

```bash
PI_THINKING_LEVEL=high pi -p "Explain quantum computing"

```

This approach is useful for CI/CD pipelines or batch processing where you want consistent reasoning depth across multiple invocations.

## How Thinking Tokens Are Processed

The thinking budget acts as a soft limit rather than a hard token ceiling. Pi leverages the provider's native "thinking" features (such as Anthropic's reasoning parameters) combined with client-side streaming buffers.

When the model generates thinking content, `TuiRenderer` receives three event types:

- `thinking_start` — Begins the reasoning block
- `thinking_delta` — Streams incremental reasoning tokens  
- `thinking_end` — Closes the reasoning block

The renderer applies theme colors based on the current level. Visual cues include gray text for `off`, light cyan for `low`, cyan for `medium`, bright cyan for `high`, and bold cyan for `max`. These definitions reside in [`packages/tui/theme/theme.ts`](https://github.com/earendil-works/pi/blob/main/packages/tui/theme/theme.ts).

Meanwhile, [`packages/coding-agent/src/core/system-prompt.ts`](https://github.com/earendil-works/pi/blob/main/packages/coding-agent/src/core/system-prompt.ts) injects a description of the current thinking level into the system prompt, ensuring the model understands its constraints.

## Combining Budgets with Output Limits

For workflows requiring strict resource constraints, combine the thinking level with the `max_output_tokens` provider option (defined in [`model-config.ts`](https://github.com/earendil-works/pi/blob/main/model-config.ts)). This creates a hard ceiling on total output while preserving the allocated reasoning budget:

```typescript
const config = {
  thinkingLevel: "high",
  max_output_tokens: 4000  // Hard limit including thinking + response
};

```

This configuration ensures the model thinks within the ~300 token budget but stops generation entirely at 4000 tokens.

## Summary

- The thinking budget in earendil pi uses five discrete levels (`off` through `max`) mapped to approximate token allowances
- Default configuration lives in `~/.pi/agent/settings.json` under the `thinkingLevel` key
- Runtime adjustments work via **Shift + Tab**, `/thinking` commands, or the `PI_THINKING_LEVEL` environment variable
- `SessionManager` stores state, [`model-resolver.ts`](https://github.com/earendil-works/pi/blob/main/model-resolver.ts) applies it to provider requests, and `TuiRenderer` displays thinking streams with theme-specific colors
- The system prompt builder automatically informs the model of its current thinking constraints

## Frequently Asked Questions

### What is the difference between thinking level and max_output_tokens?

The **thinking level** controls internal deliberation tokens before the model produces visible output, while **max_output_tokens** sets a hard ceiling on the total response length including both thinking and final output. According to the source code in [`packages/coding-agent/src/core/model-resolver.ts`](https://github.com/earendil-works/pi/blob/main/packages/coding-agent/src/core/model-resolver.ts), the thinking level is passed to the provider's thinking parameter, whereas `max_output_tokens` caps the entire generation. Use them together when you need reasoning capability with strict overall limits.

### How do I disable thinking output completely in earendil pi?

Set the thinking level to `off` either in your `~/.pi/agent/settings.json` configuration file or by typing `/thinking off` during a session. When set to `off`, the `getThinkingLevel` function in [`packages/coding-agent/src/utils/config.ts`](https://github.com/earendil-works/pi/blob/main/packages/coding-agent/src/utils/config.ts) returns zero tokens, and `TuiRenderer` displays thinking blocks in gray text (or suppresses them entirely depending on the theme configuration).

### Can I programmatically check the current thinking budget?

Yes. Import `SessionManager` from `@earendil-works/pi-coding-agent` and access the `thinkingLevel` property:

```typescript
import { SessionManager } from "@earendil-works/pi-coding-agent";

const session = SessionManager.current();
console.log(session.thinkingLevel); // "medium", "high", etc.

```

This reads the live session state maintained in [`packages/coding-agent/src/core/session-manager.ts`](https://github.com/earendil-works/pi/blob/main/packages/coding-agent/src/core/session-manager.ts), reflecting any runtime changes made via slash commands or keyboard shortcuts.

### Why does the thinking budget use approximate token counts?

The budget is a **soft limit** (as noted in the [`packages/coding-agent/src/core/model-resolver.ts`](https://github.com/earendil-works/pi/blob/main/packages/coding-agent/src/core/model-resolver.ts) implementation) because Pi delegates token enforcement to the underlying LLM provider's native reasoning features. The client buffers and displays thinking deltas but relies on the provider to halt reasoning generation. The approximate values (~50, ~150, ~300 tokens) represent heuristic targets that align with provider-specific implementations of reasoning depth controls.