# Progressive Throttling in ctx_search: Triggers, Limits, and Effects on Query Results

> Learn about progressive throttling in ctx_search. Discover triggers, limits, and how it impacts query results, ensuring efficient API usage. Protect your service today.

- Repository: [Mert Köseoğlu/context-mode](https://github.com/mksglu/context-mode)
- Tags: deep-dive
- Published: 2026-04-24

---

**Progressive throttling in `ctx_search` activates when you exceed three search calls within a rolling 60‑second window, reducing results per query from two to one after the fourth call and blocking all requests entirely after the eighth call.**

The `ctx_search` tool in the `mksglu/context-mode` repository implements progressive throttling to protect the knowledge base from context flooding and encourage efficient batch querying. This mechanism monitors call frequency within a 60‑second rolling window and automatically degrades service quality to prevent prompt overload from excessive individual search calls.

## What Triggers Progressive Throttling in ctx_search?

Throttling is driven entirely by **call volume** tracked in a rolling 60‑second window. The core logic resides in [`src/server.ts`](https://github.com/mksglu/context-mode/blob/main/src/server.ts), where the implementation declares two critical thresholds at lines 1122–1128:

- **`SEARCH_MAX_RESULTS_AFTER`**: Defaults to `3`. Calls at or below this count operate with full result limits.
- **`SEARCH_BLOCK_AFTER`**: Defaults to `8`. Calls exceeding this threshold trigger a hard block.

Each request updates an internal counter (lines 1128–1135). When the 60‑second window expires, the counter resets automatically. The trigger is purely numerical—based on how many times the tool is invoked—regardless of query complexity or result size.

## How Progressive Throttling in ctx_search Affects Query Results

The impact on your search results depends on which threshold zone your current call falls into.

### Normal Operation (Calls 1–3)

When the call count is less than or equal to `SEARCH_MAX_RESULTS_AFTER` (default `3`), the tool returns up to **2 results per query** (or fewer if you specify a lower `limit` parameter). Full context snippets are returned with no warning messages appearing.

### Throttled Mode (Calls 4–8)

Once the counter exceeds `SEARCH_MAX_RESULTS_AFTER` but remains at or below `SEARCH_BLOCK_AFTER`, the system enforces stricter limits:

- The per‑query result limit is forced to **1** (see lines 1250–1252 in [`src/server.ts`](https://github.com/mksglu/context-mode/blob/main/src/server.ts)).
- A warning banner is appended to the response indicating the current call number and suggesting batch execution (lines 1286–1290).

This represents graceful degradation. The model receives minimal functionality while being alerted to consolidate queries via the warning message:

```

⚠ search call #4/8 in this window. Results limited to 1/query.
Batch queries: search(queries: ["q1","q2","q3"]) or use batch_execute.

```

### Blocked State (Calls 9+)

When the call count exceeds `SEARCH_BLOCK_AFTER` (default `8`), the tool refuses to perform the search entirely. The implementation at lines 1136–1146 in [`src/server.ts`](https://github.com/mksglu/context-mode/blob/main/src/server.ts) returns an error response instructing the model to stop making individual search calls:

```

BLOCKED: 9 search calls in 12s. You're flooding context. STOP making individual search calls.
Use batch_execute(commands, queries) for your next research step.

```

## Code Examples: Normal Usage to Blocked Requests

### Normal Single Query (Calls 1–3)

```typescript
// Returns up to 2 results per query
await ctx_search({ queries: ["How does progressive throttling work?"] });

```

### Entering Throttled Mode (Call 4)

```typescript
// Calls 1-3 operate normally
await ctx_search({ queries: ["Q1"] });
await ctx_search({ queries: ["Q2"] });
await ctx_search({ queries: ["Q3"] });

// Call 4 triggers throttling - returns only 1 result plus warning
await ctx_search({ queries: ["Q4"] });

```

### Hitting the Block Limit (Call 9)

```typescript
// Consume the entire allowed window
for (let i = 1; i <= 8; i++) {
  await ctx_search({ queries: [`Q${i}`] });
}

// Call 9 is rejected with an error
await ctx_search({ queries: ["Q9"] });

```

### Optimal Pattern: Batch Execution

To avoid throttling entirely, batch all research questions in a single call:

```typescript
await ctx_search({
  queries: [
    "Explain progressive throttling",
    "Show the code that limits results",
    "What error is shown after the block limit?"
  ],
  limit: 3
});

```

This executes all searches in one request, staying well below any threshold and preventing context overload (each search can pull in up to ~40 KB of snippets).

## Key Implementation Files

- **[`src/server.ts`](https://github.com/mksglu/context-mode/blob/main/src/server.ts)**: Contains the throttling logic, counter management, and response shaping (lines 1122–1290).
- **[`src/store.ts`](https://github.com/mksglu/context-mode/blob/main/src/store.ts)**: Houses the underlying `searchWithFallback` method that retrieves matching sections after throttling checks pass.
- **[`src/types.ts`](https://github.com/mksglu/context-mode/blob/main/src/types.ts)**: Defines the Zod input schema for `ctx_search` including `queries`, `limit`, `source`, and `contentType` parameters.
- **[`tests/guidance-throttle.test.ts`](https://github.com/mksglu/context-mode/blob/main/tests/guidance-throttle.test.ts)**: Provides unit test patterns for verifying throttling behavior.

## Summary

- Progressive throttling in `ctx_search` triggers on **call volume** within a rolling 60‑second window, not query complexity.
- The first three calls return **up to 2 results** per query without restrictions.
- Calls four through eight are **limited to 1 result** per query and include a throttle warning banner.
- Call nine and beyond are **blocked entirely**, returning an error that forces use of `batch_execute`.
- **Batching queries** in a single `ctx_search` call avoids throttling and improves efficiency.

## Frequently Asked Questions

### What is the exact threshold before throttling starts in ctx_search?

Throttling begins after the third call. The `SEARCH_MAX_RESULTS_AFTER` constant defaults to `3`, meaning calls 1–3 operate normally with up to 2 results per query, while call 4 and beyond enter the throttled state with reduced result counts and warning banners.

### How long does the throttling window last in ctx_search?

The rolling window lasts **60 seconds**. The counter tracks how many calls occur within this period and automatically resets when the window expires, as implemented in the request handling logic in [`src/server.ts`](https://github.com/mksglu/context-mode/blob/main/src/server.ts) (lines 1128–1135).

### Can I bypass throttling by using different search parameters?

No. Throttling is based purely on **call frequency**, not the content of your queries or the `limit` parameter values. The only way to avoid throttling is to batch multiple queries into a single `ctx_search` call or use the `batch_execute` method for complex research workflows.

### Where is the throttling logic implemented in the source code?

The core throttling mechanism resides in **[`src/server.ts`](https://github.com/mksglu/context-mode/blob/main/src/server.ts)** at lines 1122–1146, which handles the counter logic and blocking, and lines 1250–1290, which enforce result limits and append warning messages. The underlying search functionality that executes after these guard checks is located in [`src/store.ts`](https://github.com/mksglu/context-mode/blob/main/src/store.ts).