# How Reasoning Effort Parameters (Low, Medium, High, Max) Affect Harvey AI Model Behavior

> Discover how Harvey AI's reasoning effort parameters (low, medium, high, max) impact model behavior. Improve complex problem-solving by understanding token allocation and latency trade-offs.

- Repository: [Harvey/harvey-labs](https://github.com/harveyai/harvey-labs)
- Tags: deep-dive
- Published: 2026-08-11

---

**Higher reasoning effort parameters make Harvey AI models allocate more tokens to internal chain-of-thought steps, improving complex problem-solving at the cost of increased latency and token usage.**

The **reasoning effort parameter** in Harvey Labs is a provider-agnostic control that lets developers tune how deeply an LLM thinks before responding. According to the harveyai/harvey-labs source code, this single parameter—set via CLI as `--reasoning-effort` or programmatically as `reasoning_effort`—gets translated by provider-specific adapters into each backend's native reasoning controls.

## How Reasoning Effort Flows Through the Harness

The Harvey Labs architecture separates generic configuration from provider-specific implementation. The **harness** ([`harness/run.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/run.py)) parses user input and stores the effort level in `args.reasoning_effort`【/cache/repos/github.com/harveyai/harvey-labs/main/harness/run.py】. This value propagates to all adapters through the base class in [`harness/adapters/base.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/adapters/base.py), which maintains the `reasoning_effort` attribute uniformly across providers.

Each adapter then maps this generic value to the exact API structure its provider expects.

### OpenAI Adapter: Native Reasoning Object

The `OpenAIAdapter` constructs a top-level `reasoning` object for the Responses API:

```python

# From harness/adapters/openai.py lines 53-55

reasoning = {
    "effort": self.reasoning_effort,  # low, medium, high, etc.

    "summary": "auto"
}

```

**Supported values:** `none`, `minimal`, `low`, `medium`, `high`, `xhigh`

The OpenAI model reserves a portion of its context window for internal reasoning tokens. Higher effort levels expand this reserved budget, enabling more thorough exploration before final output generation.

### Mistral Adapter: Conditional Reasoning Flag

The `MistralAdapter` applies reasoning effort only to designated reasoning models:

```python

# From harness/adapters/mistral.py lines 57-59

if self.model in REASONING_MODELS:
    request["reasoning_effort"] = self.reasoning_effort

```

**Supported values:** `none`, `high` (currently functional)

Models like `mistral-medium-3.5` and `mistral-small-2603` expose this flag. When enabled, the model emits intermediate thinking fragments that improve multi-step reasoning tasks.

### Anthropic Adapter: Adaptive Output Configuration

The `AnthropicAdapter` injects effort into `extra_body` for adaptive models:

```python

# From harness/adapters/anthropic.py lines 90-93

extra_body = {
    "output_config": {
        "effort": self.reasoning_effort  # low, medium, high

    }
}

```

**Supported values:** `low`, `medium`, `high` (extensible for future adaptive levels)

Anthropic's adaptive models dynamically scale their "thinking" token allocation based on this parameter. The effect is most visible on complex analytical tasks where shallow responses would otherwise miss edge cases.

### Google Adapter: Thinking Level Translation

The `GoogleAdapter` maps generic effort to Gemini-specific thinking levels via `THINKING_LEVEL_MAP`:

```python

# From harness/adapters/google.py lines 66-68

thinking_level = THINKING_LEVEL_MAP.get(
    self.reasoning_effort, "fast"
)

```

**Mapping example:** `low → "slow"` (more extensive planning), with other levels mapped accordingly

Gemini 1.5+ uses this to control its internal planning phase duration before token generation begins.

### Fireworks and Baseten Adapters

- **Fireworks ([`harness/adapters/fireworks.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/adapters/fireworks.py) lines 42-45):** Forwards `reasoning_effort` in `extra_body`
- **Baseten ([`harness/adapters/baseten.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/adapters/baseten.py) lines 61-63):** Passes as a provider-specific flag

Both follow the same pattern: higher effort triggers more internal reasoning steps within their respective inference stacks.

## Practical Usage Examples

### Programmatic API

```python
from harvey_labs.harness.run import run

run(
    task="complex_analysis",
    model="gpt-4o-mini",
    reasoning_effort="high",      # Options: low, medium, high, max, xhigh

    temperature=0.0,
)

```

### Direct Adapter Invocation

```python
from harvey_labs.harness.adapters.openai import OpenAIAdapter

adapter = OpenAIAdapter(
    model="gpt-4o-mini",
    temperature=0.0,
    reasoning_effort="medium",
)
response = adapter.chat(
    messages=[{"role": "user", "content": "Explain quantum entanglement step by step"}],
    tools=[]
)
print(response.text)

```

### Command Line Interface

```bash
python -m harvey_labs.main.harness.run \
    --task legal_analysis \
    --model gpt-4o-mini \
    --reasoning-effort high \
    --temperature 0.0

```

## Key Trade-offs When Adjusting Reasoning Effort

| Effort Level | Best For | Latency | Token Cost |
|-------------|----------|---------|-----------|
| **low** | Simple classification, factual retrieval | Fastest | Lowest |
| **medium** | Standard document analysis, summarization | Moderate | Moderate |
| **high** / **max** | Complex reasoning, legal analysis, multi-step planning | Slowest | Highest |

## Summary

- The `reasoning_effort` parameter in Harvey Labs provides unified control over model reasoning depth across OpenAI, Mistral, Anthropic, Google, and other providers
- Each adapter in `harness/adapters/` translates the generic value to provider-native API fields
- Higher effort increases internal chain-of-thought token allocation, improving answer quality on complex tasks
- Trade-offs include increased latency and token consumption, with provider-specific limits on supported values
- Use CLI `--reasoning-effort` or programmatic `reasoning_effort` for consistent behavior regardless of backend

## Frequently Asked Questions

### What happens if I set an unsupported reasoning effort value?

Adapters guard against invalid values. For example, the `MistralAdapter` only adds the `reasoning_effort` field when the model appears in `REASONING_MODELS`. Unsupported values are either ignored or mapped to defaults by the underlying provider API, though explicit errors may occur depending on the provider's validation strictness.

### Does reasoning effort work with temperature and other parameters?

Yes. `reasoning_effort` operates independently of `temperature`, `top_p`, and other sampling parameters. You can combine high reasoning effort with `temperature=0` for deterministic, thorough analysis, or with higher temperature for creative exploration with deep reasoning structure.

### Which Harvey AI models support the "max" or "xhigh" effort levels?

These extended levels are currently OpenAI-specific (`xhigh`) or supported through custom provider configurations. Check your target model's documentation: OpenAI's latest reasoning models accept `none` through `xhigh`, while Anthropic and Google currently support `low`/`medium`/`high`. The harness adapters automatically handle unsupported level gracefully where possible.