How Reasoning Effort Parameters (Low, Medium, High, Max) Affect Harvey AI Model Behavior

Higher reasoning effort parameters make Harvey AI models allocate more tokens to internal chain-of-thought steps, improving complex problem-solving at the cost of increased latency and token usage.

The reasoning effort parameter in Harvey Labs is a provider-agnostic control that lets developers tune how deeply an LLM thinks before responding. According to the harveyai/harvey-labs source code, this single parameter—set via CLI as --reasoning-effort or programmatically as reasoning_effort—gets translated by provider-specific adapters into each backend's native reasoning controls.

How Reasoning Effort Flows Through the Harness

The Harvey Labs architecture separates generic configuration from provider-specific implementation. The harness (harness/run.py) parses user input and stores the effort level in args.reasoning_effort【/cache/repos/github.com/harveyai/harvey-labs/main/harness/run.py】. This value propagates to all adapters through the base class in harness/adapters/base.py, which maintains the reasoning_effort attribute uniformly across providers.

Each adapter then maps this generic value to the exact API structure its provider expects.

OpenAI Adapter: Native Reasoning Object

The OpenAIAdapter constructs a top-level reasoning object for the Responses API:


# From harness/adapters/openai.py lines 53-55

reasoning = {
    "effort": self.reasoning_effort,  # low, medium, high, etc.

    "summary": "auto"
}

Supported values: none, minimal, low, medium, high, xhigh

The OpenAI model reserves a portion of its context window for internal reasoning tokens. Higher effort levels expand this reserved budget, enabling more thorough exploration before final output generation.

Mistral Adapter: Conditional Reasoning Flag

The MistralAdapter applies reasoning effort only to designated reasoning models:


# From harness/adapters/mistral.py lines 57-59

if self.model in REASONING_MODELS:
    request["reasoning_effort"] = self.reasoning_effort

Supported values: none, high (currently functional)

Models like mistral-medium-3.5 and mistral-small-2603 expose this flag. When enabled, the model emits intermediate thinking fragments that improve multi-step reasoning tasks.

Anthropic Adapter: Adaptive Output Configuration

The AnthropicAdapter injects effort into extra_body for adaptive models:


# From harness/adapters/anthropic.py lines 90-93

extra_body = {
    "output_config": {
        "effort": self.reasoning_effort  # low, medium, high

    }
}

Supported values: low, medium, high (extensible for future adaptive levels)

Anthropic's adaptive models dynamically scale their "thinking" token allocation based on this parameter. The effect is most visible on complex analytical tasks where shallow responses would otherwise miss edge cases.

Google Adapter: Thinking Level Translation

The GoogleAdapter maps generic effort to Gemini-specific thinking levels via THINKING_LEVEL_MAP:


# From harness/adapters/google.py lines 66-68

thinking_level = THINKING_LEVEL_MAP.get(
    self.reasoning_effort, "fast"
)

Mapping example: low → "slow" (more extensive planning), with other levels mapped accordingly

Gemini 1.5+ uses this to control its internal planning phase duration before token generation begins.

Fireworks and Baseten Adapters

Both follow the same pattern: higher effort triggers more internal reasoning steps within their respective inference stacks.

Practical Usage Examples

Programmatic API

from harvey_labs.harness.run import run

run(
    task="complex_analysis",
    model="gpt-4o-mini",
    reasoning_effort="high",      # Options: low, medium, high, max, xhigh

    temperature=0.0,
)

Direct Adapter Invocation

from harvey_labs.harness.adapters.openai import OpenAIAdapter

adapter = OpenAIAdapter(
    model="gpt-4o-mini",
    temperature=0.0,
    reasoning_effort="medium",
)
response = adapter.chat(
    messages=[{"role": "user", "content": "Explain quantum entanglement step by step"}],
    tools=[]
)
print(response.text)

Command Line Interface

python -m harvey_labs.main.harness.run \
    --task legal_analysis \
    --model gpt-4o-mini \
    --reasoning-effort high \
    --temperature 0.0

Key Trade-offs When Adjusting Reasoning Effort

Effort Level Best For Latency Token Cost
low Simple classification, factual retrieval Fastest Lowest
medium Standard document analysis, summarization Moderate Moderate
high / max Complex reasoning, legal analysis, multi-step planning Slowest Highest

Summary

  • The reasoning_effort parameter in Harvey Labs provides unified control over model reasoning depth across OpenAI, Mistral, Anthropic, Google, and other providers
  • Each adapter in harness/adapters/ translates the generic value to provider-native API fields
  • Higher effort increases internal chain-of-thought token allocation, improving answer quality on complex tasks
  • Trade-offs include increased latency and token consumption, with provider-specific limits on supported values
  • Use CLI --reasoning-effort or programmatic reasoning_effort for consistent behavior regardless of backend

Frequently Asked Questions

What happens if I set an unsupported reasoning effort value?

Adapters guard against invalid values. For example, the MistralAdapter only adds the reasoning_effort field when the model appears in REASONING_MODELS. Unsupported values are either ignored or mapped to defaults by the underlying provider API, though explicit errors may occur depending on the provider's validation strictness.

Does reasoning effort work with temperature and other parameters?

Yes. reasoning_effort operates independently of temperature, top_p, and other sampling parameters. You can combine high reasoning effort with temperature=0 for deterministic, thorough analysis, or with higher temperature for creative exploration with deep reasoning structure.

Which Harvey AI models support the "max" or "xhigh" effort levels?

These extended levels are currently OpenAI-specific (xhigh) or supported through custom provider configurations. Check your target model's documentation: OpenAI's latest reasoning models accept none through xhigh, while Anthropic and Google currently support low/medium/high. The harness adapters automatically handle unsupported level gracefully where possible.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →