Agno Reasoning and reasoning_model Capabilities: A Complete Technical Guide
TLDR: Agno's reasoning subsystem provides unified native reasoning model support with automatic provider detection, streaming and non-streaming execution, async variants, and a fallback Chain-of-Thought engine for non-native models, all managed through the ReasoningManager class.
The agno-agi/agno repository implements a sophisticated reasoning layer that enables AI agents to leverage native "thinking" capabilities from modern LLM providers. This article explores the complete capabilities of Agno's reasoning engine and reasoning_model support, covering everything from provider-specific detection logic to event-driven monitoring and tool-level reasoning persistence.
Core Architecture of the Agno Reasoning Engine
At the heart of Agno's reasoning capabilities lies the ReasoningManager class in libs/agno/agno/reasoning/manager.py. This class provides a unified API that abstracts away provider-specific implementations while handling model detection, execution orchestration, and event emission.
The manager relies on ReasoningConfig (defined in the same file at lines 76-89) to control behavior. Configuration options include min_steps, max_steps, tool enablement, tool call limits, JSON mode toggles, telemetry settings, and custom RunContext objects.
Native Reasoning Model Detection and Support
Agno automatically detects whether a supplied model instance supports native reasoning through the is_native_reasoning_model method. This detection logic delegates to provider-specific checker functions located in individual modules such as libs/agno/agno/reasoning/openai.py, anthropic.py, deepseek.py, gemini.py, groq.py, ollama.py, azure_ai_foundry.py, and vertexai.py.
Provider-Specific Implementation Details
Each provider implements a checker function that determines native reasoning eligibility based on model IDs and configuration flags:
-
OpenAI:
is_openai_reasoning_modelinlibs/agno/agno/reasoning/openai.py(lines 11-27) detects models includinggpt-4o,gpt-3.5-turbo,gpt-4.1,gpt-4.5, andgpt-5.1. -
Anthropic:
is_anthropic_reasoning_modelchecks forthinking=Trueparameter support inClaudemodels such asclaude-3-5-sonnet-20241022. -
DeepSeek: Detection relies on model ID substrings "reasoner" or "r1" (e.g.,
deepseek-deepseek-r1,deepseek-reasoner). -
Google Gemini, Groq, Ollama, Azure AI Foundry, and Vertex AI each implement similar detection logic with provider-specific flags such as
thinkingparameters or model ID patterns.
Execution Modes: Streaming, Non-Streaming, and Async
The ReasoningManager supports multiple execution patterns to accommodate different latency and interaction requirements.
Non-Streaming Native Reasoning
For synchronous execution, the manager calls provider-specific helpers such as get_openai_reasoning, get_anthropic_reasoning, or get_deepseek_reasoning. These return a ReasoningResult containing the assistant message, extracted thinking content, and a list of ReasoningStep objects. The implementation resides in libs/agno/agno/reasoning/manager.py at lines 96-132.
Streaming Native Reasoning
Streaming mode yields incremental "thinking" chunks while the model generates output. The stream_native_reasoning method (and its provider variants like get_openai_reasoning_stream) yields tuples of (delta, None) for each chunk, concluding with (None, result) when the final message is ready. This implementation appears in libs/agno/agno/reasoning/manager.py at lines 150-197.
Async Variants
All execution modes provide async equivalents: aget_native_reasoning, astream_native_reasoning, and arun_default_reasoning. These async implementations occupy lines 215-285 in libs/agno/agno/reasoning/manager.py.
Default Chain-of-Thought Fallback
When is_native_reasoning_model returns False, the ReasoningManager automatically falls back to a generic Chain-of-Thought (CoT) loop. This fallback uses a dedicated reasoning agent configured with a structured output schema based on ReasoningSteps (defined in libs/agno/agno/reasoning/step.py).
The fallback logic, implemented in _run_default_reasoning_events and run_default_reasoning (lines 286-363 of manager.py), iteratively runs the agent, yields each ReasoningStep, and terminates when the agent returns NextAction.FINAL_ANSWER or reaches the max_steps limit configured in ReasoningConfig.
Event-Driven Monitoring with ReasoningEvent
The reasoning subsystem emits a structured event stream through the ReasoningEvent class, enabling real-time monitoring and UI updates. Event types include:
started– Reasoning process initiatedcontent_delta– Incremental thinking content (streaming)step– Completion of a discrete reasoning stepcompleted– Final answer availableerror– Processing failure
These events are generated throughout ReasoningManager methods and can be consumed by higher-level Agent or Team runtimes to surface progress to users or logs.
Tool-Level Reasoning with ReasoningTools
For agents requiring explicit thought recording, Agno provides the ReasoningTools toolkit in libs/agno/agno/tools/reasoning.py. This toolkit exposes two high-level tools:
think– Records scratch-pad thoughts with title, reasoning, action, and confidenceanalyze– Evaluates previous thoughts against specified criteria
When the LLM calls these tools, Agno automatically serializes ReasoningStep objects into run_context.session_state["reasoning_steps"]. You can later retrieve and validate these using ReasoningStep.model_validate_json() for audit trails or multi-step reasoning workflows.
Configuration and Usage Examples
Detecting Native Reasoning Support
from agno.reasoning.openai import is_openai_reasoning_model
from agno.models.openai import OpenAIChat
model = OpenAIChat(id="gpt-4o")
print(is_openai_reasoning_model(model)) # → True
from agno.reasoning.anthropic import is_anthropic_reasoning_model
from agno.models.anthropic import Claude
model2 = Claude(id="claude-3-5-sonnet-20241022", thinking=True)
print(is_anthropic_reasoning_model(model2)) # → True
Using Native Reasoning with Streaming
from agno.agent import Agent
from agno.models.openai import OpenAIChat
from agno.run import Message
from agno.reasoning.manager import ReasoningEventType
model = OpenAIChat(id="gpt-4o")
agent = Agent(name="Researcher", reasoning_model=model)
messages = [
Message(role="system", content="You are a helpful assistant."),
Message(role="user", content="Explain quantum entanglement."),
]
for event in agent.reason(messages, stream=True):
if event.event_type == ReasoningEventType.content_delta:
print("Thinking:", event.reasoning_content)
elif event.event_type == ReasoningEventType.completed:
print("Final:", event.message.content)
Implementing Fallback CoT Reasoning
from agno.agent import Agent
from agno.models.base import Model
from agno.run import Message
from agno.reasoning.manager import ReasoningEventType
# Non-native model triggers automatic CoT fallback
model = Model(id="gpt-3.5-turbo-mini")
agent = Agent(name="Planner", reasoning_model=model)
for ev in agent.reason([Message(role="user", content="Plan a weekend trip.")]):
if ev.event_type == ReasoningEventType.step:
step = ev.reasoning_step
print(f"Step {step.title}: {step.reasoning}")
elif ev.event_type == ReasoningEventType.completed:
for s in ev.reasoning_steps:
print(f"- {s.title} → {s.result}")
Recording Custom Thoughts with ReasoningTools
from agno.agent import Agent
from agno.tools.reasoning import ReasoningTools
from agno.run import RunContext
from agno.reasoning.step import ReasoningStep
agent = Agent(name="Investigator")
agent.run_context = RunContext()
agent.run_context.tools = [ReasoningTools(enable_think=True, enable_analyze=True)]
# When the LLM calls think(title="...", thought="...", action="...", confidence=...)
# the tool stores the step in session_state:
run_state = agent.run_context.session_state
if "reasoning_steps" in run_state:
steps_json = run_state["reasoning_steps"].get(agent.run_context.run_id, [])
for json_str in steps_json:
step = ReasoningStep.model_validate_json(json_str)
print(f"{step.title} → {step.reasoning}")
Summary
-
Unified API: The
ReasoningManagerinlibs/agno/agno/reasoning/manager.pyprovides a single interface for all reasoning operations, abstracting provider-specific implementations. -
Automatic Provider Detection: Nine major providers (OpenAI, Anthropic, DeepSeek, Gemini, Groq, Ollama, Azure AI Foundry, Vertex AI) are supported through dedicated checker functions that detect native reasoning capabilities based on model IDs and configuration flags.
-
Flexible Execution Modes: Supports synchronous (
get_*_reasoning), streaming (stream_*_reasoning), and async variants (aget_*,astream_*) to accommodate different latency requirements. -
Robust Fallback Mechanism: Automatically falls back to a deterministic Chain-of-Thought loop using structured output schemas (
ReasoningSteps) when native reasoning is unavailable. -
Event-Driven Architecture: Emits
ReasoningEventobjects (started, content_delta, step, completed, error) for real-time monitoring and UI integration. -
Tool-Level Persistence: The
ReasoningToolstoolkit enables explicit thought recording viathinkandanalyzetools, automatically persistingReasoningStepobjects tosession_state.
Frequently Asked Questions
What is the difference between native reasoning and default CoT in Agno?
Native reasoning leverages provider-specific APIs that expose the model's internal "thinking" process, such as OpenAI's reasoning models or Anthropic's thinking=True mode. This provides access to raw reasoning tokens and streaming deltas. Default Chain-of-Thought (CoT) is Agno's fallback implementation that uses a structured output schema (ReasoningSteps) and iterative agent loops to generate reasoning steps when the underlying model does not support native thinking APIs.
How do I enable streaming for reasoning models in Agno?
Pass stream=True to the agent.reason() method or directly invoke ReasoningManager.stream_native_reasoning(). The system returns ReasoningEvent objects with event_type of content_delta for incremental thinking chunks and completed for the final answer. This works for all supported native providers including OpenAI, Anthropic, DeepSeek, and Gemini.
Which LLM providers support native reasoning in Agno?
Agno supports native reasoning detection and streaming for nine major providers: OpenAI (GPT-4o, GPT-4.1, GPT-4.5), Anthropic (Claude 3.5 Sonnet with thinking=True), DeepSeek (R1/reasoner models), Google Gemini, Groq, Ollama, Azure AI Foundry, and Vertex AI. Each provider has a dedicated checker function (e.g., is_openai_reasoning_model) in the libs/agno/agno/reasoning/ directory.
How can I persist reasoning steps across agent tool calls?
Use the ReasoningTools toolkit from libs/agno/agno/tools/reasoning.py. Enable the think and analyze tools on your agent. When the LLM invokes these tools, Agno automatically serializes ReasoningStep objects into run_context.session_state["reasoning_steps"]. You can later retrieve these using ReasoningStep.model_validate_json() to reconstruct the reasoning history for audit trails or subsequent agent runs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →