How the ReAct Loop Works in WeKnora's `internal/agent` Package

The ReAct loop in WeKnora's internal/agent package implements a classic Reason→Act pattern where AgentEngine.Execute orchestrates iterative LLM calls through executeLoop, with each round handled by runReActIteration to decide whether to continue thinking, execute tools, or return a final answer.

WeKnora, Tencent's open-source agent framework, provides a production-ready implementation of the ReAct (Reasoning and Acting) agent pattern in its internal/agent directory. The implementation follows an iterative cycle where the agent thinks, acts, and observes results until reaching a conclusion or exhausting its iteration budget. The core logic resides in three primary files: engine.go for orchestration, think.go for LLM communication, and act.go for tool execution.

Architecture of the ReAct Implementation

The ReAct loop spans three critical files within internal/agent:

  • engine.go: Contains the orchestration logic including AgentEngine.Execute, executeLoop, and runReActIteration
  • think.go: Implements resilient LLM invocation via callLLMWithRetry
  • act.go: Handles tool execution through executeToolCalls and runToolCall

These components work together to provide a robust loop with built-in context management, retry logic, and comprehensive observability via Langfuse tracing.

Entry Point: AgentEngine.Execute

The AgentEngine.Execute function in engine.go serves as the primary entry point for agent execution. It accepts a context.Context, session identifiers, the user query, previous message history, and optional image URLs.

func (e *AgentEngine) Execute(
    ctx context.Context,
    sessionID, messageID, query string,
    llmContext []chat.Message, imageURLs ...[]string,
) (*types.AgentState, error) {
    // …build system prompt, initialise messages, register tools…
    _, err := e.executeLoop(ctx, state, query, messages, tools, sessionID, messageID)
    // …final logging and return…
}

This method prepares the conversation context—including system prompts and any pinned mentions—initializes the AgentState, and guarantees resource cleanup via defer e.toolRegistry.Cleanup(ctx). After setup, it delegates control to executeLoop to run the actual ReAct cycle.

The Main Loop: executeLoop

executeLoop manages the iteration lifecycle using budget constraints and explicit control outcomes. It tracks safety variables including emptyRetries and consecutiveSameContent to prevent infinite loops.

func (e *AgentEngine) executeLoop(
    ctx context.Context,
    state *types.AgentState,
    query string,
    messages []chat.Message,
    tools []chat.Tool,
    sessionID, messageID string,
) (*types.AgentState, error) {
    emptyRetries, consecutiveSameContent := 0, 0
    lastResponseContent := ""

    for e.withinIterationBudget(state.CurrentRound) || e.allowSteerOverrun {
        outcome, iterErr := e.runReActIteration(
            ctx, state, &messages, tools,
            sessionID, messageID, query,
            &emptyRetries, &consecutiveSameContent, &lastResponseContent,
        )
        if iterErr != nil { return state, iterErr }

        switch outcome {
        case iterOutcomeContinue:   continue
        case iterOutcomeBreak:      break
        case iterOutcomeNext:       state.CurrentRound++
        }
    }
    // Handle max iterations if incomplete...
    return state, nil
}

The loop continues while withinIterationBudget returns true—checking against the configured MaxIterations—or while allowSteerOverrun permits one extra round after a user steering message. Three sentinel outcomes control flow:

  • iterOutcomeContinue: Retry the current round (typically for empty content)
  • iterOutcomeBreak: Exit the loop with final answer or fatal error
  • iterOutcomeNext: Proceed to the next ReAct round

Inside a Single ReAct Iteration

runReActIteration implements the core Think→Analyze→Act→Observe cycle. Each invocation represents one complete pass through the ReAct pattern.

Context Management and Steering

Each iteration begins with context window estimation and potential compaction via manageContextWindow. The function also drains any mid-run user messages through drainSteerMessages, allowing users to inject steering commands during execution.

The Think Phase

The engine invokes the LLM through callLLMWithRetry from think.go. If the response hits context limits and overflowRecovered is false, the engine triggers forceCompaction and retries once.

resp, err := e.callLLMWithRetry(ctx, messagesPtr, tools, state, query,
                                state.CurrentRound, sessionID)
if err != nil { return iterOutcomeNext, err }
if resp == nil { return iterOutcomeBreak, nil }

The Analyze Phase

analyzeResponse evaluates the LLM output to determine the next step. If the response contains no tool calls and signals completion (verdict.isDone), the loop prepares to break. The engine also detects stuck loops by comparing content against lastResponseContent, aborting after maxRepeatedResponseRounds identical non-tool responses.

The Act Phase

When the LLM returns tool calls, executeToolCalls routes execution to runToolCall for each tool. The engine supports parallel execution based on the ParallelToolCalls configuration flag.

The Observe Phase

Results from tool executions append back to the conversation history via appendToolResults and appendToolImages. The AgentStep struct records the thought, tool calls, and timestamps for the round history.

state.RoundSteps = append(state.RoundSteps, step)
*messagesPtr = e.appendToolResults(*messagesPtr, step)
*messagesPtr = e.appendToolImages(ctx, *messagesPtr, step)

Resilient LLM Communication with callLLMWithRetry

callLLMWithRetry in think.go provides the sole LLM interface for the engine. It sanitizes messages to remove secret data, retries transient network errors a configurable number of times, and returns nil only when the model explicitly finishes without a response. This centralized wrapper ensures every round has a well-structured request/response pair with consistent error handling.

Tool Execution and Parallelization

runToolCall in act.go handles individual tool invocations through a structured pipeline:

  1. Parses JSON arguments with repair logic for malformed JSON
  2. Resolves temporary model handles to concrete MCP targets
  3. Emits tool-hint events for UI progress feedback
  4. Creates Langfuse spans named agent.tool.<name>
  5. Executes via e.toolRegistry.ExecuteTool with toolExecutionTimeout constraints
  6. Wraps results in types.ToolCall structures and emits tool-result events

The engine refuses truncated tool calls when finish reasons indicate token limits, marking them failed with truncatedArgumentsError to prevent execution of partial arguments.

Observability and Tracing with Langfuse

Every ReAct operation generates hierarchical Langfuse spans for comprehensive tracing:

  • agent.execute: Root span for the entire execution
  • agent.round.<N>: Individual iteration spans capturing input prompts and token usage
  • agent.tool.<name>: Tool execution spans recording duration, success flags, and error information

These spans attach to the parent execution context, providing a complete audit trail of the agent's reasoning and actions.

Edge Cases and Safety Mechanisms

The implementation includes several safeguards for production reliability:

Situation Mitigation Strategy
Context overflow After first detection, forceCompaction compresses history with a single retry (overflowRecovered flag)
Empty LLM replies Up to maxEmptyResponseRetries, the engine injects continuation prompts and retries (emptyRetries counter)
Stuck loops Tracks consecutiveSameContent against lastResponseContent, aborting after maxRepeatedResponseRounds identical responses
Truncated arguments Detects token-cap finish reasons and marks tool calls failed with truncatedArgumentsError
Steer-overrun Allows one extra iteration (allowSteerOverrun) when users inject messages during the final round to prevent premature termination

Summary

  • WeKnora's ReAct loop resides in internal/agent, primarily implemented across engine.go, think.go, and act.go
  • AgentEngine.Execute initializes the loop, while executeLoop manages the iteration budget and control flow
  • Each iteration follows a strict Think→Analyze→Act→Observe cycle within runReActIteration
  • callLLMWithRetry provides resilient LLM communication with sanitization and retry logic
  • Tool execution supports both serial and parallel modes via executeToolCalls and runToolCall
  • Built-in safety mechanisms handle context overflow, empty responses, stuck loops, and truncated arguments
  • Full observability is provided through Langfuse spans for execution, rounds, and individual tool calls

Frequently Asked Questions

What is the maximum number of iterations in WeKnora's ReAct loop?

The maximum iterations are controlled by the withinIterationBudget method checking state.CurrentRound against a configured MaxIterations value, or unlimited if not specified. Additionally, the allowSteerOverrun flag permits one extra iteration after user steering messages to prevent premature termination.

How does WeKnora handle context window overflow during agent execution?

When a response hits context limits and the overflowRecovered flag is false, the engine triggers forceCompaction to compress the conversation history, sets overflowRecovered to true, and retries the LLM call exactly once. If the retry also overflows, the error propagates to the caller.

Can tool calls execute in parallel in WeKnora's agent?

Yes, the executeToolCalls function in act.go supports parallel execution when the ParallelToolCalls configuration flag is enabled. Each tool call runs through runToolCall with its own Langfuse span and timeout constraints via toolExecutionTimeout.

What happens when the LLM returns empty content repeatedly?

The engine tracks empty responses using the emptyRetries counter, allowing up to maxEmptyResponseRetries attempts while injecting prompts requesting the model provide its complete answer. For non-empty but identical content, consecutiveSameContent tracks repetitions against lastResponseContent, aborting the loop after maxRepeatedResponseRounds to prevent infinite loops.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →