# How the ReAct Loop Works in WeKnora's `internal/agent` Package

> Discover how the ReAct loop in WeKnora's internal agent package orchestrates LLM calls for iterative reasoning and action. Learn about its executeLoop and runReActIteration functions to understand its thought process and tool e...

- Repository: [Tencent/WeKnora](https://github.com/tencent/WeKnora)
- Tags: internals
- Published: 2026-09-13

---

**The ReAct loop in WeKnora's `internal/agent` package implements a classic Reason→Act pattern where `AgentEngine.Execute` orchestrates iterative LLM calls through `executeLoop`, with each round handled by `runReActIteration` to decide whether to continue thinking, execute tools, or return a final answer.**

WeKnora, Tencent's open-source agent framework, provides a production-ready implementation of the **ReAct** (Reasoning and Acting) agent pattern in its `internal/agent` directory. The implementation follows an iterative cycle where the agent thinks, acts, and observes results until reaching a conclusion or exhausting its iteration budget. The core logic resides in three primary files: [`engine.go`](https://github.com/Tencent/WeKnora/blob/main/engine.go) for orchestration, [`think.go`](https://github.com/Tencent/WeKnora/blob/main/think.go) for LLM communication, and [`act.go`](https://github.com/Tencent/WeKnora/blob/main/act.go) for tool execution.

## Architecture of the ReAct Implementation

The ReAct loop spans three critical files within `internal/agent`:

- **[`engine.go`](https://github.com/Tencent/WeKnora/blob/main/engine.go)**: Contains the orchestration logic including `AgentEngine.Execute`, `executeLoop`, and `runReActIteration`
- **[`think.go`](https://github.com/Tencent/WeKnora/blob/main/think.go)**: Implements resilient LLM invocation via `callLLMWithRetry`
- **[`act.go`](https://github.com/Tencent/WeKnora/blob/main/act.go)**: Handles tool execution through `executeToolCalls` and `runToolCall`

These components work together to provide a robust loop with built-in context management, retry logic, and comprehensive observability via Langfuse tracing.

## Entry Point: AgentEngine.Execute

The `AgentEngine.Execute` function in [`engine.go`](https://github.com/Tencent/WeKnora/blob/main/engine.go) serves as the primary entry point for agent execution. It accepts a `context.Context`, session identifiers, the user query, previous message history, and optional image URLs.

```go
func (e *AgentEngine) Execute(
    ctx context.Context,
    sessionID, messageID, query string,
    llmContext []chat.Message, imageURLs ...[]string,
) (*types.AgentState, error) {
    // …build system prompt, initialise messages, register tools…
    _, err := e.executeLoop(ctx, state, query, messages, tools, sessionID, messageID)
    // …final logging and return…
}

```

This method prepares the conversation context—including system prompts and any pinned mentions—initializes the `AgentState`, and guarantees resource cleanup via `defer e.toolRegistry.Cleanup(ctx)`. After setup, it delegates control to `executeLoop` to run the actual ReAct cycle.

## The Main Loop: executeLoop

`executeLoop` manages the iteration lifecycle using budget constraints and explicit control outcomes. It tracks safety variables including `emptyRetries` and `consecutiveSameContent` to prevent infinite loops.

```go
func (e *AgentEngine) executeLoop(
    ctx context.Context,
    state *types.AgentState,
    query string,
    messages []chat.Message,
    tools []chat.Tool,
    sessionID, messageID string,
) (*types.AgentState, error) {
    emptyRetries, consecutiveSameContent := 0, 0
    lastResponseContent := ""

    for e.withinIterationBudget(state.CurrentRound) || e.allowSteerOverrun {
        outcome, iterErr := e.runReActIteration(
            ctx, state, &messages, tools,
            sessionID, messageID, query,
            &emptyRetries, &consecutiveSameContent, &lastResponseContent,
        )
        if iterErr != nil { return state, iterErr }

        switch outcome {
        case iterOutcomeContinue:   continue
        case iterOutcomeBreak:      break
        case iterOutcomeNext:       state.CurrentRound++
        }
    }
    // Handle max iterations if incomplete...
    return state, nil
}

```

The loop continues while `withinIterationBudget` returns true—checking against the configured `MaxIterations`—or while `allowSteerOverrun` permits one extra round after a user steering message. Three sentinel outcomes control flow:

- **`iterOutcomeContinue`**: Retry the current round (typically for empty content)
- **`iterOutcomeBreak`**: Exit the loop with final answer or fatal error
- **`iterOutcomeNext`**: Proceed to the next ReAct round

## Inside a Single ReAct Iteration

`runReActIteration` implements the core Think→Analyze→Act→Observe cycle. Each invocation represents one complete pass through the ReAct pattern.

### Context Management and Steering

Each iteration begins with context window estimation and potential compaction via `manageContextWindow`. The function also drains any mid-run user messages through `drainSteerMessages`, allowing users to inject steering commands during execution.

### The Think Phase

The engine invokes the LLM through `callLLMWithRetry` from [`think.go`](https://github.com/Tencent/WeKnora/blob/main/think.go). If the response hits context limits and `overflowRecovered` is false, the engine triggers `forceCompaction` and retries once.

```go
resp, err := e.callLLMWithRetry(ctx, messagesPtr, tools, state, query,
                                state.CurrentRound, sessionID)
if err != nil { return iterOutcomeNext, err }
if resp == nil { return iterOutcomeBreak, nil }

```

### The Analyze Phase

`analyzeResponse` evaluates the LLM output to determine the next step. If the response contains no tool calls and signals completion (`verdict.isDone`), the loop prepares to break. The engine also detects stuck loops by comparing content against `lastResponseContent`, aborting after `maxRepeatedResponseRounds` identical non-tool responses.

### The Act Phase

When the LLM returns tool calls, `executeToolCalls` routes execution to `runToolCall` for each tool. The engine supports **parallel execution** based on the `ParallelToolCalls` configuration flag.

### The Observe Phase

Results from tool executions append back to the conversation history via `appendToolResults` and `appendToolImages`. The `AgentStep` struct records the thought, tool calls, and timestamps for the round history.

```go
state.RoundSteps = append(state.RoundSteps, step)
*messagesPtr = e.appendToolResults(*messagesPtr, step)
*messagesPtr = e.appendToolImages(ctx, *messagesPtr, step)

```

## Resilient LLM Communication with callLLMWithRetry

`callLLMWithRetry` in [`think.go`](https://github.com/Tencent/WeKnora/blob/main/think.go) provides the sole LLM interface for the engine. It sanitizes messages to remove secret data, retries transient network errors a configurable number of times, and returns `nil` only when the model explicitly finishes without a response. This centralized wrapper ensures every round has a well-structured request/response pair with consistent error handling.

## Tool Execution and Parallelization

`runToolCall` in [`act.go`](https://github.com/Tencent/WeKnora/blob/main/act.go) handles individual tool invocations through a structured pipeline:

1. Parses JSON arguments with repair logic for malformed JSON
2. Resolves temporary model handles to concrete MCP targets
3. Emits **tool-hint** events for UI progress feedback
4. Creates Langfuse spans named `agent.tool.<name>`
5. Executes via `e.toolRegistry.ExecuteTool` with `toolExecutionTimeout` constraints
6. Wraps results in `types.ToolCall` structures and emits **tool-result** events

The engine refuses truncated tool calls when finish reasons indicate token limits, marking them failed with `truncatedArgumentsError` to prevent execution of partial arguments.

## Observability and Tracing with Langfuse

Every ReAct operation generates hierarchical Langfuse spans for comprehensive tracing:

- **`agent.execute`**: Root span for the entire execution
- **`agent.round.<N>`**: Individual iteration spans capturing input prompts and token usage
- **`agent.tool.<name>`**: Tool execution spans recording duration, success flags, and error information

These spans attach to the parent execution context, providing a complete audit trail of the agent's reasoning and actions.

## Edge Cases and Safety Mechanisms

The implementation includes several safeguards for production reliability:

| Situation | Mitigation Strategy |
|-----------|---------------------|
| **Context overflow** | After first detection, `forceCompaction` compresses history with a single retry (`overflowRecovered` flag) |
| **Empty LLM replies** | Up to `maxEmptyResponseRetries`, the engine injects continuation prompts and retries (`emptyRetries` counter) |
| **Stuck loops** | Tracks `consecutiveSameContent` against `lastResponseContent`, aborting after `maxRepeatedResponseRounds` identical responses |
| **Truncated arguments** | Detects token-cap finish reasons and marks tool calls failed with `truncatedArgumentsError` |
| **Steer-overrun** | Allows one extra iteration (`allowSteerOverrun`) when users inject messages during the final round to prevent premature termination |

## Summary

- WeKnora's ReAct loop resides in `internal/agent`, primarily implemented across [`engine.go`](https://github.com/Tencent/WeKnora/blob/main/engine.go), [`think.go`](https://github.com/Tencent/WeKnora/blob/main/think.go), and [`act.go`](https://github.com/Tencent/WeKnora/blob/main/act.go)
- `AgentEngine.Execute` initializes the loop, while `executeLoop` manages the iteration budget and control flow
- Each iteration follows a strict Think→Analyze→Act→Observe cycle within `runReActIteration`
- `callLLMWithRetry` provides resilient LLM communication with sanitization and retry logic
- Tool execution supports both serial and parallel modes via `executeToolCalls` and `runToolCall`
- Built-in safety mechanisms handle context overflow, empty responses, stuck loops, and truncated arguments
- Full observability is provided through Langfuse spans for execution, rounds, and individual tool calls

## Frequently Asked Questions

### What is the maximum number of iterations in WeKnora's ReAct loop?

The maximum iterations are controlled by the `withinIterationBudget` method checking `state.CurrentRound` against a configured `MaxIterations` value, or unlimited if not specified. Additionally, the `allowSteerOverrun` flag permits one extra iteration after user steering messages to prevent premature termination.

### How does WeKnora handle context window overflow during agent execution?

When a response hits context limits and the `overflowRecovered` flag is false, the engine triggers `forceCompaction` to compress the conversation history, sets `overflowRecovered` to true, and retries the LLM call exactly once. If the retry also overflows, the error propagates to the caller.

### Can tool calls execute in parallel in WeKnora's agent?

Yes, the `executeToolCalls` function in [`act.go`](https://github.com/Tencent/WeKnora/blob/main/act.go) supports parallel execution when the `ParallelToolCalls` configuration flag is enabled. Each tool call runs through `runToolCall` with its own Langfuse span and timeout constraints via `toolExecutionTimeout`.

### What happens when the LLM returns empty content repeatedly?

The engine tracks empty responses using the `emptyRetries` counter, allowing up to `maxEmptyResponseRetries` attempts while injecting prompts requesting the model provide its complete answer. For non-empty but identical content, `consecutiveSameContent` tracks repetitions against `lastResponseContent`, aborting the loop after `maxRepeatedResponseRounds` to prevent infinite loops.