How to Debug Agent Loops and Prevent Infinite Spinning in 12-Factor Agents

To debug agent loops and prevent infinite spinning in 12-Factor Agents, implement step counters, validate LLM intent responses, and add structured logging around the while True loop in files like 03-agent.py and 07-agent.py.

The humanlayer/12-factor-agents repository demonstrates a canonical agent architecture where a persistent loop drives interaction with Large Language Models (LLMs). Because this pattern relies on the model explicitly emitting a terminal intent to break the cycle, misconfigurations or unexpected responses can cause the agent to spin indefinitely. Understanding how to instrument these loops ensures your agent either reaches a conclusion or fails loudly with actionable diagnostics.

Understanding the Agent Loop Architecture

At the core of every 12-Factor Agent is a loop implemented in files such as workshops/2025-07-16/walkthrough/03-agent.py and the more robust 07-agent.py. This loop follows a strict three-phase pattern:

  1. Serialize the thread – Convert the current conversation state to JSON or XML using thread.serialize_as_json() or thread.serialize_as_xml()
  2. Call the LLM client – Invoke baml_client.DetermineNextStep(thread_str) to receive the next action
  3. Interpret the intent – Check the result.intent attribute; if it equals "done_for_now", return the final message, otherwise append tool calls or clarifications and continue

The loop runs unconditionally via while True until the LLM explicitly signals completion. This design makes the agent vulnerable to infinite spinning when the model returns unrecognized intents or when the thread state stops evolving between iterations.

Common Causes of Infinite Spinning

According to the source code analysis, the typical infinite-spin scenario follows a specific failure chain:

  • The LLM returns an intent that the handler does not recognize (for example, "continue" instead of "done_for_now")
  • The elif chain falls through to a raise ValueError that is caught and silently suppressed by outer exception handlers
  • No new events are appended to the thread, causing the next DetermineNextStep call to receive identical context
  • The cycle repeats endlessly because the state never changes and the unhandled intent never triggers a hard failure

This pattern is particularly insidious in production because the agent appears active while making zero progress.

Debugging Techniques to Stop Infinite Spinning

Log Every Iteration

Visibility into the exact payload sent to and received from the LLM is the first line of defense. Before calling DetermineNextStep, print the serialized thread string. Immediately after the call, log the raw result object.

thread_str = thread.serialize_as_xml() if use_xml else thread.serialize_as_json()
print(f"📄 Serialized ({len(thread_str)} chars) →", thread_str[:200])

result = baml_client.DetermineNextStep(thread_str)
print("🤖 LLM result →", result)

This output reveals when the thread stops growing or when the model returns unexpected payloads.

Validate LLM Response Shapes

Guard against malformed responses that lack the required intent attribute. A missing attribute can trigger an AttributeError that aborts the loop or, worse, gets caught by a broad exception handler that masks the error.

if not hasattr(result, "intent"):
    raise ValueError(f"Malformed reply: {result}")

This validation ensures that any deviation from the expected schema results in an immediate, explicit failure rather than undefined behavior.

Implement a Safety Counter

Replace the unconditional while True with a bounded loop using a configurable max_steps parameter. This prevents runaway processes from consuming resources indefinitely.

def agent_loop(thread, clarification_handler, use_xml=True, max_steps=30):
    step = 0
    while step < max_steps:
        # ... existing logic ...

        step += 1
    raise RuntimeError("Agent loop exceeded safe iteration limit")

When the step counter reaches the limit, raise a RuntimeError with a clear message indicating that the agent failed to reach a terminal state.

Inspect Thread History

Persistent loops often stem from duplicated or missing thread events that confuse the model into thinking it has not finished. After each iteration, dump the complete thread state to a temporary file for offline analysis.

import pathlib
import json

dump_path = pathlib.Path("/tmp/thread_snapshot.json")
dump_path.write_text(json.dumps(thread.events, indent=2))

This allows you to diff consecutive states and identify where the conversation stagnates.

Force Breaks on Unexpected Intents

Ensure that any new intent introduced by a model update triggers a hard failure rather than silent continuation. Add an explicit else clause to your intent handling chain.

if result.intent == "done_for_now":
    return result.message
elif result.intent == "add":
    # handle tool call

    pass
else:
    raise NotImplementedError(f"Unhandled intent {result.intent}")

This pattern converts silent logic errors into immediately visible crashes.

Enable Verbose LLM Logging

If your BAML client supports it, enable debug mode during client initialization in walkthroughgen_py.py or wherever get_baml_client() is defined.

baml_client = get_baml_client()
baml_client.debug = True  # enables raw request/response logging

This exposes the raw JSON payloads transmitted between your application and the LLM provider, revealing truncation or formatting issues.

Production-Ready Debugging Implementation

Combine these techniques into a single hardened loop function. This implementation includes serialization logging, intent validation, a safety counter, and explicit error handling for unanticipated states.

def agent_loop(thread, clarification_handler, use_xml=True, max_steps=30):
    step = 0
    while step < max_steps:
        baml_client = get_baml_client()
        thread_str = thread.serialize_as_xml() if use_xml else thread.serialize_as_json()
        print(f"📄 Serialized ({len(thread_str)} chars) →", thread_str[:200])

        result = baml_client.DetermineNextStep(thread_str)
        print("🤖 LLM result →", result)

        if not hasattr(result, "intent"):
            raise ValueError(f"Malformed reply: {result}")

        if result.intent == "done_for_now":
            return result.message
        elif result.intent == "add":
            thread.add_event(result.tool_call)
        elif result.intent == "clarify":
            clarification_handler(result.question)
        else:
            raise NotImplementedError(f"Unhandled intent {result.intent}")

        step += 1
        
        # Snapshot for forensic analysis

        import json
        with open(f"/tmp/agent_step_{step}.json", "w") as f:
            json.dump(thread.events, f, indent=2)
            
    raise RuntimeError("Agent loop exceeded safe iteration limit")

This structure transforms opaque infinite spins into deterministic failures with clear tracebacks and preserved state snapshots.

Summary

  • The agent loop in 03-agent.py and 07-agent.py relies on while True and a terminal done_for_now intent; any gap in intent handling can cause infinite spinning.
  • Logging serialization and LLM responses before processing reveals when the thread stops evolving or the model returns garbage.
  • Validating result.intent with hasattr prevents AttributeError exceptions from being swallowed by outer catch blocks.
  • Implementing max_steps provides a hard ceiling on iterations, converting infinite loops into explicit RuntimeError exceptions.
  • Dumping thread history to disk enables post-mortem analysis of why the model refuses to terminate.
  • Explicit else: raise clauses ensure that new intents introduced by model updates fail fast rather than causing silent spins.

Frequently Asked Questions

How do I know if my 12-Factor Agent is stuck in an infinite loop?

If your agent process runs longer than expected while consuming steady CPU but producing no new output files or API calls to external tools, it is likely spinning. Check the logs for repeated identical thread serializations or the same LLM response appearing multiple times consecutively. If you see the same intent being processed repeatedly without new events being added to thread.events, the loop is stuck.

What is the maximum safe value for max_steps in the agent loop?

The appropriate value depends on your use case complexity, but the source examples suggest starting with 20 to 30 steps for typical workflows. Simple tasks in 03-agent.py might resolve in under 10 steps, while complex multi-tool workflows in 07-agent.py may require the full 30. Set the limit based on the longest reasonable successful execution trace you observe during testing, then add a 50% buffer.

Why does the LLM return an intent that causes infinite spinning instead of done_for_now?

The model may return an undefined intent like "continue" or "wait" if your system prompt is ambiguous or if the conversation history lacks clear termination signals. The thread serialization might also be truncated or malformed, causing the model to hallucinate intermediate states. Validate your XML/JSON serialization in thread.serialize_as_xml() and ensure your BAML schema for DetermineNextStep strictly enumerates allowed intents.

Where should I add logging to debug an agent loop effectively?

Add logging at three critical points in 07-agent.py: immediately before baml_client.DetermineNextStep() to capture the input thread, immediately after to capture the raw result, and inside the intent handling logic to confirm which branch executes. For forensic debugging, add a file dump of thread.events at the end of each loop iteration to /tmp/ or a similar volatile directory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →