How to Implement Human-in-the-Loop with LLM Tool Calls in 12-Factor Agents

Implement human-in-the-loop by exposing a structured request_human_input tool to the LLM, handling the resulting intent with a clarification handler like get_human_input(), and persisting the request-response pair in the thread history before resuming the agent loop.

The humanlayer/12-factor-agents repository provides a production-ready pattern for implementing human-in-the-loop with LLM tool calls through the Contact Humans with Tools factor. This architecture treats human interaction as a deterministic tool call, allowing the LLM to decide when it needs clarification while the agent loop handles pausing, state serialization, and seamless resumption.

The Three-Component Architecture

The implementation consists of three integrated pieces that work across any transport layer, from console applications to Slack webhooks.

Structured Tool Definition

The LLM emits a structured payload when it needs human input. According to content/factor-07-contact-humans-with-tools.md, this tool carries the question, optional context, and UI options such as urgency, format, and multiple-choice answers. The model decides when to ask, while the surrounding service handles how to deliver the question via console, Colab, email, or Slack.

Clarification Handler

The get_human_input() function in workshops/2025-07-16/walkthrough/05-main.py serves as the synchronous helper that presents the question to the user and returns the answer. This handler is injected into the agent loop, allowing the system to interrupt execution whenever the LLM emits a request_human_input intent.

Pause/Resume Agent Loop

The core implementation lives in workshops/2025-07-16/walkthrough/07-agent.py. The loop serializes the current thread (as JSON or token-efficient XML), calls the LLM via baml_client.DetermineNextStep(thread_str), and inspects the returned intent:

  • done_for_now – Return the final result and exit.
  • request_more_information – Invoke the clarification handler and append both the request and response to the thread.
  • Tool intents (e.g., add, multiply) – Execute the tool and append the outcome.

When the LLM returns an intent of request_human_input, the loop executes:

clarification = clarification_handler(result.message)   # pause

thread.events.append({"type": "clarification_request", "data": result.message})
thread.events.append({"type": "clarification_response", "data": clarification})

After recording the human response, the loop iterates again, sending the enriched context back to the LLM. Because the request and response are part of the thread history, the model retains full context of the decision.

Code Implementation Walkthrough

Below is a minimal end-to-end implementation that ties the pieces together from the repository's workshop files.

First, define the helper that interacts with the human operator:


# workshops/2025-07-16/walkthrough/05-main.py

def get_human_input(prompt: str) -> str:
    """Prompt the operator and return the answer."""
    print(f"\n🤔 {prompt}")
    # In Colab this uses the real input(); otherwise an auto-response is used for testing.

    return input("Your response: ") if IN_COLAB else "yes, proceed"

Next, implement the agent loop with the clarification handler:


# workshops/2025-07-16/walkthrough/07-agent.py

def run_agent(initial_message: str):
    thread = Thread([{"type": "user_input", "data": initial_message}])

    def clarification_handler(question: str) -> str:
        # Suspend the loop, ask the human, then resume.

        return get_human_input(f"The agent needs clarification: {question}")

    # The loop will automatically pause when the LLM emits 'request_human_input'.

    final_answer = agent_loop(thread, clarification_handler, use_xml=True)
    print(f"\n✅ Final answer: {final_answer}")

Trigger the workflow with a message that requires human approval:

run_agent("Deploy version 1.2.3 to production?")

The agent_loop function handles the serialization and state management automatically, performing these steps on each iteration:

  1. Serialize the thread to XML or JSON for token efficiency.
  2. Call baml_client.DetermineNextStep(thread_str).
  3. Route based on the returned intent.

Key Architectural Benefits

  • Clear separation of concerns – The LLM only decides when to ask; the surrounding service handles how to ask (UI, Slack, email).
  • Durable state – The request-response pair is persisted in the thread events, satisfying Factor 5 (unify execution state).
  • Outer-loop capability – Agents can be triggered by cron or webhook, pause for a human, then resume (Factor 6).
  • Scalable token usage – Serializing the thread as XML or YAML keeps the context compact while preserving structured data (Factors 3 and 4).

Summary

  • Implement human-in-the-loop with LLM tool calls by defining a structured request_human_input tool in the agent's schema.
  • Use a clarification handler like get_human_input() from workshops/2025-07-16/walkthrough/05-main.py to synchronously capture human responses.
  • Leverage the pause/resume logic in workshops/2025-07-16/walkthrough/07-agent.py to maintain conversation context across interruptions.
  • Persist all interactions as thread events to enable durable execution and multi-user coordination.
  • Support asynchronous transports like webhooks by recording the incoming response as a distinct response_from_human event.

Frequently Asked Questions

What is the human-in-the-loop pattern in 12-factor agents?

The human-in-the-loop pattern treats human interaction as a tool call that the LLM can invoke through a structured request_human_input intent. Rather than pre-defining human checkpoints in code, the LLM dynamically decides when it needs clarification, approval, or additional information based on the current context.

How does the agent loop handle pausing for human input?

The agent loop in workshops/2025-07-16/walkthrough/07-agent.py inspects the LLM's returned intent. When it detects request_more_information, it calls the injected clarification_handler function, which blocks execution until the human responds. The loop then appends both the request and response to the thread history and continues iteration, passing the full updated context back to the LLM.

Can this pattern work with asynchronous channels like Slack or email?

Yes. While the workshop implementation uses a synchronous console helper, the architecture supports asynchronous channels through webhook handlers. As documented in content/factor-07-contact-humans-with-tools.md, the serialized thread state can wait indefinitely for a webhook callback, which records the human response as a new thread event before resuming the agent loop.

How is state preserved during human-in-the-loop interruptions?

State preservation follows Factor 5 (unify execution state). The thread object containing all events—including the clarification_request and clarification_response—is serialized to JSON or XML before each LLM call. This allows the agent to be paused for seconds or days without losing context, enabling serverless deployments and durable execution across system restarts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →