# How to Implement Human-in-the-Loop with LLM Tool Calls in 12-Factor Agents

> Learn how to implement human-in-the-loop with LLM tool calls. Expose request_human_input tool, handle intent with clarification handler, and persist history for seamless agent loops. Build smarter AI applications.

- Repository: [HumanLayer/12-factor-agents](https://github.com/humanlayer/12-factor-agents)
- Tags: how-to-guide
- Published: 2026-05-19

---

**Implement human-in-the-loop by exposing a structured `request_human_input` tool to the LLM, handling the resulting intent with a clarification handler like `get_human_input()`, and persisting the request-response pair in the thread history before resuming the agent loop.**

The humanlayer/12-factor-agents repository provides a production-ready pattern for implementing human-in-the-loop with LLM tool calls through the **Contact Humans with Tools** factor. This architecture treats human interaction as a deterministic tool call, allowing the LLM to decide when it needs clarification while the agent loop handles pausing, state serialization, and seamless resumption.

## The Three-Component Architecture

The implementation consists of three integrated pieces that work across any transport layer, from console applications to Slack webhooks.

### Structured Tool Definition

The LLM emits a structured payload when it needs human input. According to [`content/factor-07-contact-humans-with-tools.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-07-contact-humans-with-tools.md), this tool carries the question, optional context, and UI options such as urgency, format, and multiple-choice answers. The model decides *when* to ask, while the surrounding service handles *how* to deliver the question via console, Colab, email, or Slack.

### Clarification Handler

The `get_human_input()` function in [`workshops/2025-07-16/walkthrough/05-main.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/walkthrough/05-main.py) serves as the synchronous helper that presents the question to the user and returns the answer. This handler is injected into the agent loop, allowing the system to interrupt execution whenever the LLM emits a `request_human_input` intent.

### Pause/Resume Agent Loop

The core implementation lives in [`workshops/2025-07-16/walkthrough/07-agent.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/walkthrough/07-agent.py). The loop serializes the current thread (as JSON or token-efficient XML), calls the LLM via `baml_client.DetermineNextStep(thread_str)`, and inspects the returned intent:

- **`done_for_now`** – Return the final result and exit.
- **`request_more_information`** – Invoke the clarification handler and append both the request and response to the thread.
- **Tool intents** (e.g., `add`, `multiply`) – Execute the tool and append the outcome.

When the LLM returns an intent of `request_human_input`, the loop executes:

```python
clarification = clarification_handler(result.message)   # pause

thread.events.append({"type": "clarification_request", "data": result.message})
thread.events.append({"type": "clarification_response", "data": clarification})

```

After recording the human response, the loop iterates again, sending the enriched context back to the LLM. Because the request and response are part of the thread history, the model retains full context of the decision.

## Code Implementation Walkthrough

Below is a minimal end-to-end implementation that ties the pieces together from the repository's workshop files.

First, define the helper that interacts with the human operator:

```python

# workshops/2025-07-16/walkthrough/05-main.py

def get_human_input(prompt: str) -> str:
    """Prompt the operator and return the answer."""
    print(f"\n🤔 {prompt}")
    # In Colab this uses the real input(); otherwise an auto-response is used for testing.

    return input("Your response: ") if IN_COLAB else "yes, proceed"

```

Next, implement the agent loop with the clarification handler:

```python

# workshops/2025-07-16/walkthrough/07-agent.py

def run_agent(initial_message: str):
    thread = Thread([{"type": "user_input", "data": initial_message}])

    def clarification_handler(question: str) -> str:
        # Suspend the loop, ask the human, then resume.

        return get_human_input(f"The agent needs clarification: {question}")

    # The loop will automatically pause when the LLM emits 'request_human_input'.

    final_answer = agent_loop(thread, clarification_handler, use_xml=True)
    print(f"\n✅ Final answer: {final_answer}")

```

Trigger the workflow with a message that requires human approval:

```python
run_agent("Deploy version 1.2.3 to production?")

```

The `agent_loop` function handles the serialization and state management automatically, performing these steps on each iteration:

1. Serialize the thread to XML or JSON for token efficiency.
2. Call `baml_client.DetermineNextStep(thread_str)`.
3. Route based on the returned intent.

## Key Architectural Benefits

- **Clear separation of concerns** – The LLM only decides *when* to ask; the surrounding service handles *how* to ask (UI, Slack, email).
- **Durable state** – The request-response pair is persisted in the thread events, satisfying Factor 5 (unify execution state).
- **Outer-loop capability** – Agents can be triggered by cron or webhook, pause for a human, then resume (Factor 6).
- **Scalable token usage** – Serializing the thread as XML or YAML keeps the context compact while preserving structured data (Factors 3 and 4).

## Summary

- Implement human-in-the-loop with LLM tool calls by defining a structured `request_human_input` tool in the agent's schema.
- Use a clarification handler like `get_human_input()` from [`workshops/2025-07-16/walkthrough/05-main.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/walkthrough/05-main.py) to synchronously capture human responses.
- Leverage the pause/resume logic in [`workshops/2025-07-16/walkthrough/07-agent.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/walkthrough/07-agent.py) to maintain conversation context across interruptions.
- Persist all interactions as thread events to enable durable execution and multi-user coordination.
- Support asynchronous transports like webhooks by recording the incoming response as a distinct `response_from_human` event.

## Frequently Asked Questions

### What is the human-in-the-loop pattern in 12-factor agents?

The human-in-the-loop pattern treats human interaction as a tool call that the LLM can invoke through a structured `request_human_input` intent. Rather than pre-defining human checkpoints in code, the LLM dynamically decides when it needs clarification, approval, or additional information based on the current context.

### How does the agent loop handle pausing for human input?

The agent loop in [`workshops/2025-07-16/walkthrough/07-agent.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/walkthrough/07-agent.py) inspects the LLM's returned intent. When it detects `request_more_information`, it calls the injected `clarification_handler` function, which blocks execution until the human responds. The loop then appends both the request and response to the thread history and continues iteration, passing the full updated context back to the LLM.

### Can this pattern work with asynchronous channels like Slack or email?

Yes. While the workshop implementation uses a synchronous console helper, the architecture supports asynchronous channels through webhook handlers. As documented in [`content/factor-07-contact-humans-with-tools.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-07-contact-humans-with-tools.md), the serialized thread state can wait indefinitely for a webhook callback, which records the human response as a new thread event before resuming the agent loop.

### How is state preserved during human-in-the-loop interruptions?

State preservation follows Factor 5 (unify execution state). The thread object containing all events—including the `clarification_request` and `clarification_response`—is serialized to JSON or XML before each LLM call. This allows the agent to be paused for seconds or days without losing context, enabling serverless deployments and durable execution across system restarts.