# How to Test and Evaluate LLM Agent Behavior Reliably: A 12-Factor Agents Guide

> Learn to reliably test and evaluate LLM agent behavior using the 12-Factor Agents framework. This guide details a reproducible notebook workflow for isolated prompt execution and tool calls in sandboxed environments.

- Repository: [HumanLayer/12-factor-agents](https://github.com/humanlayer/12-factor-agents)
- Tags: how-to-guide
- Published: 2026-05-19

---

**The 12-Factor Agents framework enables reliable LLM agent testing through a reproducible notebook-based workflow that isolates prompt execution, tool calls, and control flow in sandboxed environments.**

Testing LLM agents presents unique challenges because non-deterministic language models interact with external tools, APIs, and complex control flows. The `humanlayer/12-factor-agents` repository solves this by implementing a deterministic, four-step testing loop that validates both the **LLM prompt** (Factor 2) and the **control-flow loop** (Factor 8). This approach uses Jupyter notebooks as executable test artifacts, ensuring that agent behavior remains consistent and debuggable across commits.

## The 4-Step Testing Loop

The framework follows a strict **generate → execute → analyse → report** cycle. This separation of concerns isolates the agent’s reasoning from the execution environment and makes failures easy to detect and reproduce.

### 1. Generate a Test Notebook

Start by defining your test scenario as a minimal JSON representation of a Jupyter notebook. The repository provides [`workshops/2025-07-16/walkthroughgen_py.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/walkthroughgen_py.py), which converts high-level YAML scenarios into runnable notebooks. This ensures test cases are declarative and version-controlled.

The generator creates cells that define the agent’s task—such as prompting the LLM, invoking tool calls, or handling error conditions. By externalizing test definitions into YAML, you can add new regression tests without modifying Python code.

### 2. Execute in a Clean Sandbox

The shell script [`workshops/2025-07-16/test_notebook_colab_sim.sh`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/test_notebook_colab_sim.sh) creates a fresh Python virtual environment for each test run, guaranteeing identical dependencies. It installs Jupyter, copies the generated notebook, and executes it using `nbconvert`’s `ExecutePreprocessor` in a Google-Colab-style sandbox.

This script prints a clear **PASS/FAIL** banner and preserves the temporary directory (`tmp/test_…`) on disk for manual inspection. By running in a clean venv, the framework eliminates "works on my machine" inconsistencies that plague LLM agent development.

### 3. Analyze Executed Outputs

After execution, the notebook contains cell outputs including structured JSON responses from the LLM and BAML-generated logs. The utility [`workshops/2025-07-16/hack/inspect_notebook.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/hack/inspect_notebook.py) traverses the notebook structure and prints concise summaries of each code cell.

The analyzer highlights patterns such as `BAML`, `Parsed`, or error traces. You can filter outputs by specific keywords to verify that the agent emitted the expected tool call or handled errors according to Factor 9 (*Compact Errors*).

### 4. Report Results for CI/CD

A Bash wrapper ([`test_log_capture.sh`](https://github.com/humanlayer/12-factor-agents/blob/main/test_log_capture.sh) in the same directory) combines execution and analysis, exiting with non-zero status if expected patterns are missing. This makes the pipeline CI-friendly: failing tests abort the build, while passing tests produce deterministic provenance via saved notebooks and captured logs.

## Architectural Benefits of Notebook-Based Testing

This approach provides specific advantages for LLM agent evaluation:

- **Deterministic environment**: Fresh virtual environments (`python -m venv`) per test guarantee identical dependencies and eliminate state leakage between runs.
- **Full-stack visibility**: Executed notebooks save all outputs, including raw LLM responses. The [`inspect_notebook.py`](https://github.com/humanlayer/12-factor-agents/blob/main/inspect_notebook.py) utility filters by keyword or error type to surface exactly what the agent did.
- **Fast feedback loops**: The Bash wrapper exits early on first failure. CI systems can parallelize multiple YAML scenarios to test different agent configurations simultaneously.
- **Extensible scenarios**: New test cases require only new YAML files; the same [`walkthroughgen_py.py`](https://github.com/humanlayer/12-factor-agents/blob/main/walkthroughgen_py.py) runner consumes them without code changes.
- **Self-healing validation**: When the LLM returns an error, the notebook can contain follow-up correction steps. The analyzer captures both attempts for review, supporting robust error handling validation.

## Practical Implementation Examples

Below is a minimal "hello-world" test that verifies whether the LLM emits a JSON tool call:

```bash

# Create a simple notebook (JSON file)

cat > hello_test.ipynb <<'EOF'
{
  "cells": [
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "# Prompt the LLM to add two numbers\n",

        "print('Ask LLM: 2 + 2')\n",
        "result = client.DetermineNextStep('What is 2+2?')\n",
        "print('LLM response:', result)\n"
      ]
    }
  ],
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 4
}
EOF

# Run the test in the sandbox

./test_notebook_colab_sim.sh hello_test.ipynb

# Inspect outputs (looking for the expected JSON tool call)

python3 inspect_notebook.py ./tmp/test_$(date +%Y%m%d_%H%M%S)/test_notebook.ipynb "DetermineNextStep"

```

For BAML log capture validation, generate the notebook from a YAML scenario and verify structured output parsing:

```bash

# Generate notebook from YAML scenario

uv run python walkthroughgen_py.py simple_log_test.yaml -o test_capture.ipynb

# Execute and analyze

./test_notebook_colab_sim.sh test_capture.ipynb
python3 inspect_notebook.py ./tmp/test_$(date +%Y%m%d_%H%M%S)/test_notebook.ipynb "run_with_baml_logs"

```

If the log pattern `---Parsed Response (class DoneForNow)---` is present, the script exits with status `0`; otherwise it returns `1`, making it compatible with GitHub Actions, GitLab CI, and other pipeline tools.

## Key Source Files

The testing framework relies on these specific components in the `humanlayer/12-factor-agents` repository:

- [`workshops/2025-07-16/testing.md`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/testing.md): High-level documentation describing the testing philosophy and integration with 12-Factor principles.
- [`workshops/2025-07-16/test_notebook_colab_sim.sh`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/test_notebook_colab_sim.sh): Shell script that creates isolated Python environments and executes notebooks via `nbconvert`.
- [`workshops/2025-07-16/hack/inspect_notebook.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/hack/inspect_notebook.py): Python utility for post-execution analysis and log pattern matching.
- [`workshops/2025-07-16/walkthroughgen_py.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/walkthroughgen_py.py): Converts YAML test scenarios into executable Jupyter notebooks.
- [`content/factor-02-own-your-prompts.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-02-own-your-prompts.md): Defines the requirement to version and test LLM prompts as code.
- [`content/factor-08-own-your-control-flow.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-08-own-your-control-flow.md): Specifies deterministic control flow requirements that the test loop validates.
- [`content/factor-10-small-focused-agents.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md): Explains why keeping agents small reduces context window overflow and test flakiness.

## Summary

- **Use the generate-execute-analyze-report loop** to isolate LLM reasoning from execution environments and create reproducible test cases.
- **Leverage [`test_notebook_colab_sim.sh`](https://github.com/humanlayer/12-factor-agents/blob/main/test_notebook_colab_sim.sh)** to run tests in clean virtual environments, ensuring deterministic behavior across different machines.
- **Analyze outputs with [`inspect_notebook.py`](https://github.com/humanlayer/12-factor-agents/blob/main/inspect_notebook.py)** to verify structured JSON tool calls, BAML parsing, and error handling patterns.
- **Store tests as YAML scenarios** consumed by [`walkthroughgen_py.py`](https://github.com/humanlayer/12-factor-agents/blob/main/walkthroughgen_py.py) to maintain declarative, version-controlled test definitions.
- **Validate both prompts (Factor 2) and control flow (Factor 8)** to ensure the entire agent stack behaves reliably, not just the underlying language model.

## Frequently Asked Questions

### What makes notebook-based testing suitable for LLM agents?

Notebooks capture the complete execution context—including LLM responses, tool outputs, and intermediate state—in a single JSON file. This provides immutable provenance of exactly what the agent did, making it possible to debug non-deterministic behavior by inspecting the saved `tmp/test_…` directory after a failed run.

### How does the 12-Factor Agents framework prevent flaky tests?

The framework enforces **small, focused agents** (Factor 10) that stay within LLM context windows, reducing variability from truncated inputs. It also uses fresh virtual environments for every test execution via [`test_notebook_colab_sim.sh`](https://github.com/humanlayer/12-factor-agents/blob/main/test_notebook_colab_sim.sh), eliminating dependency drift that causes inconsistent behavior.

### Can this testing approach integrate with existing CI/CD pipelines?

Yes. The [`test_log_capture.sh`](https://github.com/humanlayer/12-factor-agents/blob/main/test_log_capture.sh) wrapper exits with status `0` for passing tests and `1` for failures, conforming to standard Unix conventions. Because tests run in isolated sandboxes without requiring external Jupyter servers, they execute reliably in GitHub Actions, GitLab CI, and local pre-commit hooks.

### What is BAML and why is it used in the test examples?

BAML (BoundaryML) is a framework for building typed LLM parsers. The test examples look for BAML log patterns like `---Parsed Response (class DoneForNow)---` to verify that the agent correctly structured its output according to predefined schemas. This validates that natural language inputs produce machine-readable tool calls, which is the core requirement outlined in [`content/factor-01-natural-language-to-tool-calls.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-01-natural-language-to-tool-calls.md).