# How the TinyAgents Harness Drives Checkpointed Graph Execution for Agent Turns in OpenHuman

> Learn how the TinyAgents harness drives checkpointed graph execution for agent turns in OpenHuman. Discover seamless resumption and precise state management for robust agent behavior.

- Repository: [Tiny Humans/openhuman](https://github.com/tinyhumansai/openhuman)
- Tags: internals
- Published: 2026-09-01

---

**The TinyAgents harness orchestrates agent turns as deterministic directed-acyclic graphs (DAGs), persisting execution state via `SqliteCheckpointer` and journaling events through `tinyagents_harness::observability`, enabling precise resumption after crashes, user approvals, or sub-agent delegation.**

In the OpenHuman repository, the `tinyagents_harness` crate serves as the core execution engine that transforms large language model (LLM) outputs into resilient, multi-step workflows. By implementing **checkpointed graph execution**, the harness ensures that every tool invocation, sub-agent call, and model interaction within an agent turn is tracked, persisted, and recoverable from the exact point of interruption.

## Architectural Overview

The harness architecture separates concerns between runtime orchestration, durable persistence, and observability. This separation allows the system to freeze and resume complex agent graphs without losing execution context or workspace state.

### The Runtime Engine

The primary entry point resides in [`src/openhuman/agent/tinyagents/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/mod.rs), which exposes the `AgentHarness::run` method. This method constructs a `RunConfig` from the incoming request and generates a `RunPolicy` via `run_policy_for` defined in [`src/openhuman/agent/tinyagents/model.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/model.rs). The policy encodes hard constraints including `max_wall_clock_ms`, `max_output_tokens`, and `unknown_tool_policy`, ensuring that resource limits are enforced before the graph executes.

### The Checkpointing Layer

Durability relies on the `tinyflows_sqlite::checkpoint::SqliteCheckpointer`, initialized in [`src/openhuman/flows/tinyflows/caps/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/flows/tinyflows/caps/ops.rs) through the `open_flow_checkpointer` function. This component serializes the partial graph, current `RunContext`, and workspace-relative files to `<workspace>/flows/checkpoints.db`. When a turn pauses or crashes, the harness reloads this checkpoint to reconstruct the exact graph state.

### Observability and Journaling

Execution transparency comes from `tinyagents_harness::observability::JournalSink`, implemented in [`src/openhuman/agent/tinyagents/replay/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/replay/ops.rs). This sink captures every `AgentEvent`—including node starts, completions, and errors—and writes them to the session database. The journal provides an append-only log that supports deterministic replay for debugging and real-time UI progress reporting via [`src/openhuman/web_chat/progress_bridge.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/web_chat/progress_bridge.rs).

## Step-by-Step Checkpointed Execution Flow

### 1. Policy Initialization and Graph Construction

When a turn begins, `AgentHarness::run` first builds a `RunPolicy` that governs the entire execution. The harness then parses the LLM's tool-call output into a `tinyagents_graph` DAG, where:
- **Nodes** represent tool executions, model inferences, or sub-agent launches
- **Edges** carry data dependencies between operations

### 2. Initial Checkpoint Creation

Before executing any node, the harness invokes `open_flow_checkpointer` to create a `Checkpoint` entry. This initial checkpoint persists the graph topology and starting `RunContext`, establishing a recovery baseline. If the process terminates immediately after this write, the turn can resume from the beginning without data loss.

### 3. Middleware Pipeline Execution

Each graph node traverses a middleware stack defined in [`src/openhuman/agent/tinyagents/middleware.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/middleware.rs). The pipeline includes:
- **`MessageTrimMiddleware`** – Truncates context to enforce `max_output_token` budgets
- **`ContextCompressionMiddleware`** – Summarizes historical messages when limits approach
- **`ModelFallbackMiddleware`** – Retries failed inferences with fallback models per the `RunPolicy`

### 4. Event Journaling

Upon node completion, the harness emits an `AgentEvent` captured by the `JournalSink`. In [`src/openhuman/agent/tinyagents/replay/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/replay/ops.rs), these events are written to the session store, creating a granular audit trail. The UI reads this journal to display real-time progress without blocking the execution thread.

### 5. Sub-Agent Delegation

When the graph invokes a sub-agent, [`src/openhuman/web_chat/progress_bridge.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/web_chat/progress_bridge.rs) spawns a new run via `subagent_runner::run_subagent`. This child execution receives its own `RunPolicy` and writes checkpoints to `.openhuman/subagent_checkpoints`. The parent graph stores the child `run_id`, enabling hierarchical resume semantics where parent and child states are restored atomically.

### 6. Approval Gates and Pausing

Destructive operations trigger approval gates handled in [`src/openhuman/agent/tinyagents/host/budget_gate.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/host/budget_gate.rs). When a gate activates, the harness:
1. Writes a checkpoint with status `AwaitingUser`
2. Persists the current node index and input arguments
3. Suspends execution and releases resources

The `SqliteCheckpointer` ensures the graph remains in a consistent state during the pause.

### 7. Resumption and Termination

On resume, the harness reloads the checkpoint from `flows/checkpoints.db`, replays the journal to reconstruct transient state, and continues execution from the saved node index. Upon reaching a terminal node (no further tool calls), the harness finalizes the checkpoint, flushes the final transcript to the session database, and returns the aggregated response to the caller.

## Key Source Files and Responsibilities

| Component | Source File | Responsibility |
|-----------|-------------|----------------|
| **Harness Entry** | [`src/openhuman/agent/tinyagents/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/mod.rs) | `AgentHarness::run` and runtime initialization |
| **Policy & Model** | [`src/openhuman/agent/tinyagents/model.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/model.rs) | `RunPolicy` construction via `run_policy_for` |
| **Checkpoint Store** | [`src/openhuman/flows/tinyflows/caps/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/flows/tinyflows/caps/ops.rs) | `open_flow_checkpointer` and `SqliteCheckpointer` implementation |
| **Event Journal** | [`src/openhuman/agent/tinyagents/replay/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/replay/ops.rs) | `JournalSink` implementation for `AgentEvent` recording |
| **Middleware** | [`src/openhuman/agent/tinyagents/middleware.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/middleware.rs) | `MessageTrimMiddleware`, `ContextCompressionMiddleware`, `ModelFallbackMiddleware` |
| **Budget Control** | [`src/openhuman/agent/tinyagents/host/budget_gate.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/host/budget_gate.rs) | Enforcement of wall-clock and token limits |
| **Sub-Agent Bridge** | [`src/openhuman/web_chat/progress_bridge.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/web_chat/progress_bridge.rs) | `subagent_runner` integration and checkpoint coordination |
| **Context Management** | [`src/openhuman/agent/tinyagents/thread_context.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/agent/tinyagents/thread_context.rs) | Thread-local `RunContext` and workspace handling |

## Summary

- The TinyAgents harness treats each turn as a **DAG** of operations, enabling parallel execution and dependency tracking
- **Checkpointing** via `SqliteCheckpointer` in [`flows/tinyflows/caps/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/flows/tinyflows/caps/ops.rs) provides durable, SQLite-backed state persistence
- **Journaling** through [`replay/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/replay/ops.rs) creates an immutable event log for debugging and UI progress
- **Middleware** components enforce resource budgets and reliability policies without modifying graph logic
- **Sub-agent runs** maintain separate checkpoints linked to parent graphs, supporting complex hierarchical workflows
- **Approval gates** leverage checkpoint writes to pause execution safely, resuming deterministically when the user approves

## Frequently Asked Questions

### How does the TinyAgents harness handle crashes during tool execution?

When a crash occurs mid-node, the harness relies on the checkpoint written at the start of that node. Upon restart, `AgentHarness::run` reloads the checkpoint from `flows/checkpoints.db` and replays the journal to reconstruct the exact pre-crash state, then retries the failed node or applies the `ModelFallbackMiddleware` retry policy as configured in the `RunPolicy`.

### What database schema stores the checkpointed graph state?

The `SqliteCheckpointer` defined in [`src/openhuman/flows/tinyflows/caps/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/flows/tinyflows/caps/ops.rs) persists data to `<workspace>/flows/checkpoints.db`. This SQLite database stores serialized graph structures, `RunContext` objects, and execution metadata, shared between the flows engine and the tinyagents harness for unified state management.

### How are sub-agent checkpoints linked to parent agent turns?

When a sub-agent launches via [`progress_bridge.rs`](https://github.com/tinyhumansai/openhuman/blob/main/progress_bridge.rs), it receives a unique `run_id` and writes checkpoints to `.openhuman/subagent_checkpoints`. The parent graph stores this `run_id` in its own checkpoint; during resumption, the parent reloads its state, then loads the child checkpoint by ID, ensuring both graphs resume in sync.

### Can checkpointed execution resume across different process restarts?

Yes. Because checkpoints are written to durable SQLite storage in `flows/checkpoints.db` rather than process memory, a turn can survive complete application restarts. The harness initializes the `SqliteCheckpointer` on startup, reads the latest checkpoint for the session, and reconstructs the graph and journal state from disk, enabling robust long-running agent workflows.