How the TinyAgents Harness Drives Checkpointed Graph Execution for Agent Turns in OpenHuman

The TinyAgents harness orchestrates agent turns as deterministic directed-acyclic graphs (DAGs), persisting execution state via SqliteCheckpointer and journaling events through tinyagents_harness::observability, enabling precise resumption after crashes, user approvals, or sub-agent delegation.

In the OpenHuman repository, the tinyagents_harness crate serves as the core execution engine that transforms large language model (LLM) outputs into resilient, multi-step workflows. By implementing checkpointed graph execution, the harness ensures that every tool invocation, sub-agent call, and model interaction within an agent turn is tracked, persisted, and recoverable from the exact point of interruption.

Architectural Overview

The harness architecture separates concerns between runtime orchestration, durable persistence, and observability. This separation allows the system to freeze and resume complex agent graphs without losing execution context or workspace state.

The Runtime Engine

The primary entry point resides in src/openhuman/agent/tinyagents/mod.rs, which exposes the AgentHarness::run method. This method constructs a RunConfig from the incoming request and generates a RunPolicy via run_policy_for defined in src/openhuman/agent/tinyagents/model.rs. The policy encodes hard constraints including max_wall_clock_ms, max_output_tokens, and unknown_tool_policy, ensuring that resource limits are enforced before the graph executes.

The Checkpointing Layer

Durability relies on the tinyflows_sqlite::checkpoint::SqliteCheckpointer, initialized in src/openhuman/flows/tinyflows/caps/ops.rs through the open_flow_checkpointer function. This component serializes the partial graph, current RunContext, and workspace-relative files to <workspace>/flows/checkpoints.db. When a turn pauses or crashes, the harness reloads this checkpoint to reconstruct the exact graph state.

Observability and Journaling

Execution transparency comes from tinyagents_harness::observability::JournalSink, implemented in src/openhuman/agent/tinyagents/replay/ops.rs. This sink captures every AgentEvent—including node starts, completions, and errors—and writes them to the session database. The journal provides an append-only log that supports deterministic replay for debugging and real-time UI progress reporting via src/openhuman/web_chat/progress_bridge.rs.

Step-by-Step Checkpointed Execution Flow

1. Policy Initialization and Graph Construction

When a turn begins, AgentHarness::run first builds a RunPolicy that governs the entire execution. The harness then parses the LLM's tool-call output into a tinyagents_graph DAG, where:

  • Nodes represent tool executions, model inferences, or sub-agent launches
  • Edges carry data dependencies between operations

2. Initial Checkpoint Creation

Before executing any node, the harness invokes open_flow_checkpointer to create a Checkpoint entry. This initial checkpoint persists the graph topology and starting RunContext, establishing a recovery baseline. If the process terminates immediately after this write, the turn can resume from the beginning without data loss.

3. Middleware Pipeline Execution

Each graph node traverses a middleware stack defined in src/openhuman/agent/tinyagents/middleware.rs. The pipeline includes:

  • MessageTrimMiddleware – Truncates context to enforce max_output_token budgets
  • ContextCompressionMiddleware – Summarizes historical messages when limits approach
  • ModelFallbackMiddleware – Retries failed inferences with fallback models per the RunPolicy

4. Event Journaling

Upon node completion, the harness emits an AgentEvent captured by the JournalSink. In src/openhuman/agent/tinyagents/replay/ops.rs, these events are written to the session store, creating a granular audit trail. The UI reads this journal to display real-time progress without blocking the execution thread.

5. Sub-Agent Delegation

When the graph invokes a sub-agent, src/openhuman/web_chat/progress_bridge.rs spawns a new run via subagent_runner::run_subagent. This child execution receives its own RunPolicy and writes checkpoints to .openhuman/subagent_checkpoints. The parent graph stores the child run_id, enabling hierarchical resume semantics where parent and child states are restored atomically.

6. Approval Gates and Pausing

Destructive operations trigger approval gates handled in src/openhuman/agent/tinyagents/host/budget_gate.rs. When a gate activates, the harness:

  1. Writes a checkpoint with status AwaitingUser
  2. Persists the current node index and input arguments
  3. Suspends execution and releases resources

The SqliteCheckpointer ensures the graph remains in a consistent state during the pause.

7. Resumption and Termination

On resume, the harness reloads the checkpoint from flows/checkpoints.db, replays the journal to reconstruct transient state, and continues execution from the saved node index. Upon reaching a terminal node (no further tool calls), the harness finalizes the checkpoint, flushes the final transcript to the session database, and returns the aggregated response to the caller.

Key Source Files and Responsibilities

Component Source File Responsibility
Harness Entry src/openhuman/agent/tinyagents/mod.rs AgentHarness::run and runtime initialization
Policy & Model src/openhuman/agent/tinyagents/model.rs RunPolicy construction via run_policy_for
Checkpoint Store src/openhuman/flows/tinyflows/caps/ops.rs open_flow_checkpointer and SqliteCheckpointer implementation
Event Journal src/openhuman/agent/tinyagents/replay/ops.rs JournalSink implementation for AgentEvent recording
Middleware src/openhuman/agent/tinyagents/middleware.rs MessageTrimMiddleware, ContextCompressionMiddleware, ModelFallbackMiddleware
Budget Control src/openhuman/agent/tinyagents/host/budget_gate.rs Enforcement of wall-clock and token limits
Sub-Agent Bridge src/openhuman/web_chat/progress_bridge.rs subagent_runner integration and checkpoint coordination
Context Management src/openhuman/agent/tinyagents/thread_context.rs Thread-local RunContext and workspace handling

Summary

  • The TinyAgents harness treats each turn as a DAG of operations, enabling parallel execution and dependency tracking
  • Checkpointing via SqliteCheckpointer in flows/tinyflows/caps/ops.rs provides durable, SQLite-backed state persistence
  • Journaling through replay/ops.rs creates an immutable event log for debugging and UI progress
  • Middleware components enforce resource budgets and reliability policies without modifying graph logic
  • Sub-agent runs maintain separate checkpoints linked to parent graphs, supporting complex hierarchical workflows
  • Approval gates leverage checkpoint writes to pause execution safely, resuming deterministically when the user approves

Frequently Asked Questions

How does the TinyAgents harness handle crashes during tool execution?

When a crash occurs mid-node, the harness relies on the checkpoint written at the start of that node. Upon restart, AgentHarness::run reloads the checkpoint from flows/checkpoints.db and replays the journal to reconstruct the exact pre-crash state, then retries the failed node or applies the ModelFallbackMiddleware retry policy as configured in the RunPolicy.

What database schema stores the checkpointed graph state?

The SqliteCheckpointer defined in src/openhuman/flows/tinyflows/caps/ops.rs persists data to <workspace>/flows/checkpoints.db. This SQLite database stores serialized graph structures, RunContext objects, and execution metadata, shared between the flows engine and the tinyagents harness for unified state management.

How are sub-agent checkpoints linked to parent agent turns?

When a sub-agent launches via progress_bridge.rs, it receives a unique run_id and writes checkpoints to .openhuman/subagent_checkpoints. The parent graph stores this run_id in its own checkpoint; during resumption, the parent reloads its state, then loads the child checkpoint by ID, ensuring both graphs resume in sync.

Can checkpointed execution resume across different process restarts?

Yes. Because checkpoints are written to durable SQLite storage in flows/checkpoints.db rather than process memory, a turn can survive complete application restarts. The harness initializes the SqliteCheckpointer on startup, reads the latest checkpoint for the session, and reconstructs the graph and journal state from disk, enabling robust long-running agent workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →