How Maka Handles Tool Execution Using the T1/T2 Protocol: A Complete Guide
Maka orchestrates durable tool execution through a two-phase checkpoint protocol called T1/T2 that persists pre-execution state to SQLite before running side-effects and commits results only after successful completion, enabling crash recovery without duplicate operations.
The Apache Maka project provides a runtime environment designed for reliable AI agent execution, where tools must produce consistent, recoverable side-effects even when processes crash unexpectedly. At the heart of this reliability mechanism lies the T1/T2 protocol, a transactional approach to tool invocation implemented across the runtime ledger and storage layers. This protocol ensures that every tool execution is bounded by durable checkpoints that guard against data loss and state corruption.
Understanding the T1/T2 Protocol Architecture
The T1/T2 protocol treats tool execution as a distributed transaction requiring two distinct persistence events. T1 captures the intent and state before any side-effects occur, while T2 finalizes the operation after the tool completes successfully. This bifurcation creates a recoverable boundary around unsafe operations, allowing the runtime to distinguish between "in-flight" and "completed" work across process restarts.
According to the architecture documentation in docs/architecture/runtime-resume-architecture.md, this design guarantees that the system can resume from exactly one of two valid states: either prior to T1 (tool not started) or after T2 (tool completed), never in between.
Phase 1 – The T1 Pre-Execution Checkpoint
Before invoking any tool, Maka writes a pre-execution checkpoint to the SQLite-backed ledger. This entry captures the execution context including the run-id, invocation-id, and turn-id, along with the serialized tool input and a snapshot of the model state. The checkpoint serves as an explicit promise that the system acknowledges responsibility for this operation.
In packages/runtime/src/tool-runtime.ts, the ToolRuntime class validates these identifiers before recording the start event:
// T1 – create pre‑execution checkpoint
if (!runId) throw new Error('Durable tool execution requires a run id');
if (!invocationId) throw new Error('Durable tool execution requires an invocation id');
await ledger.recordToolStart({ runId, invocationId, turnId, toolInput });
This write operation persists to packages/runtime/src/ai-sdk-backend.ts, where the storage layer manages the transactional boundary. Once T1 completes successfully, the runtime considers the operation "in-progress" and can recover it even if the process terminates immediately after.
Phase 2 – Tool Execution and Watchdog Management
With the T1 checkpoint secured, the runtime invokes the actual tool through a sandboxed execution environment. During this phase, the system pauses the stream watchdog—implemented in packages/runtime/src/builtin-tools.ts—to prevent timeout interruptions from corrupting long-running side-effects like database writes or API calls.
The tool executes with its inputs captured at T1, producing results that are held in memory but not yet exposed to the model or persisted to the ledger. This isolation ensures that partial failures do not contaminate the runtime state.
Phase 3 – The T2 Post-Execution Commit
Upon successful tool completion, Maka writes the T2 checkpoint, finalizing the side-effect and advancing the protocol version. This commit includes the tool's output and transitions the operation from "in-flight" to "completed." The T2 entry in packages/runtime/src/ai-sdk-backend.ts updates the protocol_version column to signal state advancement:
// T2 – commit post‑execution result
await ledger.recordToolFinish({
runId,
invocationId,
turnId,
toolOutput: toolResult,
});
// Underlying SQLite operation advancing protocol state
await db.run(
`INSERT INTO runtime_storage_root_binding(singleton, root_id, protocol_version)
VALUES (1, ?, ?)`,
[rootId, newProtocolVersion]
);
Only after T2 completes does the tool's effect become visible to subsequent model turns, guaranteeing atomic visibility of side-effects.
Crash Recovery and Durability Guarantees
If a crash occurs between T1 and T2, Maka's recovery logic—defined in docs/architecture/runtime-resume-phase0-crash-contract.md—replays the operation using the checkpointed inputs. Because tools are required to be idempotent, re-execution produces identical results without duplicating side-effects.
The SQLite ledger guarantees exactly-once semantics for T2 writes through unique constraint enforcement on the (runId, invocationId) tuple. As documented in the crash contract, the runtime identifies incomplete T1 entries during startup and re-executes them until they reach T2, ensuring no operation remains stranded in an intermediate state.
Isolation from Model Inference
The T1/T2 boundaries remain strictly separated from diagnostic operations. As implemented in packages/runtime/src/run-trace.ts, tracing, cost accounting, and logging occur in diagnostic-only sections that never perturb the durability boundary. This separation makes tool execution purely side-effecting from the model's perspective while maintaining observability for operators.
Summary
- T1 checkpoints capture execution context (run-id, invocation-id, turn-id) and tool inputs to SQLite before any side-effects occur, enabling recovery.
- T2 commits finalize tool outputs and advance the
protocol_version, making results visible only after durable persistence. - The ToolRuntime in
packages/runtime/src/tool-runtime.tsorchestrates the protocol throughrecordToolStart()andrecordToolFinish()calls. - Crash recovery relies on idempotent tool re-execution using T1 state, with guarantees defined in
runtime-resume-phase0-crash-contract.md. - Diagnostic tracing remains isolated from T1/T2 boundaries to prevent observability code from affecting durability semantics.
Frequently Asked Questions
What does T1/T2 stand for in Maka's protocol?
T1 and T2 represent the two transactional phases of tool execution: T1 is the pre-execution checkpoint that records intent and input state, while T2 is the post-execution commit that finalizes outputs. Though sometimes written as "TI/T2" in documentation, the nomenclature refers to these distinct persistence boundaries that bracket the actual tool invocation.
How does Maka ensure tool execution is idempotent?
The runtime enforces idempotency by design: if a crash occurs after T1 but before T2, the recovery logic retrieves the original inputs from the checkpoint and re-invokes the tool with identical parameters. Tools are implemented to produce the same side-effects when called with the same invocation-id, ensuring that duplicate T2 entries never create divergent state.
What happens if a crash occurs between T1 and T2?
The runtime detects incomplete operations during startup by querying the ledger for T1 entries lacking corresponding T2 commits. According to runtime-resume-phase0-crash-contract.md, the system re-executes these tools using the checkpointed context until they successfully reach T2. The SQLite storage layer guarantees that only one T2 entry can exist per invocation-id, preventing duplicate effects.
Where is the T1/T2 protocol implemented in the codebase?
The core implementation resides in packages/runtime/src/tool-runtime.ts, which manages the checkpoint lifecycle. Supporting components include packages/runtime/src/ai-sdk-backend.ts for SQLite persistence of protocol_version, packages/runtime/src/builtin-tools.ts for watchdog management during execution, and packages/runtime/src/run-trace.ts for non-invasive diagnostics. Architecture specifications live in docs/architecture/runtime-resume-architecture.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →