How Apache Maka's Runtime Ledger Repair System Handles Event Log Corruption
Apache Maka's runtime ledger repair system detects corruption through diagnostic codes like incomplete_event, reconstructs missing terminal events by backfilling from stored conversation messages, validates recovered data against legacy turn state, and atomically appends missing events to restore ledger consistency.
Apache Maka maintains an append-only ledger of RuntimeEvents that records every interaction within a session. When this ledger becomes inconsistent—whether through missing terminal events, mismatched turn states, or unreconstructable steering messages—the runtime ledger repair subsystem rebuilds a coherent view of the session. Implemented primarily in packages/runtime/src/runtime-ledger-repair.ts, this system ensures downstream components can operate on reliable, consistent data.
Detecting Corruption in the Runtime Ledger
Identifying Corrupted Runs via Diagnostics
The repair process begins when AgentRunStore emits diagnostics containing a runId and the corruption code incomplete_event. The method firstRuntimeRepairRunId, located in runtime-ledger-repair.ts (lines 82-104), scans these diagnostics to locate the first un-repaired run matching known corruption patterns. This detection mechanism serves as the entry point for all subsequent repair workflows.
Reconstructing Missing Terminal Events
Loading and Validating Run State
The repairMissingTerminalFactOnce method initiates recovery by loading the run via runStore.readRun and forwarding it to repairRunTerminalFact. This validation layer confirms that the target run is actually in a terminal state before attempting reconstruction, preventing invalid repairs on active sessions.
Backfilling Events from Stored Messages
The core reconstruction logic resides in repairRunTerminalFactFromSnapshot (lines 202-226 in runtime-ledger-repair.ts). This method delegates to backfillRuntimeEventsFromStoredMessages (found in runtime-event-backfill.ts), which replays original conversation messages to generate a fresh list of RuntimeEvents. The system can utilize either full modelHistory or simplified conversation_text contexts to recreate the event stream.
Trustworthiness Verification
Before appending recovered data, the system invokes isTrustworthyRecoveredTerminal (lines 664-682). This function compares the reconstructed terminal against the legacy latestTurnState, verifying that statuses match and that failure or abort metadata align perfectly. Only trustworthy recoveries proceed to the persistence phase, preventing the propagation of inconsistent state.
Persisting Repairs to the Durable Ledger
Appending Missing Events
The missingRecoveredRuntimeEvents method (lines 363-394) walks the recovered event list and deduplicates entries against existing ledger records. It returns only truly missing events, which are then appended to the durable store via runtimeEventStore.appendRuntimeEvent.
Updating Run Headers and Metadata
Once terminal events are verified, repairRunHeaderFromExistingTerminal (lines 256-282) derives the correct run status, timestamps, and failure or abort metadata. It then calls commitTerminalRunWithRuntimeFact (defined in terminal-run-commit.ts) to persist the corrected run header, ensuring the run record reflects the repaired terminal state.
Synchronizing Turn State
To guarantee UI consistency, appendTerminalTurnStateIfNeeded (lines 398-409) writes a fresh turn_state message when the existing one is missing or out-of-sync. This step ensures that user interface components see a consistent view of the session's final state.
Handling Specialized Corruption Scenarios
Repairing Transcript Runs
For corrupted transcript runs—such as those imported from previous sessions—materializeTranscriptLedger (lines 85-133) derives turn records from ledger-only messages and creates a synthetic transcriptRunHeader. It then materializes the run via materializeTranscriptRun, restoring any missing runtime events and turn states for historic data.
Recovering Steering Messages
The repairSteeringMessagesOnce method (lines 136-155) iterates over all RuntimeEvents belonging to inline runs, extracts missing steering user messages via steeringMessageFromRuntimeEvent, and appends them to the message store. This ensures that user steering inputs remain accessible for future analysis or session continuation.
Concurrency Control and Repair Queueing
All repair operations are serialized per sessionId:runId through the withRepairQueue method (lines 223-243). This queueing mechanism prevents race conditions when multiple repair triggers occur concurrently, ensuring that backfill operations and ledger appends remain atomic and consistent even under parallel execution.
Practical Implementation Example
The following TypeScript example demonstrates how client components interact with the RuntimeLedgerRepair class:
import { RuntimeLedgerRepair } from '@maka/runtime';
import { createRuntimeLedgerRepairDeps } from './runtime-deps'; // custom factory for the required stores
// 1️⃣ Initialise the repair helper
const deps = createRuntimeLedgerRepairDeps(); // provides runStore, runtimeEventStore, etc.
const ledgerRepair = new RuntimeLedgerRepair(deps);
// 2️⃣ Repair a specific run that is known to be corrupted
async function fixRun(sessionId: string, runId: string) {
const repaired = await ledgerRepair.repairMissingTerminalFactOnce(sessionId, runId);
console.log(`Run ${runId} repaired: ${repaired}`);
}
// 3️⃣ Repair the whole transcript (e.g. after importing a historic session)
async function fixTranscript(header) {
await ledgerRepair.materializeTranscriptLedger(header);
}
// 4️⃣ Repair steering messages that may be missing
async function fixSteering(sessionId: string) {
const count = await ledgerRepair.repairSteeringMessagesOnce(sessionId);
console.log(`Recovered ${count} steering messages`);
}
The helper automatically serialises work per session/run and updates both the durable event store and the in-memory message store.
Summary
- Detection: The system identifies corruption via
firstRuntimeRepairRunIdscanning forincomplete_eventdiagnostics. - Reconstruction:
repairRunTerminalFactFromSnapshotbackfills events usingbackfillRuntimeEventsFromStoredMessagesto replay conversation history. - Validation:
isTrustworthyRecoveredTerminalverifies recovered terminals against legacy turn state before acceptance. - Persistence: Missing events are deduplicated and appended via
runtimeEventStore.appendRuntimeEvent, followed by header updates throughcommitTerminalRunWithRuntimeFact. - Specialized Repairs: Transcript runs and steering messages are handled by
materializeTranscriptLedgerandrepairSteeringMessagesOncerespectively. - Safety:
withRepairQueueensures atomic, race-condition-free repairs persessionId:runId.
Frequently Asked Questions
What triggers the runtime ledger repair system in Apache Maka?
The system triggers when diagnostics indicate specific corruption codes—particularly incomplete_event—or when the ledger contains missing terminal facts, mismatched turn states, or unreconstructable steering messages. The firstRuntimeRepairRunId method continuously monitors for these conditions.
How does Apache Maka validate that a repaired terminal event is accurate?
The isTrustworthyRecoveredTerminal function validates recovered terminals against the legacy latestTurnState, ensuring that run statuses match exactly and that failure or abort metadata align. Only events passing this verification are appended to the durable ledger.
Can the repair system handle concurrent repair attempts on the same run?
Yes, the withRepairQueue method serializes all repair work per sessionId:runId, preventing race conditions and ensuring atomic updates to the ledger even when multiple repair processes trigger simultaneously.
What happens to steering messages during ledger repair?
The repairSteeringMessagesOnce method scans all RuntimeEvents for inline runs, extracts any missing steering user messages via steeringMessageFromRuntimeEvent, and appends them to the message store. This ensures complete conversation history preservation even when original steering records are corrupted.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →