Recovery and Checkpoint Mechanisms in Reasonix for Rewinding Agent State
Reasonix implements a snapshot-based recovery system that captures pre-edit file states for every user turn, enabling transactional rewinds to any previous checkpoint through content-addressed storage and conflict-aware planning.
The DeepSeek-Reasonix agent framework provides deterministic recovery capabilities through a sophisticated checkpoint architecture. This system persists workspace snapshots at turn boundaries, allowing users to rewind both conversation history and file system state to arbitrary points. Understanding these recovery and checkpoint mechanisms reveals how Reasonix maintains data integrity during complex multi-turn agent workflows without relying on external version control.
Core Architecture: Store, Snapshots, and Rewind Plans
The recovery system rests on three fundamental abstractions implemented in internal/checkpoint/checkpoint.go.
Checkpoint Store
The checkpoint.Store type maintains an in-memory list of checkpoints and handles persistence to disk. Key methods include New for initialization, Begin to start a turn, List to enumerate available checkpoints, and TruncateFrom to remove future state after a rewind. The store serializes each checkpoint to a dedicated JSON file within a session-specific directory.
File Snapshots
checkpoint.FileSnap structures capture a file’s pre-edit state—including content, mode, SHA hash, and blob reference—the first time a path is touched during a turn. The store exposes CaptureBefore and CaptureBeforeFromChange to record these snapshots before any mutation occurs, ensuring the system can always revert to the exact prior state.
Rewind Transactions
The checkpoint.RewindPlan type prepares safe restore operations by analyzing coverage, detecting conflicts, and planning compensation steps. The store coordinates execution through PrepareRewind, CommitRewindWithForward, and RestoreCode, treating each rewind as an atomic transaction with forward-undo capability.
Checkpoint Lifecycle from Begin to Persist
Each user turn triggers a predictable sequence of checkpoint operations.
Starting a Turn with Store.Begin
When a new turn initiates, Store.Begin(turn, prompt, msgIndex) creates a fresh Checkpoint object and persists the previous one. According to the source in internal/checkpoint/checkpoint.go:23-33, this method establishes the boundary for all subsequent file operations.
Capturing Pre-Edit State
As writer tools such as fileedit or notebookedit execute, they invoke Store.CaptureBeforeFromChange or Store.CaptureBefore before modifying any file. The first touch of each path in the current turn is recorded as a FileSnap, as implemented in internal/checkpoint/checkpoint.go:63-84.
Recording Post-Mutation Fingerprints
After a mutation succeeds, Store.CaptureAfter stores the post-edit SHA, mode, and existence flag. This fingerprinting, found in internal/checkpoint/checkpoint.go:102-126, enables the system to track which checkpoint "owns" each file version.
Persistence and Garbage Collection
Every checkpoint serializes to <session>.ckpt/turn-N.json. A background garbage collector (gcLocked) expires older payloads and prunes unused blobs based on configurable retainN and blobQuota parameters, as shown in internal/checkpoint/checkpoint.go:66-97.
Rewind and Recovery Mechanics
Rewinding involves planning, validation, and transactional execution to ensure workspace consistency.
Preparing a Rewind Plan
Store.PrepareRewind(fromTurn, scope, ...) constructs a RewindPlan detailing which files can be restored, identifying coverage gaps, and flagging required confirmations. The controller calls this from internal/control/rewind.go:99-121 to initiate the recovery workflow.
Handling Coverage and Gaps
Checkpoints report coverage as complete, partial, or none (CoverageComplete, CoveragePartial, CoverageNone). Partial coverage arises from missing file payloads or recorded CoverageGap entries. The system blocks rewinds unless the planner confirms missing data is acceptable or the user explicitly approves the risk, as defined in internal/checkpoint/checkpoint.go:92-108.
Committing Transactional Restores
Store.CommitRewindWithForward(planID, ...) executes the restore transactionally. It writes files, deletes removed ones, and records a forward-undo transaction for later reversal. If any step fails, a compensation routine restores the original workspace state, implemented in internal/checkpoint/checkpoint.go:96-101.
Truncating Future State
After a successful rewind, Store.TruncateFrom(fromTurn) removes checkpoint files for all subsequent turns. This pruning, located in internal/checkpoint/checkpoint.go:120-133, ensures the agent cannot clash with stale snapshots when resuming operation.
Controller Integration and API Surface
The agent controller in internal/control/rewind.go orchestrates the recovery flow. It exposes PrepareRewind to validate boundaries and build plans, then coordinates with Store.RestoreCode—a high-level helper in internal/checkpoint/checkpoint.go:168-190 that validates legacy checkpoints, checks for conflicts, and commits the transaction.
The controller also surfaces checkpoint metadata via an HTTP endpoint. The /checkpoints handler in internal/serve/serve.go:1209-1215 returns Meta objects generated by Store.List, powering the UI picker for turn selection.
Practical Implementation Examples
Initializing a Checkpoint Store
// Initialise a checkpoint store for a workspace at "/tmp/workspace"
store := checkpoint.New("/tmp/workspace.ckpt", "/tmp/workspace")
// Begin turn 0 (first user turn)
store.Begin(0, "User asks for a file creation", 0)
// The file-editing tool will call CaptureBeforeFromChange before mutating.
store.CaptureBeforeFromChange(diff.Change{
Path: "hello.txt",
Kind: diff.Create,
OldText: "", // non-existent file
}, checkpoint.CaptureBeforeOpts{Source: checkpoint.CapturePreviewer})
// After the tool writes the file, it records the after fingerprint.
store.CaptureAfter("hello.txt", checkpoint.CaptureAfterOpts{Seq: 1})
Executing a Rewind Operation
// Suppose the user wants to go back to turn 0.
plan, err := store.PrepareRewind(0, checkpoint.RewindCode, 0, 0, false)
if err != nil {
log.Fatalf("cannot prepare rewind: %v", err)
}
// If the plan requires confirmation (partial coverage), handle it here…
if checkpoint.RewindPlanRequiresConfirmation(plan) {
// ask the user and obtain consent before proceeding
}
// Commit the rewind – this restores files and records a forward transaction.
result, err := store.CommitRewindWithForward(plan.PlanID, nil, nil, nil)
if err != nil {
log.Fatalf("rewind failed: %v", err)
}
fmt.Printf("rewound! written=%v deleted=%v\n", result.Written, result.Deleted)
Cleaning Up Future Turns After Rewind
// Remove all checkpoints from turn 1 onward (the agent rewound to turn 0)
if err := store.TruncateFrom(1); err != nil {
log.Fatalf("truncate failed: %v", err)
}
Listing Available Checkpoints
for _, meta := range store.List() {
fmt.Printf("Turn %d – %s (%d files, coverage=%s)\n",
meta.Turn, meta.Prompt, len(meta.Paths), meta.Coverage)
}
Summary
- Snapshot-based persistence: Reasonix captures
FileSnaprecords before any edit, storing content-addressed blobs and metadata incheckpoint.Store. - Turn-boundary lifecycle: Each turn begins with
Store.Begin, captures pre-edit state viaCaptureBeforeFromChange, and seals mutations withCaptureAfter. - Transactional rewinds: The
RewindPlansystem evaluates coverage gaps and conflicts beforeCommitRewindWithForwardexecutes atomic restores with compensation support. - Garbage collection: Background
gcLockedprunes expired checkpoints based on retention policies and blob quotas to manage disk usage. - Controller integration: The
internal/control/rewind.gocontroller coordinates planning and execution, whileinternal/serve/serve.goexposes checkpoint metadata to frontend pickers.
Frequently Asked Questions
What triggers a checkpoint creation in Reasonix?
A checkpoint is created at the start of every user turn when the controller calls Store.Begin. This initializes a new Checkpoint object in memory and persists the previous turn's state to disk as JSON. The system captures file snapshots lazily as tools touch paths during the turn, rather than snapshotting the entire workspace upfront.
How does Reasonix handle partial checkpoint coverage during rewinds?
The PrepareRewind method evaluates coverage levels (CoverageComplete, CoveragePartial, or CoverageNone). When coverage is partial due to missing file payloads or gaps, the RewindPlan flags the condition and requires explicit user confirmation before CommitRewindWithForward executes. This prevents accidental data loss from incomplete snapshot history.
Can a rewind operation be undone after it is committed?
Yes. CommitRewindWithForward records a forward-undo transaction during execution. If the user needs to reverse the rewind, the system can replay this transaction to restore the workspace to its pre-rewind state. If the commit itself fails midway, an automatic compensation routine reverts all partial changes to maintain consistency.
Where are checkpoint snapshots physically stored?
Checkpoints serialize to individual JSON files located at <session>.ckpt/turn-N.json within the workspace directory. File content (blobs) is stored separately in a content-addressed manner to deduplicate identical files across turns. The checkpoint.Store manages these paths based on the root and session parameters passed to checkpoint.New.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →