How Openship Backup and Restore Works for Databases and Docker Volumes
Openship implements a complete, database-aware backup and restore subsystem using four tightly-coupled tables and a finite-state-machine orchestrator that streams PostgreSQL dumps and Docker volume archives to configurable destinations while maintaining durable execution state.
Openship's infrastructure treats every backup job as a durable row in the backup_run table, enabling safe recovery from crashes and supporting multiple storage backends including S3, SFTP, and local filesystems. The entire lifecycle—from scheduling policies to atomic volume restoration—is implemented in the oblien/openship repository through a combination of database schemas, orchestrators, and volume management utilities.
Core Data Model
The backup subsystem persists configuration and execution state across four primary tables defined in packages/db/src/schema/backup.ts:
-
backup_destination – Stores user-scoped storage endpoints (S3, SFTP, or local filesystem) with fields including
organizationId,name, andpathPrefix. -
backup_policy – Defines per-project schedules and triggers, containing
projectId, optionalserviceId,destinationId,enabledstatus,cronExpression, andpayloadKind. -
backup_run – Tracks individual backup executions with a status field that transitions through states (queued → preparing → snapshotting → uploading → verifying → succeeded/failed), plus
bytesTransferredandmanifestKey. -
backup_restore – Records restore operations with states (queued → preparing → prepared → applying → terminal) and links to the original
backup_runviarunId.
Backup Flow and State Machine
The BackupOrchestrator in apps/api/src/modules/backups/backup.orchestrator.ts drives each backup_run through a deterministic finite-state machine, persisting transitions via repos.backupRun.transition() in packages/db/src/repos/backup.repo.ts.
The flow proceeds through these phases:
-
Enqueue – An API endpoint creates a
backup_runrow in the queued state and hands the ID to the job runner (BullMQ or in-process fallback). -
Preparing – The orchestrator loads the policy, destination, and service handle, then verifies the destination using
destination.preflight(). -
Snapshotting – An optional pre-hook command executes inside the target container to prepare consistent state.
-
Uploading – The orchestrator iterates the producer adapters (
packages/adapters/src/backup/*), yielding artifacts:- Database dump – Streams
pg_dumpoutput for PostgreSQL services. - Volume archive – The
volume-transferadapter creates a tar archive of Docker volume contents.
Each artifact pipes through a
HashingPassthroughthat simultaneously streams to the destination and computes a SHA-256 checksum viauploadArtifact. - Database dump – Streams
-
Verifying – After all artifacts upload, the orchestrator writes a manifest JSON file containing metadata, artifact lists, and environment variable keys.
-
Post-hook – An optional command executes; failures are logged but do not abort the run.
-
Final transition – The row marks as succeeded (or failed on error), and the
notification-dispatcher.tsemits completion events.
If a crash occurs mid-execution, the run remains in its last in-flight status. The sweepStaleRuns function detects these on next boot and transitions them to server_error.
Restore Flow and Volume Swapping
Restores operate as the logical inverse of backups, processed by the RestoreOrchestrator using the same repository patterns in packages/db/src/repos/backup.repo.ts (specifically createBackupRestoreRepo, transition, and listInFlightByProject).
The restore state machine progresses as follows:
-
Queued → Preparing – Locate the original
backup_run, verify the destination, and download the manifest. -
Prepared – Extract artifacts into a staging area (Docker named volume or cloud workspace sub-path).
-
Applying – The orchestrator stops the target service, swaps the staged volume into the service's mount point, and restarts the container.
After the applying phase completes, the service runs with restored volume contents and database state.
Docker Volume Namespace Isolation
Openship isolates project volumes by prefixing named volumes with the project slug using the format openship-<slug>-<name>. This scoping prevents cross-project volume collisions and enables secure multi-tenancy.
The runtime helper volume-namespace in packages/adapters/src/runtime/volume-namespace.ts provides:
scopeVolumeBinds()– Rewrites raw Compose volume specifications into the namespaced Docker volume name.- Volume-transfer adapter – Streams the tar archive directly from the namespaced volume to the backup destination without intermediate disk writes.
Practical Implementation Examples
Triggering a Manual Backup
Use the backupOrchestrator to enqueue a backup outside the cron schedule:
import { backupOrchestrator } from "@repo/api/modules/backups/backup.orchestrator";
import { repos } from "@repo/db";
// Locate the policy for a PostgreSQL service
const policy = await repos.backupPolicy.findById("policy_123");
// Enqueue manual trigger
const { runId } = await backupOrchestrator.enqueue({
policyId: policy.id,
trigger: { source: "manual", userId: "user_42", clientIp: "10.0.0.5" },
});
console.log(`Backup queued, run ID = ${runId}`);
Inspect the resulting backup_run row via API or direct database query: SELECT * FROM backup_run WHERE id = $runId.
Restoring a Volume from Backup
To restore a specific backup run:
import { repos } from "@repo/db";
import { restoreOrchestrator } from "@repo/api/modules/backups/restore.orchestrator";
// Locate the backup run to restore
const run = await repos.backupRun.findById("bkr_abcdef");
// Create and enqueue restore record
const restore = await repos.backupRestore.create({
id: `rst_${crypto.randomUUID()}`,
runId: run.id,
organizationId: run.organizationId,
projectId: run.projectId,
status: "queued",
startedAt: new Date(),
});
// Execute the restore flow
await restoreOrchestrator.execute(restore.id);
Summary
- Openship manages backups through four tables (
backup_destination,backup_policy,backup_run,backup_restore) that track destinations, schedules, executions, and restores. - The BackupOrchestrator drives jobs through a durable finite-state machine with states from queued to succeeded, using
repos.backupRun.transition()for persistence. - Database dumps stream via
pg_dumpwhile volume archives transfer through thevolume-transferadapter, both passing through aHashingPassthroughfor SHA-256 verification. - Restore operations stage artifacts in temporary volumes before atomically swapping them into live services.
- Volume isolation uses the
openship-<slug>-<name>prefix viascopeVolumeBinds()inpackages/adapters/src/runtime/volume-namespace.ts. - Crashes are handled safely by
sweepStaleRuns(), which transitions orphaned in-flight runs toserver_erroron system boot.
Frequently Asked Questions
How does Openship handle backup failures during execution?
When an error occurs during any state, the orchestrator catches the exception and transitions the backup_run status to failed via repos.backupRun.transition(). A crash during execution leaves the run in its current in-flight state (e.g., snapshotting or uploading), which the sweepStaleRuns() function detects on the next orchestrator boot and moves to server_error. All state transitions emit events through the notification dispatcher for observability.
What storage backends does Openship support for backups?
According to the backup_destination schema in packages/db/src/schema/backup.ts, Openship supports S3-compatible storage, SFTP servers, and local filesystem paths. Each destination configuration includes organizationId scoping and an optional pathPrefix for organizing backup artifacts within the target storage.
How are Docker volumes isolated between different projects?
Openship enforces isolation by prefixing all named volumes with the project slug using the pattern openship-<slug>-<volume-name>. The scopeVolumeBinds() function in packages/adapters/src/runtime/volume-namespace.ts automatically rewrites raw Compose volume specifications to these namespaced identifiers, ensuring that Project A cannot access Project B's volume data during backup or restore operations.
Can I trigger backups manually, or only via scheduled cron expressions?
Both methods are supported. The backup_policy table accepts a cronExpression for automated scheduling, but the backupOrchestrator.enqueue() method allows manual triggering via the API by specifying trigger: { source: "manual", userId: "...", clientIp: "..." }. This creates a backup_run row that enters the same FSM workflow as scheduled jobs, enabling on-demand backups before deployments or maintenance windows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →