Openship PID Tracking and Process Lifecycle Management: How PGlite Locks Prevent Data Corruption
Openship ensures exclusive access to PGlite data directories by writing the current process PID, host identity, and machine ID to an atomic lock file that self-heals after crashes and supports controlled takeovers during development.
Openship's process lifecycle management centers on a robust PID-based locking system that coordinates exclusive access to PGlite databases across the entire runtime. In the oblien/openship repository, the acquirePgliteLock function in packages/db/src/pglite-lock.ts serves as the gatekeeper, preventing concurrent processes from corrupting shared data directories. This mechanism extends beyond the database layer into Docker Compose namespaces, SSH adapters, and the runtime supervisor, creating a unified model for process isolation and resource cleanup.
How the PGlite Lock Manages PID Tracking
Lock File Creation with lockPathFor
When a process requests exclusive access, it calls acquirePgliteLock(dataDir, options). The helper lockPathFor computes a sibling path (<dataDir>.lock) and attempts an atomic open(..., "wx") operation. If successful, the function writes a JSON record containing process.pid, a timestamp, the hostname, and a stable machine identifier.
Machine Identity Validation
The machineId() function prefers platform-specific UUIDs—macOS IOPlatformUUID, Windows MachineGuid, or Linux /etc/machine-id. If these lookups fail, it falls back to the hostname and marks the result as unstable. This stability flag determines whether a lock can be safely evaluated across reboots or physical machine boundaries.
Reading and Validating Existing Locks
readLock parses the JSON record and validates required fields such as pid and host. If parsing fails, the lock is treated as unreadable and removed immediately. For legacy lock files that stored only a hostname, holderIsLegacyHostname detects the absence of a machine ID and falls back to PID-liveness checks rather than incorrectly flagging a cross-machine conflict.
Liveness Checks and Stale-Lock Reclamation
isProcessAlive(pid) probes the operating system with process.kill(pid, 0). A successful call means the holder exists; an ESRCH error means the PID is dead; EPERM indicates the process exists but belongs to another user. If acquirePgliteLock finds a lock whose PID is dead, it removes the stale file and retries acquisition.
Cross-Machine Safety
When a lock contains a stable machineId, the current process compares it against its own ID. If the values differ, acquisition aborts with a clear error because PGlite data directories cannot be safely shared across physical machines. This check protects against networked filesystem scenarios where two hosts might attempt concurrent access.
Development Hot-Reload Takeover
When running with --watch or OPENSHIP_DEV_LOCK_TAKEOVER=true, a new process may start while the previous one still holds the lock. In this mode, acquirePgliteLock can take over the existing lock: it sends SIGTERM to the old holder, waits for a grace period, follows with SIGKILL if necessary, and then loops to reacquire the lock.
Graceful Release via releasePgliteLock
On normal exit, the lock is released through releasePgliteLock. This function double-checks ownership by verifying both the PID and the machine ID before removing the file, ensuring that a takeover lock is not unintentionally deleted by a terminated predecessor.
Integration with Openship Subsystems
Docker Compose PID Namespaces
The same PID-namespace concept is exposed to Docker Compose through the pidMode field ("service:<name>"), allowing services to share a PID namespace for coordinated signaling. The ComposeNamespaceField definition in packages/core/src/compose-namespace.ts validates these values and enforces namespace constraints.
Runtime Supervisor and Remote Journal
The runtime supervisor in packages/adapters/src/runtime/supervisor/systemd.ts spawns background processes that rely on PID-aware commands like execReliable. Temporary socket paths in packages/adapters/src/system/system-ssh.ts embed the current PID—for example, /tmp/openship-ssh-${process.pid}-...—keeping resources scoped to the owning process and making cleanup deterministic.
System-SSH Resource Scoping
In packages/adapters/src/system/system-ssh-executor.ts, the executor creates PTY markers and forwarding sockets named with the current PID. This convention ensures that parallel executions do not clash and that orphaned resources can be mapped directly to their originating process. Deployment errors in packages/db/src/schema/deployment.ts also include offending PIDs for diagnostic clarity.
Practical Code Examples
Acquiring and Releasing the Database Lock
import { acquirePgliteLock, releasePgliteLock } from "packages/db/src/pglite-lock";
async function startDatabase(dataDir: string) {
// Wait for any previous holder to exit; allow takeover in dev
await acquirePgliteLock(dataDir, { takeover: true });
// ... open your PGlite instance here ...
}
process.once("SIGINT", () => {
releasePgliteLock();
process.exit(0);
});
Enabling Dev Mode Takeover
if (process.env.OPENSHIP_DEV_LOCK_TAKEOVER === "true" || process.execArgv.includes("--watch")) {
await acquirePgliteLock("/var/lib/openship/db", { takeover: true });
}
Diagnostic Lock Inspection
import { readFileSync } from "fs";
import { lockPathFor } from "packages/db/src/pglite-lock";
const lockPath = lockPathFor("/var/lib/openship/db");
const lockContents = readFileSync(lockPath, "utf8");
console.log("Current lock record:", JSON.parse(lockContents));
Summary
- Atomic lock files: Openship uses
open(..., "wx")inpackages/db/src/pglite-lock.tsto guarantee exclusive PGlite access via PID-tracked lock files. - Self-healing reclamation: Dead PIDs detected by
isProcessAlivetrigger automatic stale-lock removal, preventing corruption after crashes. - Machine-aware safety: Stable
machineIdcomparisons block cross-machine access attempts, while legacy hostname-only locks fall back to PID checks. - Dev hot-reload support: The takeover path sends
SIGTERMandSIGKILLto prior holders whenOPENSHIP_DEV_LOCK_TAKEOVERis enabled. - Subsystem integration: PID scoping extends to Docker Compose namespaces, SSH sockets, and runtime supervisors across the
oblien/openshipcodebase.
Frequently Asked Questions
What happens if the lock holder crashes without releasing the lock?
Openship detects the dead PID on the next acquisition attempt through isProcessAlive. If process.kill(pid, 0) returns ESRCH, the lock file is treated as stale and removed before the new process retries. This self-healing mechanism prevents permanent data directory lockouts after unexpected terminations.
Can two machines safely share the same PGlite data directory?
No. The acquirePgliteLock function compares the stable machineId stored in the lock against the local machine's identity. If the IDs differ and the stored ID is marked stable, acquisition aborts immediately. PGlite data directories are not designed for concurrent access across physical hosts.
How does Openship handle hot reloading during development?
When OPENSHIP_DEV_LOCK_TAKEOVER is set or Node.js runs with --watch, acquirePgliteLock enters takeover mode. It signals the previous holder with SIGTERM, grants a grace period, escalates to SIGKILL if needed, and then claims the lock. This allows rapid iteration without manual lock cleanup.
Where else does Openship use PID tracking outside of PGlite locks?
The runtime supervisor in packages/adapters/src/runtime/supervisor/systemd.ts relies on PID-aware execution via execReliable. Additionally, packages/adapters/src/system/system-ssh.ts and system-ssh-executor.ts embed process.pid into temporary socket and PTY paths, ensuring resources remain scoped to their owning process and can be cleaned up deterministically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →