Openship Process Group Cleanup: How It Eliminates Orphaned Processes

Openship prevents orphaned processes by using either a NohupSupervisor with setsid and kill -- -PID to terminate entire process groups, or a SystemdSupervisor that leverages systemd cgroups for automatic cleanup.

Openship is an open-source deployment platform that runs user services inside a process supervisor abstraction. To ensure robust Openship process group cleanup, the platform ships with two concrete supervisor implementations that guarantee no stray processes remain after a deployment stops, even if the host loses connectivity mid-operation.

Process Supervisors in Openship

The supervisor architecture is defined in packages/adapters/src/runtime/supervisor/types.ts, which exports the ProcessSupervisor interface. This abstraction allows the rest of the codebase to call stop(), destroy(), and other lifecycle methods without knowing the underlying platform-specific implementation.

Depending on the host environment, Openship automatically selects between two strategies:

NohupSupervisor for macOS and Generic Linux

The NohupSupervisor handles environments without systemd, including macOS and minimal containers. It spawns commands using nohup sh -lc … & and, on Linux, prefixes the launch with setsid to create a new process group whose leader PID equals the forked process ID.

When stop() is invoked in packages/adapters/src/runtime/supervisor/nohup.ts, the supervisor reads the saved PID and executes kill -- -${pid}. The double-dash and negative PID signal the kernel to target the entire process group (PGID), not just the leader. If the group does not terminate within 10 seconds, a forced kill -9 -- -${pid} is issued.

// NohupSupervisor.stop (excerpt from lines 65-73, 70-82)
const pid = await this.readPid(deploymentId);
if (!pid) return;

if (await this.isAlive(pid)) {
  // Kill entire process group (PGID = leader PID from setsid)
  await this.executor.exec(`kill -- -${pid} 2>/dev/null || kill ${pid} 2>/dev/null || true`);
  // Wait 10s, then SIGKILL if needed...
}

The setsid call occurs at launch time (lines 13-17), ensuring the service and all its descendants share a distinct process group that survives SSH disconnects. The leader PID is persisted to a .pid file in the deployment working directory, allowing reliable lookup during cleanup.

SystemdSupervisor for Systemd-Enabled Linux

On hosts running systemd, the SystemdSupervisor writes transient unit files (openship-<deployment>.service) and delegates cleanup to the init system. Systemd tracks the main process and its children via cgroups, automatically reaping all members when the unit stops.

The stop() method in packages/adapters/src/runtime/supervisor/systemd.ts (lines 85-90) simply executes systemctl stop openship-<deployment>.service. Tearing down the unit’s cgroup guarantees that no orphaned processes remain, regardless of how many child processes the service forked.

How Process Group Cleanup Works

Creating the Process Group with setsid

When NohupSupervisor launches a deployment on Linux, it uses setsid to create a new session and process group. This isolates the service from the parent shell’s process group, preventing SIGHUP propagation and ensuring the group remains intact even if the supervising SSH session drops.

Terminating the Entire Group

The cleanup logic relies on negative PID signaling. In Unix systems, kill -TERM -- -PGID delivers the signal to every process in the group. Openship implements this in NohupSupervisor.stop() with a graceful-to-forceful escalation:

  1. Send SIGTERM to the whole group via kill -- -${pid}
  2. Poll for 10 seconds checking process existence
  3. If processes persist, send SIGKILL via kill -9 -- -${pid}

For systemd-managed deployments, the cgroup destruction performed by systemctl stop accomplishes the same result without manual signal management.

Code Examples

Stopping a Deployment with NohupSupervisor

import { NohupSupervisor } from "@repo/adapters/runtime/supervisor/nohup";
import { exec } from "@repo/adapters/system/remote-journal";

const executor = /* CommandExecutor implementation */;
const workDir = "/var/openship/projects/my-app";
const supervisor = new NohupSupervisor(executor, workDir);

// Graceful shutdown (kills the whole process group)
await supervisor.stop("deployment-123");

// Force-kill and remove artifacts
await supervisor.destroy("deployment-123");

Stopping a Deployment with SystemdSupervisor

import { SystemdSupervisor } from "@repo/adapters/runtime/supervisor/systemd";

const executor = /* CommandExecutor implementation */;
const workDir = "/var/openship/projects/my-app";
const supervisor = new SystemdSupervisor(executor, workDir);

// Systemd stops the unit and destroys its cgroup
await supervisor.stop("deployment-456");

Key Implementation Files

Summary

  • Process group isolation is achieved via setsid (nohup) or systemd cgroups, ensuring services survive supervisor disconnects but remain controllable.
  • Cleanup completeness is guaranteed by targeting the entire PGID with kill -- -PID (nohup) or cgroup teardown (systemd).
  • Escalation logic waits 10 seconds before force-killing stubborn processes in the nohup implementation.
  • Abstraction layer in types.ts allows the rest of Openship to remain agnostic of the underlying supervisor strategy.

Frequently Asked Questions

How does Openship prevent zombie processes if the SSH connection drops?

Openship uses setsid in NohupSupervisor to create a new process group independent of the SSH session. Because the service runs in its own session, it does not receive SIGHUP when the parent terminal closes. Later, the supervisor can still locate the group via the saved PID file and terminate it with kill -- -PID, ensuring no orphaned processes remain.

What is the difference between NohupSupervisor and SystemdSupervisor cleanup?

NohupSupervisor manually manages process groups using Unix signals (kill -- -PID) and implements a 10-second graceful shutdown timeout before force-killing. SystemdSupervisor delegates cleanup to systemd, which tracks processes via cgroups; when the unit stops, systemd automatically reaps all processes in the cgroup without requiring manual signal logic.

Where does Openship store the process ID for later cleanup?

The NohupSupervisor writes the leader PID to a .pid file stored in the deployment’s working directory (specified during supervisor construction). The stop() method reads this file to determine the process group ID (PGID) to terminate. SystemdSupervisor does not require a PID file because it queries systemd directly via systemctl.

What happens if a process refuses to terminate during stop()?

In packages/adapters/src/runtime/supervisor/nohup.ts, the supervisor first sends SIGTERM to the entire process group. It then polls for up to 10 seconds. If processes remain alive after the timeout, it escalates to SIGKILL (kill -9) to guarantee termination. The SystemdSupervisor relies on systemd’s internal timeout and kill semantics, which provide similar escalation behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →