How Claude-Mem Health Monitoring Detects and Recovers From Worker Failures

Claude-Mem continuously monitors worker liveness through PID-file validation, HTTP health probes, and version compatibility checks, automatically spawning replacement processes when failures or mismatches are detected.

Claude-Mem, an open-source memory system for AI coding assistants maintained at thedotmack/claude-mem, relies on a long-lived worker process to host its HTTP API. To ensure high availability for the Claude-Code hook, Cursor integration, and Viewer UI, the system implements a resilient health monitoring subsystem that detects dead or outdated workers and recovers without manual intervention.

The Health Monitoring Architecture

The detection-and-recovery loop spans three coordinated modules in the src/services directory:

Detecting Worker Failures

Claude-Mem employs a three-layer detection strategy to identify unresponsive workers.

PID-File Sanity Checks

When ensureWorkerStarted() initiates, it invokes cleanStalePidFile() via the ProcessManager module. This function reads the stored PID from ~/.claude-mem/worker.pid and issues a non-blocking process.kill(pid, 0) signal. If the process no longer exists, the PID file is immediately removed, signaling that the worker has died.

// Conceptual flow from ProcessManager.ts
await cleanStalePidFile(); // Removes ~/.claude-mem/worker.pid if process is dead

Port Probes

The isPortInUse(port) function in HealthMonitor.ts performs an HTTP GET request to http://127.0.0.1:<port>/api/health. A successful 200 OK response confirms the worker is listening; any connection refusal or timeout indicates the port is free, suggesting the worker has crashed or shut down.

// From HealthMonitor.ts
const isHealthy = await isPortInUse(37666); // Returns true if /api/health responds

Health-Loop Verification

When a port appears occupied, waitForHealth(port, timeout) enters a polling loop, querying the health endpoint every 500 milliseconds until it receives a 200 response or the timeout expires. This handles race conditions where the process exists but the HTTP server hasn't finished initializing.

Detecting Version Mismatches

A responsive worker may still be incompatible after a Claude-Mem plugin upgrade. The checkVersionMatch(port) function addresses this by comparing the installed plugin version (read from the marketplace package.json via getInstalledPluginVersion()) against the running worker's version (fetched from GET http://127.0.0.1:<port>/api/version via getRunningWorkerVersion()).

If the versions differ, the system returns { matches: false }, triggering a controlled restart to ensure the worker matches the installed codebase.

Recovery Mechanisms

When detection layers identify a failure or mismatch, Claude-Mem executes a coordinated recovery sequence.

Graceful Shutdown and Port Release

Upon version mismatch detection, ensureWorkerStarted() calls httpShutdown(port) to request a graceful stop via the worker's /api/admin/shutdown endpoint, followed by waitForPortFree(port, timeout) to ensure the TCP socket is fully released before spawning a replacement.

Daemon Spawning

The spawnDaemon(__filename, port) function in ProcessManager.ts constructs platform-specific commands to launch a detached worker process. On Windows, it uses PowerShell; on Unix systems, it employs setsid or plain detached spawn to ensure the worker survives the parent's termination.

Windows Cooldown Protection

To prevent rapid-fire spawn attempts from creating multiple console windows on Windows, shouldSkipSpawnOnWindows() checks for a lock file .worker-start-attempted and enforces a 2-minute cooldown period between spawn attempts.

Post-Spawn Verification

After spawning, ensureWorkerStarted() re-enters the health verification loop with an extended timeout (HOOK_TIMEOUTS.POST_SPAWN_WAIT). If the new worker responds to health checks, the system clears the spawn lock file via clearWorkerSpawnAttempted() and returns success. If health checks fail, the PID file is removed and the function returns false, allowing the CLI to report an error and the Claude-Code plugin wrapper to retry on the next hook invocation.

End-to-End Health Check Flow

The complete detection and recovery sequence follows this simplified flow:

ensureWorkerStarted(port)
 ├─ cleanStalePidFile()                ← remove dead PID file
 ├─ isPortInUse(port) ?
 │   └─ yes → waitForHealth()
 │          ├─ health OK → checkVersionMatch()
 │          │   ├─ match → ready
 │          │   └─ mismatch → httpShutdown() → waitForPortFree()
 │          └─ health timeout → spawnDaemon()
 └─ no  → spawnDaemon()
        └─ waitForHealth() → success ? ready : error

Code Examples

Programmatic Health Check

Use the same function the CLI relies on to verify worker status:

import { ensureWorkerStarted } from './services/worker-service';
import { getWorkerPort } from './shared/worker-utils';

async function verifyWorkerHealth() {
  const port = getWorkerPort();
  const isHealthy = await ensureWorkerStarted(port);
  
  if (isHealthy) {
    console.log('✅ Worker is up-to-date and responsive');
  } else {
    console.error('❌ Worker failed health verification');
  }
}

Source: ensureWorkerStarted implementation in src/services/worker-service.ts.

Manual Health Endpoint Testing

Test the worker's health and version endpoints directly:


# Check liveness

curl -s http://127.0.0.1:37666/api/health

# Verify version compatibility

curl -s http://127.0.0.1:37666/api/version

Source: Health probe functions in src/services/infrastructure/HealthMonitor.ts.

Forcing a Version-Mismatch Recovery

Simulate the automatic restart logic when versions differ:

import { httpShutdown, waitForPortFree } from './services/infrastructure/HealthMonitor';
import { getWorkerPort } from './shared/worker-utils';

async function gracefulRestart() {
  const port = getWorkerPort();
  
  // Request graceful shutdown
  await httpShutdown(port);
  
  // Wait for port release
  await waitForPortFree(port, 10000);
  
  // Next call to ensureWorkerStarted() will spawn fresh instance
}

Source: Version-mismatch handling in src/services/worker-service.ts.

Summary

  • Three-layer detection: Claude-Mem validates worker liveness through PID-file verification, HTTP port probes, and iterative health-loop polling.
  • Version safety: The system compares the installed plugin version against the running worker's reported version via checkVersionMatch(), triggering restarts on mismatches.
  • Automatic recovery: The ensureWorkerStarted() orchestrator handles stale PID cleanup, graceful shutdowns, platform-specific daemon spawning, and post-spawn verification without user intervention.
  • Platform protection: Windows-specific cooldown logic prevents rapid spawn attempts from creating multiple console windows.

Frequently Asked Questions

How does Claude-Mem know if the worker process has crashed?

Claude-Mem checks the PID file stored at ~/.claude-mem/worker.pid using a non-blocking process.kill(pid, 0) signal. If the process no longer exists, cleanStalePidFile() removes the stale PID file, signaling that the worker has died and needs restart.

What happens if the worker is running but unresponsive to HTTP requests?

The isPortInUse() function attempts a GET request to /api/health. If the connection is refused or times out, the system treats the port as free. Subsequently, ensureWorkerStarted() will spawn a new daemon, and the old unresponsive process will be replaced once the new worker binds to the port.

How does the system handle plugin updates without manual restarts?

When ensureWorkerStarted() detects a running worker, it calls checkVersionMatch() to compare the installed plugin version (from package.json) against the worker's reported version via /api/version. If versions differ, the system sends an httpShutdown() request, waits for the port to free with waitForPortFree(), and spawns a new worker matching the updated codebase.

Why does Claude-Mem use a cooldown period on Windows?

To prevent rapid spawn attempts from creating multiple PowerShell console windows when the worker fails to start immediately, shouldSkipSpawnOnWindows() checks for a .worker-start-attempted lock file and enforces a 2-minute cooldown. This ensures the system doesn't overwhelm the user with console pop-ups during recovery attempts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →