# How cloud-mcp-health Detects and Fixes Stale OpenWork Cloud MCP Connections

> Learn how cloud-mcp-health detects and fixes stale OpenWork Cloud MCP connections. It compares live MCP status to persisted delivery state and automatically re-registers stale workspaces.

- Repository: [Different AI/openwork](https://github.com/different-ai/openwork)
- Tags: how-to-guide
- Published: 2026-08-22

---

**The `cloud-mcp-health` module detects stale connections by comparing the OpenWork engine's live MCP status against persisted delivery state, and automatically repairs them by marking workspaces as "stale" to trigger re-registration.**

The `cloud-mcp-health` component in the `different-ai/openwork` repository provides autonomous health monitoring for OpenWork Cloud MCP (Model Context Protocol) integrations. By continuously validating the connection between the OpenWork engine and cloud-hosted MCP servers, this system identifies configuration drift and expired tokens without manual intervention.

## Detection Logic

The detection mechanism relies on cross-referencing the engine's real-time inspection data with the local delivery state store. When the OpenWork engine reports that a cloud MCP server is no longer present while the local state still indicates an active connection, the system flags the discrepancy as a stale entry.

### Engine Inspection and State Validation

The health probe queries the engine's MCP status through `engineInspection`, specifically checking the `cloudPresent` boolean. This property indicates whether the engine currently recognizes the cloud MCP server as active and reachable.

According to the source code in [`apps/server/src/cloud-mcp-health.ts`](https://github.com/different-ai/openwork/blob/main/apps/server/src/cloud-mcp-health.ts), the detection logic evaluates two conditions simultaneously:

```typescript
// apps/server/src/cloud-mcp-health.ts
if (engineInspection?.cloudPresent === false && delivery.state !== "stale") {
  // If the engine says cloud is missing but we still have a delivery entry,
  // mark it as stale to trigger re-registration.
  cloudMcpDeliveryState.markWorkspaceStale(workspace, directory);
}

```

This check ensures that only non-stale entries trigger the marking process, preventing redundant state updates.

### Identifying Configuration Drift

A stale connection typically manifests when the cloud MCP token expires, the server endpoint changes, or the engine's runtime configuration is reset. The `delivery.state` property tracks the last known status of the connection, with valid states including `"ready"`, `"registering"`, `"failed"`, and `"stale"`.

## Marking Workspaces as Stale

Once detected, the system transitions the workspace entry into a stale state through the `CloudMcpDeliveryStateStore` class. This transition clears applied revisions and timestamps, effectively invalidating the cached connection metadata.

### The markWorkspaceStale Implementation

The `markWorkspaceStale` method updates the in-memory state map for the specific workspace and directory combination:

```typescript
// apps/server/src/cloud-mcp-health.ts
markWorkspaceStale(workspace: WorkspaceInfo, directory: string | null): void {
  const entry = this.entries.get(this.key(workspace.id, directory));
  if (!entry) return;
  this.entries.set(this.key(workspace.id, directory), {
    ...entry,
    state: "stale",
    appliedRevision: undefined,
    appliedAt: undefined,
    updatedAt: Date.now(),
  });
}

```

This method preserves the workspace identifier while clearing the `appliedRevision` and `appliedAt` fields, ensuring the next health check treats the connection as requiring full re-registration.

## Automatic Repair Through Re-registration

After marking a connection as stale, the `cloud-mcp-health` system initiates a self-healing sequence on the subsequent health probe cycle. The stale state acts as a signal to bypass cached credentials and rebuild the MCP connection from scratch.

### Re-registration Flow

When the health probe encounters a delivery entry with `state: "stale"`, it invokes the registration routine through the runtime registrar. This process re-establishes the model-runtime entries, refreshes authentication tokens, and regenerates tool identifiers. The registrar coordinates with the `CloudMcpDeliveryStateStore` to track progress through distinct phases.

The state transitions follow this sequence:

- **markRegistering**: Indicates active re-registration attempt
- **markReady**: Confirms successful connection restoration  
- **markFailed**: Captures permanent connection failures

### State Transition Safety

The system prevents concurrent registration attempts by checking the current state before transitioning. Only entries explicitly marked as `"stale"` or `"failed"` trigger new registration flows, avoiding race conditions during the repair process.

## Practical Implementation Examples

### Detecting Stale Connections Manually

To programmatically check for stale connections in a specific workspace:

```typescript
import { cloudMcpDeliveryState } from "@/cloud-mcp-health";
import { getWorkspaceInfo } from "@/utils";

const workspace = await getWorkspaceInfo("workspace-123");
const directory = null;

// Check if the connection requires repair
const snapshot = cloudMcpDeliveryState.snapshot(workspace, directory, null);
if (snapshot.state === "stale") {
  console.log("Connection requires re-registration");
}

```

### Triggering Stale State Manually

For testing or administrative purposes, you can force a workspace into stale state:

```typescript
import { cloudMcpDeliveryState } from "@/cloud-mcp-health";

const workspace = { id: "workspace-123", name: "Production" };
cloudMcpDeliveryState.markWorkspaceStale(workspace, null);

```

This immediately invalidates the cached connection, forcing the next health probe to execute the full registration sequence.

### Running the Health Probe

The server automatically executes the health probe, but you can invoke it directly:

```typescript
import { runCloudMcpHealthProbe } from "@/cloud-mcp-health";

await runCloudMcpHealthProbe("workspace-123", { directory: null });

```

This function orchestrates the detection and repair logic, returning only after attempting to resolve any stale states.

## Key Source Files and Architecture

The `cloud-mcp-health` system spans several critical files in the `apps/server/src` directory:

- **[`apps/server/src/cloud-mcp-health.ts`](https://github.com/different-ai/openwork/blob/main/apps/server/src/cloud-mcp-health.ts)**: Core implementation containing `CloudMcpDeliveryStateStore`, the stale detection logic, and the registration orchestration
- **[`apps/server/src/runtime-opencode-config-store.ts`](https://github.com/different-ai/openwork/blob/main/apps/server/src/runtime-opencode-config-store.ts)**: Manages the persisted OpenCode configuration including MCP endpoints and tokens
- **[`apps/server/src/server-fetch.ts`](https://github.com/different-ai/openwork/blob/main/apps/server/src/server-fetch.ts)**: Low-level HTTP client for direct MCP protocol operations
- **[`apps/server/src/mcp.ts`](https://github.com/different-ai/openwork/blob/main/apps/server/src/mcp.ts)**: Utility functions for MCP tool management and error diagnostics

## Summary

- **Detection**: The system compares `engineInspection.cloudPresent` against the local `delivery.state` to identify when the engine no longer recognizes a cloud MCP server
- **Marking**: The `markWorkspaceStale` method transitions entries to the `"stale"` state, clearing revision metadata while preserving workspace identifiers
- **Repair**: Stale states trigger automatic re-registration, which refreshes tokens and rebuilds runtime configurations
- **Safety**: Independent projection filters prevent stale entries from being used during model execution, ensuring only validated connections serve production traffic

## Frequently Asked Questions

### What causes an OpenWork Cloud MCP connection to become stale?

A connection becomes stale when the OpenWork engine loses recognition of the cloud MCP server due to token expiration, endpoint changes, or runtime restarts. The system detects this when `engineInspection.cloudPresent` returns `false` while a local delivery entry remains in `"ready"` state, indicating the local cache is out of sync with engine reality.

### How does cloud-mcp-health prevent using stale connections during model execution?

Independent projection filters within the delivery state store ensure that only entries with `state: "ready"` are eligible for model-runtime mapping. When an entry is marked as `"stale"` in [`apps/server/src/cloud-mcp-health.ts`](https://github.com/different-ai/openwork/blob/main/apps/server/src/cloud-mcp-health.ts), the projection layer excludes it from active use until re-registration completes and the state transitions back to `"ready"`.

### Can I manually trigger the stale detection process for testing purposes?

Yes. You can invoke `cloudMcpDeliveryState.markWorkspaceStale(workspace, directory)` directly to force a workspace into stale state. This is useful for testing recovery procedures or immediately invalidating compromised credentials without restarting the server, as the next health probe will automatically initiate re-registration.

### What is the difference between "stale" and "failed" states in CloudMcpDeliveryStateStore?

The `"stale"` state indicates that the connection metadata is outdated but recoverable, triggering automatic re-registration on the next health check. The `"failed"` state indicates that re-registration attempts have exhausted retry limits or encountered permanent configuration errors, requiring manual intervention or configuration changes before the connection can be restored.