How cloud-mcp-health Detects and Fixes Stale OpenWork Cloud MCP Connections
The cloud-mcp-health module detects stale connections by comparing the OpenWork engine's live MCP status against persisted delivery state, and automatically repairs them by marking workspaces as "stale" to trigger re-registration.
The cloud-mcp-health component in the different-ai/openwork repository provides autonomous health monitoring for OpenWork Cloud MCP (Model Context Protocol) integrations. By continuously validating the connection between the OpenWork engine and cloud-hosted MCP servers, this system identifies configuration drift and expired tokens without manual intervention.
Detection Logic
The detection mechanism relies on cross-referencing the engine's real-time inspection data with the local delivery state store. When the OpenWork engine reports that a cloud MCP server is no longer present while the local state still indicates an active connection, the system flags the discrepancy as a stale entry.
Engine Inspection and State Validation
The health probe queries the engine's MCP status through engineInspection, specifically checking the cloudPresent boolean. This property indicates whether the engine currently recognizes the cloud MCP server as active and reachable.
According to the source code in apps/server/src/cloud-mcp-health.ts, the detection logic evaluates two conditions simultaneously:
// apps/server/src/cloud-mcp-health.ts
if (engineInspection?.cloudPresent === false && delivery.state !== "stale") {
// If the engine says cloud is missing but we still have a delivery entry,
// mark it as stale to trigger re-registration.
cloudMcpDeliveryState.markWorkspaceStale(workspace, directory);
}
This check ensures that only non-stale entries trigger the marking process, preventing redundant state updates.
Identifying Configuration Drift
A stale connection typically manifests when the cloud MCP token expires, the server endpoint changes, or the engine's runtime configuration is reset. The delivery.state property tracks the last known status of the connection, with valid states including "ready", "registering", "failed", and "stale".
Marking Workspaces as Stale
Once detected, the system transitions the workspace entry into a stale state through the CloudMcpDeliveryStateStore class. This transition clears applied revisions and timestamps, effectively invalidating the cached connection metadata.
The markWorkspaceStale Implementation
The markWorkspaceStale method updates the in-memory state map for the specific workspace and directory combination:
// apps/server/src/cloud-mcp-health.ts
markWorkspaceStale(workspace: WorkspaceInfo, directory: string | null): void {
const entry = this.entries.get(this.key(workspace.id, directory));
if (!entry) return;
this.entries.set(this.key(workspace.id, directory), {
...entry,
state: "stale",
appliedRevision: undefined,
appliedAt: undefined,
updatedAt: Date.now(),
});
}
This method preserves the workspace identifier while clearing the appliedRevision and appliedAt fields, ensuring the next health check treats the connection as requiring full re-registration.
Automatic Repair Through Re-registration
After marking a connection as stale, the cloud-mcp-health system initiates a self-healing sequence on the subsequent health probe cycle. The stale state acts as a signal to bypass cached credentials and rebuild the MCP connection from scratch.
Re-registration Flow
When the health probe encounters a delivery entry with state: "stale", it invokes the registration routine through the runtime registrar. This process re-establishes the model-runtime entries, refreshes authentication tokens, and regenerates tool identifiers. The registrar coordinates with the CloudMcpDeliveryStateStore to track progress through distinct phases.
The state transitions follow this sequence:
- markRegistering: Indicates active re-registration attempt
- markReady: Confirms successful connection restoration
- markFailed: Captures permanent connection failures
State Transition Safety
The system prevents concurrent registration attempts by checking the current state before transitioning. Only entries explicitly marked as "stale" or "failed" trigger new registration flows, avoiding race conditions during the repair process.
Practical Implementation Examples
Detecting Stale Connections Manually
To programmatically check for stale connections in a specific workspace:
import { cloudMcpDeliveryState } from "@/cloud-mcp-health";
import { getWorkspaceInfo } from "@/utils";
const workspace = await getWorkspaceInfo("workspace-123");
const directory = null;
// Check if the connection requires repair
const snapshot = cloudMcpDeliveryState.snapshot(workspace, directory, null);
if (snapshot.state === "stale") {
console.log("Connection requires re-registration");
}
Triggering Stale State Manually
For testing or administrative purposes, you can force a workspace into stale state:
import { cloudMcpDeliveryState } from "@/cloud-mcp-health";
const workspace = { id: "workspace-123", name: "Production" };
cloudMcpDeliveryState.markWorkspaceStale(workspace, null);
This immediately invalidates the cached connection, forcing the next health probe to execute the full registration sequence.
Running the Health Probe
The server automatically executes the health probe, but you can invoke it directly:
import { runCloudMcpHealthProbe } from "@/cloud-mcp-health";
await runCloudMcpHealthProbe("workspace-123", { directory: null });
This function orchestrates the detection and repair logic, returning only after attempting to resolve any stale states.
Key Source Files and Architecture
The cloud-mcp-health system spans several critical files in the apps/server/src directory:
apps/server/src/cloud-mcp-health.ts: Core implementation containingCloudMcpDeliveryStateStore, the stale detection logic, and the registration orchestrationapps/server/src/runtime-opencode-config-store.ts: Manages the persisted OpenCode configuration including MCP endpoints and tokensapps/server/src/server-fetch.ts: Low-level HTTP client for direct MCP protocol operationsapps/server/src/mcp.ts: Utility functions for MCP tool management and error diagnostics
Summary
- Detection: The system compares
engineInspection.cloudPresentagainst the localdelivery.stateto identify when the engine no longer recognizes a cloud MCP server - Marking: The
markWorkspaceStalemethod transitions entries to the"stale"state, clearing revision metadata while preserving workspace identifiers - Repair: Stale states trigger automatic re-registration, which refreshes tokens and rebuilds runtime configurations
- Safety: Independent projection filters prevent stale entries from being used during model execution, ensuring only validated connections serve production traffic
Frequently Asked Questions
What causes an OpenWork Cloud MCP connection to become stale?
A connection becomes stale when the OpenWork engine loses recognition of the cloud MCP server due to token expiration, endpoint changes, or runtime restarts. The system detects this when engineInspection.cloudPresent returns false while a local delivery entry remains in "ready" state, indicating the local cache is out of sync with engine reality.
How does cloud-mcp-health prevent using stale connections during model execution?
Independent projection filters within the delivery state store ensure that only entries with state: "ready" are eligible for model-runtime mapping. When an entry is marked as "stale" in apps/server/src/cloud-mcp-health.ts, the projection layer excludes it from active use until re-registration completes and the state transitions back to "ready".
Can I manually trigger the stale detection process for testing purposes?
Yes. You can invoke cloudMcpDeliveryState.markWorkspaceStale(workspace, directory) directly to force a workspace into stale state. This is useful for testing recovery procedures or immediately invalidating compromised credentials without restarting the server, as the next health probe will automatically initiate re-registration.
What is the difference between "stale" and "failed" states in CloudMcpDeliveryStateStore?
The "stale" state indicates that the connection metadata is outdated but recoverable, triggering automatic re-registration on the next health check. The "failed" state indicates that re-registration attempts have exhausted retry limits or encountered permanent configuration errors, requiring manual intervention or configuration changes before the connection can be restored.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →