# What Is the Trust Boundary in the MCP-Shared Gatekeeper?

> Understand the trust boundary in the MCP-Shared Gatekeeper. Discover how untrusted tool descriptions become trusted policy decisions at a single point in cloudflare-os.

- Repository: [Cloudflare/cloudflare-os](https://github.com/cloudflare/cloudflare-os)
- Tags: deep-dive
- Published: 2026-09-04

---

**The trust boundary in the `mcp-shared` Gatekeeper is a single, well-defined point in [`packages/mcp-shared/src/tools.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/mcp-shared/src/tools.ts) where untrusted tool descriptions from MCP endpoints are converted into trusted policy decisions, ensuring no other code can access tool annotations directly.**

The `cloudflare/cloudflare-os` repository implements a security-critical Gatekeeper pattern in its `mcp-shared` package to isolate potentially malicious Model Context Protocol (MCP) server metadata from the Gadget runtime. Understanding this boundary is essential for deploying safe, auto-approvable AI tools.

## Where the Trust Boundary Is Defined

The entire trust boundary lives in one module: [[`packages/mcp-shared/src/tools.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/mcp-shared/src/tools.ts)](https://github.com/cloudflare/cloudflare-os/blob/main/packages/mcp-shared/src/tools.ts). 

The file opens with an explicit comment that establishes the contract:

```ts
// The trust boundary: what an MCP server says about its own tools becomes what a Gadget may do.
// Nothing outside this file reads a tool's `annotations`.

```

This means **no code outside [`tools.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/tools.ts) is permitted to read a tool's `annotations` property**. All policy decisions—whether a tool is read-only, destructive, or idempotent—must flow through the classification functions exported by this module. This architectural constraint prevents downstream components from accidentally trusting unverified server metadata.

## How Trust Levels Are Classified

The boundary enforces trust through the **`ServerTrust`** type, which defines two distinct classification levels:

```ts
/** How far an endpoint's self-description is trusted. */
export type ServerTrust = "vetted" | "byo";

```

- **`vetted`** – The deployment administrator has manually asserted that the endpoint's annotations are reliable. This classification enables **auto-approval** for actions marked as safe by the server.
- **`byo`** – "Bring-your-own" indicates a user-provided URL without administrative vetting. In this mode, only the `readOnlyHint` annotation is honored; all other safety claims are ignored for auto-approval purposes.

This distinction allows the system to apply stricter scrutiny to arbitrary endpoints while permitting seamless automation for trusted infrastructure.

## The Policy Decision Point

The function **`classifyTool`** serves as the sole gateway between untrusted server descriptions and trusted runtime policy:

```ts
export function classifyTool(tool: McpTool, trust: ServerTrust): ClassifiedTool { … }

```

This function consumes a raw `McpTool` object and the chosen `ServerTrust` level, then returns a `ClassifiedTool` containing:
- The execution mode (`read` vs `action`)
- Whether the operation is auto-approvable
- Which source made the classification (`classifiedBy`)

All downstream helpers—`toolInfo()`, `toolSummary()`, and `describeCall()`—derive their outputs from this classification object. Because `classifyTool` is the only function that inspects the raw annotations, it acts as an **impassable guard rail**: malicious endpoints cannot inject arbitrary policy logic or bypass safety checks by manipulating their self-description.

## Why the Boundary Matters

Isolating trust decisions in [`tools.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/tools.ts) provides three critical security guarantees:

1. **Injection Prevention** – A compromised MCP server cannot influence policy logic outside the classification function, as no other module reads annotation data.
2. **Dynamic Trust Adjustment** – Administrators can change a server's trust level from `byo` to `vetted` (or vice versa) without restarting the connection, because the system reads the `trust` value fresh on every classification call.
3. **Stable Catalog Fingerprints** – Only annotations that affect policy decisions (`readOnlyHint`, `destructiveHint`, `idempotentHint`) contribute to the `catalogRevision` hash. This ensures the frontend can detect meaningful endpoint updates while ignoring superficial changes.

## Working with the Trust Boundary

When integrating with the Gatekeeper, always import classification utilities from the trust boundary module:

```ts
import { classifyTool, toolInfo, toolSummary } from "./tools.js";
import type { McpTool } from "./client.js";

// Example: an endpoint description received from the server
const serverTool: McpTool = {
  name: "search",
  title: "Web Search",
  description: "Search the web via a third‑party API.",
  annotations: { readOnlyHint: false, destructiveHint: false, idempotentHint: true },
  inputSchema: { /* JSON schema ... */ },
};

// The deployment decides how much it trusts this endpoint
const trust: ServerTrust = "vetted";   // or "byo"

// Classify the tool – this is the trust‑boundary call
const classified = classifyTool(serverTool, trust);

// Use the classified result in the UI or for policy checks
console.log(toolInfo(classified));      // ↦ name, title, mode, etc.
console.log(toolSummary(classified));   // ↦ a minimal catalog entry

```

To generate a catalog fingerprint for detecting endpoint updates:

```ts
import { catalogRevision } from "./tools.js";

const fingerprint = await catalogRevision([serverTool]);
// `fingerprint` changes only when the tool name or its policy‑affecting
// annotations change, letting the frontend detect endpoint updates.

```

For rendering approval prompts, use `describeCall`, which also respects the boundary:

```ts
import { describeCall } from "./tools.js";

const prompt = describeCall({
  serverName: "Acme Search",
  endpoint: "https://search.example.com",
  tool: serverTool,
  toolArgs: { query: "cloudflare" },
  mode: classified.mode,
  classifiedBy: classified.classifiedBy,
});

console.log(prompt.title);       // "Acme Search: search"
console.log(prompt.description); // Markdown shown to the user before approval

```

## Summary

- The **trust boundary** is strictly confined to [`packages/mcp-shared/src/tools.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/mcp-shared/src/tools.ts) in the `cloudflare/cloudflare-os` repository.
- The **`ServerTrust`** type distinguishes between `vetted` (administrator-approved) and `byo` (unverified user-provided) endpoints.
- The **`classifyTool`** function is the sole entry point for converting untrusted `McpTool` annotations into trusted `ClassifiedTool` policy objects.
- Downstream code relies on classified results rather than raw annotations, preventing malicious servers from bypassing safety controls.
- The boundary supports dynamic trust changes and generates stable catalog fingerprints via **`catalogRevision`**.

## Frequently Asked Questions

### What happens if code outside tools.ts tries to read tool annotations?

Any attempt to access `tool.annotations` outside of [`packages/mcp-shared/src/tools.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/mcp-shared/src/tools.ts) violates the architectural contract established in the codebase. The comment header in [`tools.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/tools.ts) explicitly states that "Nothing outside this file reads a tool's `annotations`," and code review processes (documented in [`REVIEW.md`](https://github.com/cloudflare/cloudflare-os/blob/main/REVIEW.md)) enforce this boundary to prevent security regressions.

### Can I change the trust level of a running MCP connection without restarting?

Yes. The `ServerTrust` value is passed as a parameter to `classifyTool` and evaluated fresh on each classification call. This design allows administrators to upgrade an endpoint from `byo` to `vetted` (or downgrade it) dynamically, immediately affecting auto-approval behavior without requiring a connection restart or service redeployment.

### Which annotations actually affect the catalog fingerprint?

Only annotations that influence policy decisions—specifically **`readOnlyHint`**, **`destructiveHint`**, and **`idempotentHint`**—are included in the `catalogRevision` calculation. Superficial changes to descriptions, titles, or input schemas do not alter the fingerprint, allowing the frontend to distinguish between cosmetic updates and security-relevant endpoint changes.

### How does the "byo" trust level protect against malicious servers?

When a server is classified as `byo` (bring-your-own), the Gatekeeper assumes the endpoint is untrusted and ignores all safety annotations except `readOnlyHint`. This prevents a malicious server from claiming destructive operations are safe or idempotent, effectively disabling auto-approval for any action that could modify state. Only `vetted` servers can trigger auto-approval for non-read operations based on their self-described annotations.