# RLM Runtime Kernel Host Request Protocol: How the Python REPL Communicates with the TypeScript Host

> Discover the RLM runtime kernel host request protocol. Learn how the Python REPL uses JSON-over-pipes to communicate typed requests and responses with the TypeScript host over stdin/stdout.

- Repository: [Prime Intellect/prime-agent](https://github.com/PrimeIntellect-ai/prime-agent)
- Tags: internals
- Published: 2026-09-06

---

**The RLM runtime uses a bidirectional JSON-over-pipes protocol that lets the Python REPL kernel exchange typed requests and responses with the surrounding TypeScript host via newline-delimited messages over stdin/stdout.**

The **RLM (Recursive Large-Model) runtime** is the execution sandbox at the heart of Prime Intellect's agent framework. Its kernel host request protocol enables Python code running inside the REPL to delegate privileged operations—like spawning sub-agents or querying model availability—to the outer TypeScript host. This article breaks down the wire format, message flow, and implementation details based on the source code in `PrimeIntellect-ai/prime-agent`.

## Wire Format: Line-Based JSON Framing

The RLM runtime kernel host request protocol transmits **newline-delimited JSON objects**, UTF-8 encoded, with one complete message per line. This framing design eliminates length-prefix complexity while remaining trivial to parse.

In [`prime-agent-runtime/src/rlm/repl.py`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/prime-agent-runtime/src/rlm/repl.py) (lines 44-46), a **process-wide write lock** (`_write_lock`) serializes all writes to the event channel. This guarantees **atomic frame delivery** even when multiple threads emit events concurrently.

The runtime announces **protocol version 3** in its initial `ready` event, as specified in [`repl.md`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/repl.md) line 5. Host implementations should verify this version before accepting further traffic.

## Channel Architecture

| File Descriptor | Direction | Purpose |
|---------------|-----------|---------|
| **fd 0 (stdin)** | Host → REPL | Request ingestion: `execute`, `interrupt`, `host_reply`, `snapshot`, `restore`, `list_names`, `shutdown` |
| **fd 1 (stdout)** — duped | REPL → Host | **Event emission**: `ready`, `stdout`, `stderr`, `result`, `display`, `host_request`, `error`, `done` |
| **fd 2 (stderr)** | REPL → Host | Captured and forwarded as `stderr` events via a dedicated pump thread |

The asymmetric design separates **commands** (host-to-kernel) from **events** (kernel-to-host). This unidirectional flow per descriptor simplifies reasoning about backpressure and buffering.

## Request Types and Semantics

The RLM runtime kernel host request protocol supports seven request types. Only `interrupt` and `host_reply` are **fire-and-forget**; all other requests process **strictly in order, one at a time** per the *Requests* section of [`repl.md`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/repl.md).

| Type | Payload Structure | Reply Behavior |
|------|-------------------|--------------|
| `execute` | `{"type":"execute","id":str,"code":str}` | Emits `result`/`display`/`error` events, terminates with `done` |
| `interrupt` | `{"type":"interrupt","id"?:str}` | No reply; cancels matching request by ID |
| `host_reply` | `{"type":"host_reply","id":str,"data":{"status":"ok","result":…}}` | No reply; completes pending `host_request` |
| `snapshot` | `{"type":"snapshot","id":str,…}` | `done` with snapshot statistics |
| `restore` | `{"type":"restore","id":str,…}` | `done` with restore report |
| `list_names` | `{"type":"list_names","id":str}` | `done` containing `names` list |
| `shutdown` | `{"type":"shutdown","id"?:str}` | `done` then process exit |

## Host Request Flow: The Full Round-Trip

The **type-safe, asynchronous host request mechanism** lets Python code invoke host-side handlers and await results. Here's the complete flow as implemented in the Prime Agent source:

### Step 1: Python Initiates

In [`prime-agent-runtime/src/rlm/__init__.py`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/prime-agent-runtime/src/rlm/__init__.py) (lines 66-83), the `host_request` coroutine provides the public API:

```python
from rlm import host_request

# Ask the host to list available sub-agents

reply = await host_request("rlm.list_subagents")
print(reply)  # -> {"subagents": [...]}

```

### Step 2: Runtime Routes and Registers

The call forwards to `rlm.repl.host_request` in [`repl.py`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/repl.py) (lines 101-113):

1. Generate a **runtime-minted request ID**: `rid = uuid.uuid4().hex`
2. Register a `Future` in the global `_pending_host` map
3. Emit a `host_request` event: `_send({"event":"host_request","id":rid,"data":data})`
4. Suspend await on the `Future`

### Step 3: Host Handles and Replies

The TypeScript host receives the event via [`packages/coding-agent/src/core/kernel/repl-manager.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/core/kernel/repl-manager.ts), dispatches to a registered handler by type, and returns a `host_reply` request:

```ts
// TypeScript handler registration (simplified)
import { registerHostHandler } from "./hostBridge";

registerHostHandler("rlm.list_subagents", async (payload) => {
  const subagents = await session.listSubagents();
  return { status: "ok", result: { subagents } };
});

```

### Step 4: Runtime Resolves

The REPL's reader thread invokes `_resolve_host_reply` ([`repl.py`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/repl.py) lines 27-36), which:
- Matches the incoming ID to the pending `Future`
- Resolves with the raw reply dictionary
- Consumes the ID (single-use guarantee)

Finally, `_parse_host_reply` in [`__init__.py`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/__init__.py) (lines 57-63) validates the status envelope, raising `RuntimeError` on `"error"` or malformed responses.

### Complete Example: Spawning Child Agents

Python side invoking `rlm.run`:

```python
from rlm import host_request, run

# Spawn a child agent and receive its admission handle

handle = await run("Explain quantum entanglement.", model="gpt-4")
print(handle.rlm_child_id)  # Populated via host request

```

Corresponding TypeScript handler:

```ts
registerHostHandler("rlm.run", async ({prompt, kwargs}) => {
  const child = await spawnChildAgent(prompt, kwargs);
  return {
    status: "ok",
    result: {
      rlm_child_id: child.id,
      name: child.name,
      session_dir: child.dir,
      model: child.model
    }
  };
});

```

## Error Handling and Shutdown Guarantees

The RLM runtime kernel host request protocol includes robust failure modes:

- **Inactive runtime**: If `_loop is None` or `_host_closed` is true, `host_request` raises `RuntimeError` immediately ([`repl.py`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/repl.py) lines 103-106)
- **Graceful shutdown**: `_fail_pending_host_requests` (lines 17-25) marks every waiting future with an error, preventing cell hangs
- **Protocol violations**: Malformed lines emit an `error` event with `ename:"ProtocolError"`, then continue serving—no fatal crash

## Key Implementation Files

| Path | Responsibility |
|------|---------------|
| [`prime-agent-runtime/src/rlm/repl.py`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/prime-agent-runtime/src/rlm/repl.py) | Core REPL, event framing, `_write_lock`, `_pending_host` registry, request ID generation |
| [`prime-agent-runtime/src/rlm/repl.md`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/prime-agent-runtime/src/rlm/repl.md) | Authoritative protocol specification, version declaration, request/event schema |
| [`prime-agent-runtime/src/rlm/__init__.py`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/prime-agent-runtime/src/rlm/__init__.py) | Public Python API (`host_request`, `run`, `find_models`), reply parsing and validation |
| [`packages/coding-agent/src/core/kernel/repl-manager.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/core/kernel/repl-manager.ts) | TypeScript host bridge, handler registration, `host_request` event dispatch |

## Summary

The RLM runtime kernel host request protocol delivers:

- **Lightweight JSON-over-pipes transport** with atomic, line-delimited framing
- **Birectional request/response pattern** (`host_request` ↔ `host_reply`) layered on unidirectional stdio channels
- **Strict ordering and single-use IDs** ensuring deterministic, race-free execution
- **Typed host delegation** that keeps the Python kernel sandboxed while accessing privileged host capabilities

## Frequently Asked Questions

### What transport does the RLM runtime kernel host request protocol use?

The protocol runs over **anonymous pipes** (stdin/stdout/stderr) between the TypeScript host process and the Python REPL child process. All messages are **newline-delimited JSON** with UTF-8 encoding. This design avoids network stack complexity while maintaining clean process isolation.

### How are concurrent host requests prevented from interleaving?

Two mechanisms enforce ordering: **(1)** the process-wide `_write_lock` in [`repl.py`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/repl.py) guarantees atomic frame writes, and **(2)** the runtime processes non-fire-and-forget requests **strictly sequentially**—the reader thread won't dispatch a new request until the previous cycle completes. Each `host_request` ID is also **consumed exactly once**, eliminating replay or collision risks.

### What happens if the host crashes while a host_request is pending?

The Python side detects pipe closure via `_host_closed` and raises `RuntimeError` on any new `host_request` calls. For in-flight requests, the shutdown path through `_fail_pending_host_requests` (lines 17-25) preemptively resolves all pending futures with an error, preventing indefinite awaits and allowing exception handling in user code.