# RLM Deployment Strategies: How to Deploy Recursive Language Models in Local and Cloud Environments

> Explore RLM deployment strategies from local in-process execution to cloud sandboxes like Modal, Prime, Docker, e2b, and Daytona. Deploy Recursive Language Models with flexibility.

- Repository: [az/rlm](https://github.com/alexzhang13/rlm)
- Tags: deployment-strategies
- Published: 2026-06-18

---

**RLM supports six deployment strategies ranging from in-process local execution to fully isolated cloud sandboxes including Modal, Prime, Docker, e2b, and Daytona, all communicating via a central LM Handler TCP server.**

The Recursive Language Model (RLM) framework by `alexzhang13/rlm` provides a flexible architecture for running language-model-backed REPLs across diverse execution contexts. Understanding the available deployment strategies for RLM allows you to choose the right balance of isolation, scalability, and latency for your specific use case—from quick local prototyping to production-grade cloud execution.

## Core Architecture Overview

RLM’s architecture consists of three distinct layers that enable consistent deployment across environments:

1. **Core Engine** – [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py) orchestrates the REPL loop and tracks usage metadata for each execution.
2. **LM Handler** – [`rlm/core/lm_handler.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/lm_handler.py) operates as a multi-threaded TCP server exposing a length-prefixed JSON protocol. All REPL environments, whether local or remote, dispatch LM requests to this handler via the communication utilities in [`rlm/core/comms_utils.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/comms_utils.py).
3. **Execution Environments** – Concrete implementations of `NonIsolatedEnv` or `IsolatedEnv` (defined in [`rlm/environments/base_env.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/base_env.py)) provide the actual code execution sandbox.

This separation of concerns ensures that the same user code runs identically regardless of whether you choose a local or cloud deployment strategy.

## Available RLM Deployment Strategies

Depending on your isolation requirements and infrastructure preferences, RLM offers six distinct deployment strategies:

### Local REPL

**Local REPL** runs in the same process as the caller with no OS-level isolation, utilizing Python’s built-in sandbox mechanisms. This strategy minimizes latency and requires zero configuration.

- **Implementation**: [`rlm/environments/local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/local_repl.py)
- **Best for**: Quick prototyping, Jupyter notebooks, and unit testing in the `tests/` directory
- **Isolation**: None (in-process)

Instantiate `LocalREPL` with a handler address (or let it start automatically), and execute code immediately without network overhead.

### Modal REPL

**Modal REPL** deploys to Modal’s serverless cloud sandbox, providing full OS isolation while communicating with the host via an HTTP broker pattern.

- **Implementation**: [`rlm/environments/modal_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/modal_repl.py)
- **Best for**: Scalable cloud execution with cheap on-demand compute
- **Isolation**: Full OS isolation via Modal containers

Use [`examples/modal_repl_example.py`](https://github.com/alexzhang13/rlm/blob/main/examples/modal_repl_example.py) to create a Modal sandbox, start the broker server, and execute code. The broker automatically installs packages listed in [`rlm/environments/constants.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/constants.py), including scientific libraries like `numpy`, `pandas`, and `sympy`.

### Prime REPL

**Prime REPL** functions similarly to Modal, deploying to Prime’s managed container infrastructure for workloads requiring Prime’s specific platform capabilities.

- **Implementation**: [`rlm/environments/prime_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/prime_repl.py)
- **Best for**: Deployments requiring Prime’s managed containers
- **Isolation**: Full OS isolation via Prime’s tunnel API

The `PrimeREPL` class handles tunnel creation through Prime’s SDK, maintaining the same broker architecture as Modal.

### Docker REPL

**Docker REPL** launches containers locally or on remote hosts, offering container-level isolation through Docker’s port-forwarding mechanisms.

- **Implementation**: [`rlm/environments/docker_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/docker_repl.py)
- **Best for**: Self-hosted, reproducible environments with custom images
- **Isolation**: Container-level via Docker

Run [`examples/docker_repl_example.py`](https://github.com/alexzhang13/rlm/blob/main/examples/docker_repl_example.py) to build an image from the included `Dockerfile` and launch a container running the broker server. The host connects to the container’s exposed port to forward LM requests.

### e2b REPL

**e2b REPL** leverages the Execution-as-a-Service platform, providing cloud sandboxes with persistent state across multiple executions.

- **Implementation**: [`rlm/environments/e2b_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/e2b_repl.py)
- **Best for**: Long-running workloads without infrastructure management
- **Isolation**: Cloud sandbox with persistent state

The `E2BRepl` class manages authentication to the e2b platform and runs the broker inside an e2b sandbox.

### Daytona REPL

**Daytona REPL** integrates with the Daytona remote development platform, offering container-level isolation with live VSCode-like editing capabilities.

- **Implementation**: [`rlm/environments/daytona_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/daytona_repl.py)
- **Best for**: Interactive development on remote clusters
- **Isolation**: Container-level with forwarded ports

`DaytonaRepl` spins up a remote development container, exposing the broker over a port that Daytona forwards to your local machine.

## How Isolated REPLs Communicate

All isolated environments (Modal, Prime, Docker, e2b, Daytona) share a common broker pattern architecture documented in [`AGENTS.md`](https://github.com/alexzhang13/rlm/blob/main/AGENTS.md):

1. **Broker Server** – A lightweight Flask app inside the sandbox provides three endpoints: `/enqueue`, `/pending`, and `/respond`. The sandbox’s `llm_query` functions POST to `/enqueue` and block until receiving a response.
2. **Poller Thread** – The host process running the LM Handler repeatedly calls `/pending` to collect requests, forwards each to the `LMHandler` via TCP socket, and POSTs responses back to `/respond`.
3. **Secure Tunneling** – Communication occurs over secure tunnels (Modal’s `encrypted_ports` or Prime’s tunnel API), ensuring the sandbox never requires direct network access to the LM Handler.

This architecture satisfies strict security policies while maintaining seamless integration between isolated code and language models.

## LM Client Configuration

RLM ships with client wrappers in `rlm/clients/` for OpenAI, Azure OpenAI, Anthropic, Gemini, and Portkey. Each client reads credentials from environment variables:

- `OPENAI_API_KEY`
- `AZURE_OPENAI_DEPLOYMENT`
- `ANTHROPIC_API_KEY`

When launching any REPL deployment strategy, the LM Handler loads the appropriate client based on the `model` argument passed to `llm_query` or `rlm_query`.

## Practical Deployment Example

A typical cloud deployment using Modal follows this pattern:

```python
from rlm.environments.modal_repl import ModalREPL
from rlm.core.lm_handler import LMHandler

# 1. Start the LM handler (once per host)

handler = LMHandler()
handler.start()  # listens on a free port, e.g. (host, 12345)

# 2. Launch the Modal REPL, pointing it at the handler

repl = ModalREPL(lm_handler_address=handler.address)
with repl:
    # 3. Execute code inside the sandbox

    result = repl.execute_code("""
answer["content"] = "Hello from Modal!"
answer["ready"] = True
""")
print(result.final_answer)   # → Hello from Modal!

```

This same code structure works with `DockerREPL`, `PrimeREPL`, or `E2BRepl` by simply changing the class import. All REPLs expose identical Python globals (`llm_query`, `rlm_query`, `SHOW_VARS`, `answer`, `context`, `history`), ensuring code portability across deployment strategies.

## Summary

- **Six deployment strategies** cover the full spectrum from local prototyping to enterprise cloud: Local, Modal, Prime, Docker, e2b, and Daytona.
- **Centralized LM Handler** in [`rlm/core/lm_handler.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/lm_handler.py) provides a consistent TCP interface regardless of where code executes.
- **Broker architecture** enables secure communication between isolated sandboxes and the LM Handler without direct network access.
- **Environment parity** ensures that code written for `LocalREPL` runs unchanged in any cloud sandbox, with automatic package provisioning via [`rlm/environments/constants.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/constants.py).
- **Client flexibility** supports major LLM providers through environment-variable-based configuration in `rlm/clients/`.

## Frequently Asked Questions

### What is the fastest deployment strategy for RLM development and testing?

**Local REPL** provides the fastest iteration cycle because it runs in-process with no network latency or container startup overhead. Use [`rlm/environments/local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/local_repl.py) for unit testing and notebook-based prototyping where isolation is not required.

### How does RLM maintain security when running code in cloud sandboxes?

Isolated REPLs use a broker pattern where the sandbox communicates only with a Flask broker server inside the container, while the host process polls via secure tunnels (Modal’s `encrypted_ports` or Prime’s tunnel API). The sandbox never possesses direct network access to the LM Handler or your API keys, satisfying strict security policies.

### Can I switch between deployment strategies without modifying my RLM code?

Yes. All REPL implementations expose identical Python globals (`llm_query`, `rlm_query`, `answer`, etc.) and implement the same interface from [`rlm/environments/base_env.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/base_env.py). You can switch from `LocalREPL` to `ModalREPL` or `DockerREPL` by changing only the import statement and class instantiation.

### Which deployment strategy should I choose for long-running production workloads?

**e2b REPL** or **Daytona REPL** are optimal for long-running jobs because they maintain persistent sandbox state across multiple calls without requiring you to manage container lifecycles. For scalable serverless execution, choose **Modal REPL** or **Prime REPL**, which automatically shut down when idle to minimize costs.