RLM Deployment Strategies: How to Deploy Recursive Language Models in Local and Cloud Environments
RLM supports six deployment strategies ranging from in-process local execution to fully isolated cloud sandboxes including Modal, Prime, Docker, e2b, and Daytona, all communicating via a central LM Handler TCP server.
The Recursive Language Model (RLM) framework by alexzhang13/rlm provides a flexible architecture for running language-model-backed REPLs across diverse execution contexts. Understanding the available deployment strategies for RLM allows you to choose the right balance of isolation, scalability, and latency for your specific use case—from quick local prototyping to production-grade cloud execution.
Core Architecture Overview
RLM’s architecture consists of three distinct layers that enable consistent deployment across environments:
- Core Engine –
rlm/core/rlm.pyorchestrates the REPL loop and tracks usage metadata for each execution. - LM Handler –
rlm/core/lm_handler.pyoperates as a multi-threaded TCP server exposing a length-prefixed JSON protocol. All REPL environments, whether local or remote, dispatch LM requests to this handler via the communication utilities inrlm/core/comms_utils.py. - Execution Environments – Concrete implementations of
NonIsolatedEnvorIsolatedEnv(defined inrlm/environments/base_env.py) provide the actual code execution sandbox.
This separation of concerns ensures that the same user code runs identically regardless of whether you choose a local or cloud deployment strategy.
Available RLM Deployment Strategies
Depending on your isolation requirements and infrastructure preferences, RLM offers six distinct deployment strategies:
Local REPL
Local REPL runs in the same process as the caller with no OS-level isolation, utilizing Python’s built-in sandbox mechanisms. This strategy minimizes latency and requires zero configuration.
- Implementation:
rlm/environments/local_repl.py - Best for: Quick prototyping, Jupyter notebooks, and unit testing in the
tests/directory - Isolation: None (in-process)
Instantiate LocalREPL with a handler address (or let it start automatically), and execute code immediately without network overhead.
Modal REPL
Modal REPL deploys to Modal’s serverless cloud sandbox, providing full OS isolation while communicating with the host via an HTTP broker pattern.
- Implementation:
rlm/environments/modal_repl.py - Best for: Scalable cloud execution with cheap on-demand compute
- Isolation: Full OS isolation via Modal containers
Use examples/modal_repl_example.py to create a Modal sandbox, start the broker server, and execute code. The broker automatically installs packages listed in rlm/environments/constants.py, including scientific libraries like numpy, pandas, and sympy.
Prime REPL
Prime REPL functions similarly to Modal, deploying to Prime’s managed container infrastructure for workloads requiring Prime’s specific platform capabilities.
- Implementation:
rlm/environments/prime_repl.py - Best for: Deployments requiring Prime’s managed containers
- Isolation: Full OS isolation via Prime’s tunnel API
The PrimeREPL class handles tunnel creation through Prime’s SDK, maintaining the same broker architecture as Modal.
Docker REPL
Docker REPL launches containers locally or on remote hosts, offering container-level isolation through Docker’s port-forwarding mechanisms.
- Implementation:
rlm/environments/docker_repl.py - Best for: Self-hosted, reproducible environments with custom images
- Isolation: Container-level via Docker
Run examples/docker_repl_example.py to build an image from the included Dockerfile and launch a container running the broker server. The host connects to the container’s exposed port to forward LM requests.
e2b REPL
e2b REPL leverages the Execution-as-a-Service platform, providing cloud sandboxes with persistent state across multiple executions.
- Implementation:
rlm/environments/e2b_repl.py - Best for: Long-running workloads without infrastructure management
- Isolation: Cloud sandbox with persistent state
The E2BRepl class manages authentication to the e2b platform and runs the broker inside an e2b sandbox.
Daytona REPL
Daytona REPL integrates with the Daytona remote development platform, offering container-level isolation with live VSCode-like editing capabilities.
- Implementation:
rlm/environments/daytona_repl.py - Best for: Interactive development on remote clusters
- Isolation: Container-level with forwarded ports
DaytonaRepl spins up a remote development container, exposing the broker over a port that Daytona forwards to your local machine.
How Isolated REPLs Communicate
All isolated environments (Modal, Prime, Docker, e2b, Daytona) share a common broker pattern architecture documented in AGENTS.md:
- Broker Server – A lightweight Flask app inside the sandbox provides three endpoints:
/enqueue,/pending, and/respond. The sandbox’sllm_queryfunctions POST to/enqueueand block until receiving a response. - Poller Thread – The host process running the LM Handler repeatedly calls
/pendingto collect requests, forwards each to theLMHandlervia TCP socket, and POSTs responses back to/respond. - Secure Tunneling – Communication occurs over secure tunnels (Modal’s
encrypted_portsor Prime’s tunnel API), ensuring the sandbox never requires direct network access to the LM Handler.
This architecture satisfies strict security policies while maintaining seamless integration between isolated code and language models.
LM Client Configuration
RLM ships with client wrappers in rlm/clients/ for OpenAI, Azure OpenAI, Anthropic, Gemini, and Portkey. Each client reads credentials from environment variables:
OPENAI_API_KEYAZURE_OPENAI_DEPLOYMENTANTHROPIC_API_KEY
When launching any REPL deployment strategy, the LM Handler loads the appropriate client based on the model argument passed to llm_query or rlm_query.
Practical Deployment Example
A typical cloud deployment using Modal follows this pattern:
from rlm.environments.modal_repl import ModalREPL
from rlm.core.lm_handler import LMHandler
# 1. Start the LM handler (once per host)
handler = LMHandler()
handler.start() # listens on a free port, e.g. (host, 12345)
# 2. Launch the Modal REPL, pointing it at the handler
repl = ModalREPL(lm_handler_address=handler.address)
with repl:
# 3. Execute code inside the sandbox
result = repl.execute_code("""
answer["content"] = "Hello from Modal!"
answer["ready"] = True
""")
print(result.final_answer) # → Hello from Modal!
This same code structure works with DockerREPL, PrimeREPL, or E2BRepl by simply changing the class import. All REPLs expose identical Python globals (llm_query, rlm_query, SHOW_VARS, answer, context, history), ensuring code portability across deployment strategies.
Summary
- Six deployment strategies cover the full spectrum from local prototyping to enterprise cloud: Local, Modal, Prime, Docker, e2b, and Daytona.
- Centralized LM Handler in
rlm/core/lm_handler.pyprovides a consistent TCP interface regardless of where code executes. - Broker architecture enables secure communication between isolated sandboxes and the LM Handler without direct network access.
- Environment parity ensures that code written for
LocalREPLruns unchanged in any cloud sandbox, with automatic package provisioning viarlm/environments/constants.py. - Client flexibility supports major LLM providers through environment-variable-based configuration in
rlm/clients/.
Frequently Asked Questions
What is the fastest deployment strategy for RLM development and testing?
Local REPL provides the fastest iteration cycle because it runs in-process with no network latency or container startup overhead. Use rlm/environments/local_repl.py for unit testing and notebook-based prototyping where isolation is not required.
How does RLM maintain security when running code in cloud sandboxes?
Isolated REPLs use a broker pattern where the sandbox communicates only with a Flask broker server inside the container, while the host process polls via secure tunnels (Modal’s encrypted_ports or Prime’s tunnel API). The sandbox never possesses direct network access to the LM Handler or your API keys, satisfying strict security policies.
Can I switch between deployment strategies without modifying my RLM code?
Yes. All REPL implementations expose identical Python globals (llm_query, rlm_query, answer, etc.) and implement the same interface from rlm/environments/base_env.py. You can switch from LocalREPL to ModalREPL or DockerREPL by changing only the import statement and class instantiation.
Which deployment strategy should I choose for long-running production workloads?
e2b REPL or Daytona REPL are optimal for long-running jobs because they maintain persistent sandbox state across multiple calls without requiring you to manage container lifecycles. For scalable serverless execution, choose Modal REPL or Prime REPL, which automatically shut down when idle to minimize costs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →