Main Architectural Components of the RLM Repository: Core Engine, Environment Layer, and Client Adapters
The RLM repository is organized into four distinct pillars: a Core Engine that orchestrates REPL execution and LLM communication, an Environment Layer providing sandboxed execution contexts, a Client Layer standardizing access to multiple LLM providers, and Utility modules handling parsing, logging, and token management.
The alexzhang13/rlm (Recursive Language Models) codebase implements a modular architecture designed to separate model-agnostic orchestration from provider-specific logic and execution sandboxing. This design allows developers to swap between local Python execution and cloud-isolated sandboxes like Modal or Prime without changing the underlying LLM interaction patterns. Understanding these main architectural components is essential for extending the framework or debugging recursive query flows.
Core Engine: The Orchestration Layer
The Core Engine serves as the central nervous system of the RLM repository, managing the REPL execution loop, mediating communication between environments and language models, and tracking usage statistics across sessions.
Main RLM Class
At rlm/core/rlm.py, the RLM class provides the primary entry point for executing recursive language model queries. This class initializes execution environments, registers LLM clients, and exposes the rlm_query() and rlm_query_batched() methods. The engine watches for the special answer dictionary populated by executing REPL code to determine when a turn completes, enabling recursive chains where model outputs generate subsequent code execution.
Communication Infrastructure
The engine relies on two critical communication modules located in rlm/core/lm_handler.py and rlm/core/comms_utils.py. The LM handler implements a socket-based server that forwards LLM requests from sandboxed environments to the appropriate client adapters. For isolated environments such as Modal or Prime, this handler acts as an HTTP broker, ensuring that network-restricted sandboxes can still access LLM services without exposing API keys directly.
The communication utilities (rlm/core/comms_utils.py) implement a length-prefixed JSON messaging protocol that ensures reliable data transmission between the REPL environment and the handler, preventing truncation issues common in socket-based inter-process communication.
Type Definitions
Data structures defining the contract between components reside in rlm/core/types.py. Key classes include:
REPLResult– encapsulates execution output, return values, and error states from sandboxed codeModelUsageSummary– aggregates token counts and cost data across multiple LLM callsUsageSummary– provides session-level statistics for budget tracking and optimization
Environment Layer: Sandboxed Execution
The Environment Layer abstracts the execution context where user code runs, supporting everything from local unconstrained execution to fully isolated cloud sandboxes.
Base Environment Interface
All environments inherit from rlm/environments/base_env.py, which defines the BaseEnv abstract class specifying the required interface for REPL initialization, code execution, and communication setup. This abstraction ensures that the Core Engine remains agnostic to whether code executes locally or in a remote container.
Concrete Environment Implementations
The repository provides seven distinct environment implementations:
rlm/environments/local_repl.py– Executes code in the current Python process without isolation, suitable for trusted workflowsrlm/environments/modal_repl.py– Spins up ephemeral Modal sandboxes with HTTP broker communicationrlm/environments/prime_repl.py– Integrates with Prime cloud infrastructure for isolated executionrlm/environments/ipython_repl.py– Leverages IPython kernels for enhanced interactive featuresrlm/environments/e2b_repl.py– Uses E2B's sandbox-as-a-service for secure code executionrlm/environments/docker_repl.py– Containerizes execution using local Docker daemonrlm/environments/daytona_repl.py– Integrates with Daytona development environments
Each implementation handles the specific setup required for llm_query() and rlm_query() availability within the sandbox, ensuring these functions route through the appropriate communication channel back to the Core Engine.
Client Layer: LLM Provider Abstraction
The Client Layer standardizes access to diverse LLM providers through a unified interface, enabling seamless switching between OpenAI, Anthropic, Gemini, and other backends.
Base Client Interface
rlm/clients/base_lm.py defines the BaseLM abstract class, which mandates methods for completion generation, token counting, and cost tracking via _track_cost(). This interface ensures that the Core Engine can treat all providers identically regardless of their underlying API differences.
Provider Implementations
Concrete clients adapter various cloud services:
rlm/clients/openai.py– OpenAI API integration (GPT-4, GPT-4o-mini)rlm/clients/anthropic.py– Anthropic Claude series supportrlm/clients/gemini.py– Google Gemini API adapterrlm/clients/azure_openai.py– Azure OpenAI Service implementationrlm/clients/portkey.py– Portkey AI gateway for unified multi-provider access
Each client handles provider-specific authentication, request formatting, and response parsing while exposing standardized token usage statistics that feed into the ModelUsageSummary aggregations.
Utility and Support Modules
Supporting functionality spans parsing, prompt construction, logging, and error handling across rlm/utils/ and rlm/logger/.
Parsing and Prompt Management
rlm/utils/parsing.py extracts structured data from LLM responses, handling various output formats and metadata extraction. rlm/utils/prompts.py provides pre-defined system prompts and utilities for constructing message arrays that conform to each provider's expected format.
Token and Logging Infrastructure
rlm/utils/token_utils.py implements token counting algorithms matching provider-specific tokenizers, while rlm/logger/rlm_logger.py provides structured logging for debugging recursive execution flows. Error handling consolidates in rlm/utils/exceptions.py, defining custom exception types for environment crashes, LLM timeouts, and communication failures.
How the Architecture Fits Together
The data flow through the main architectural components follows a precise sequence:
- Initialization – The
RLMclass instantiates an environment (e.g.,LocalREPLorModalREPL) and registers one or moreBaseLMclients - Execution – User code runs inside the environment and calls
llm_query()orrlm_query() - Routing – The call routes through
rlm/core/comms_utils.pyto thelm_handler.pysocket server - Completion – The handler selects the appropriate client (e.g.,
OpenAIfromrlm/clients/openai.py) and obtains the LLM response - Return – Results flow back to the executing REPL, populating the
answerdictionary that signals turn completion to the Core Engine - Tracking – Clients report usage via
_track_cost(), aggregating inModelUsageSummaryinstances for final reporting
This pipeline ensures that adding a new provider requires only subclassing BaseLM, while new execution environments extend BaseEnv without touching orchestration logic.
Practical Implementation Examples
Below are runnable patterns demonstrating the architectural interactions.
Local Execution with OpenAI
from rlm.core.rlm import RLM
from rlm.clients.openai import OpenAI
from rlm.environments.local_repl import LocalREPL
# Initialize LLM client
openai_client = OpenAI(api_key="YOUR_OPENAI_KEY", model_name="gpt-4o-mini")
# Select local execution environment
env = LocalREPL()
# Build RLM engine
rlm = RLM(env=env, lm_clients=[openai_client])
# Execute recursive query
result = rlm.rlm_query("Write a Python function that returns the nth Fibonacci number.")
print(result.final_answer)
Cloud-Isolated Environment with Modal
from rlm.core.rlm import RLM
from rlm.clients.portkey import Portkey
from rlm.environments.modal_repl import ModalREPL
client = Portkey(api_key="PORTKEY_KEY", model_name="gpt-4")
env = ModalREPL() # Spins up Modal sandbox automatically
rlm = RLM(env=env, lm_clients=[client])
prompts = [
"Summarize the Python GIL in two sentences.",
"Explain why recursion can cause a stack overflow."
]
batch_result = rlm.rlm_query_batched(prompts)
for answer in batch_result:
print(answer.final_answer)
Inspecting Usage Statistics
usage = rlm.get_usage_summary()
print(f"Total input tokens: {usage.input_tokens}")
print(f"Total output tokens: {usage.output_tokens}")
Summary
- Four Pillars: The RLM architecture consists of the Core Engine (
rlm/core/), Environment Layer (rlm/environments/), Client Layer (rlm/clients/), and Utility modules (rlm/utils/) - Core Engine:
rlm/core/rlm.pyorchestrates execution whilerlm/core/lm_handler.pymanages socket-based LLM communication via length-prefixed JSON protocols - Environment Abstraction: Seven implementations from
LocalREPLtoModalREPLprovide flexible sandboxing strategies without changing the query interface - Provider Agnostic: The
BaseLMinterface inrlm/clients/base_lm.pystandardizes OpenAI, Anthropic, Gemini, and other providers with unified cost tracking - Extensible Design: New providers subclass
BaseLMand new environments extendBaseEnv, allowing the Core Engine to remain unchanged when adding functionality
Frequently Asked Questions
How does the RLM repository handle communication between isolated environments and the LLM handler?
Isolated environments like ModalREPL or PrimeREPL communicate with the Core Engine through an HTTP broker pattern implemented in rlm/core/lm_handler.py. The environment-side code sends requests via HTTP to a host-side socket server, which then forwards the request to the appropriate LLM client. This architecture prevents API keys from residing in sandboxed containers while maintaining low-latency communication through the length-prefixed JSON protocol defined in rlm/core/comms_utils.py.
What is the difference between rlm_query() and llm_query() in the RLM architecture?
rlm_query() is the high-level method exposed by the RLM class in rlm/core/rlm.py that manages the full recursive execution loop, including environment setup, REPL monitoring, and multiple turns of interaction. llm_query() is the lower-level function available within the execution environment itself that dispatches single completion requests to the LM handler. While rlm_query() orchestrates the entire session, llm_query() provides direct access to the LLM client from within sandboxed code.
Which environment should I use for secure execution of untrusted code in RLM?
For untrusted code, use isolated environments such as ModalREPL (rlm/environments/modal_repl.py), E2BREPL (rlm/environments/e2b_repl.py), or DockerREPL (rlm/environments/docker_repl.py). These implementations spin up ephemeral sandboxes with restricted network access and filesystem isolation. Avoid LocalREPL for untrusted code as it executes in the host Python process without sandboxing.
How do I add a new LLM provider to the RLM client layer?
Create a new file in rlm/clients/ (e.g., rlm/clients/new_provider.py) and subclass BaseLM from rlm/clients/base_lm.py. Implement the required abstract methods for completion generation, token counting, and cost tracking via _track_cost(). Register the new client class when initializing the RLM engine. The Core Engine will automatically route requests to your new provider using the same communication infrastructure without requiring changes to the environment or handler code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →