Main Architectural Components of the RLM Repository: Core Engine, Environment Layer, and Client Adapters

The RLM repository is organized into four distinct pillars: a Core Engine that orchestrates REPL execution and LLM communication, an Environment Layer providing sandboxed execution contexts, a Client Layer standardizing access to multiple LLM providers, and Utility modules handling parsing, logging, and token management.

The alexzhang13/rlm (Recursive Language Models) codebase implements a modular architecture designed to separate model-agnostic orchestration from provider-specific logic and execution sandboxing. This design allows developers to swap between local Python execution and cloud-isolated sandboxes like Modal or Prime without changing the underlying LLM interaction patterns. Understanding these main architectural components is essential for extending the framework or debugging recursive query flows.

Core Engine: The Orchestration Layer

The Core Engine serves as the central nervous system of the RLM repository, managing the REPL execution loop, mediating communication between environments and language models, and tracking usage statistics across sessions.

Main RLM Class

At rlm/core/rlm.py, the RLM class provides the primary entry point for executing recursive language model queries. This class initializes execution environments, registers LLM clients, and exposes the rlm_query() and rlm_query_batched() methods. The engine watches for the special answer dictionary populated by executing REPL code to determine when a turn completes, enabling recursive chains where model outputs generate subsequent code execution.

Communication Infrastructure

The engine relies on two critical communication modules located in rlm/core/lm_handler.py and rlm/core/comms_utils.py. The LM handler implements a socket-based server that forwards LLM requests from sandboxed environments to the appropriate client adapters. For isolated environments such as Modal or Prime, this handler acts as an HTTP broker, ensuring that network-restricted sandboxes can still access LLM services without exposing API keys directly.

The communication utilities (rlm/core/comms_utils.py) implement a length-prefixed JSON messaging protocol that ensures reliable data transmission between the REPL environment and the handler, preventing truncation issues common in socket-based inter-process communication.

Type Definitions

Data structures defining the contract between components reside in rlm/core/types.py. Key classes include:

  • REPLResult – encapsulates execution output, return values, and error states from sandboxed code
  • ModelUsageSummary – aggregates token counts and cost data across multiple LLM calls
  • UsageSummary – provides session-level statistics for budget tracking and optimization

Environment Layer: Sandboxed Execution

The Environment Layer abstracts the execution context where user code runs, supporting everything from local unconstrained execution to fully isolated cloud sandboxes.

Base Environment Interface

All environments inherit from rlm/environments/base_env.py, which defines the BaseEnv abstract class specifying the required interface for REPL initialization, code execution, and communication setup. This abstraction ensures that the Core Engine remains agnostic to whether code executes locally or in a remote container.

Concrete Environment Implementations

The repository provides seven distinct environment implementations:

Each implementation handles the specific setup required for llm_query() and rlm_query() availability within the sandbox, ensuring these functions route through the appropriate communication channel back to the Core Engine.

Client Layer: LLM Provider Abstraction

The Client Layer standardizes access to diverse LLM providers through a unified interface, enabling seamless switching between OpenAI, Anthropic, Gemini, and other backends.

Base Client Interface

rlm/clients/base_lm.py defines the BaseLM abstract class, which mandates methods for completion generation, token counting, and cost tracking via _track_cost(). This interface ensures that the Core Engine can treat all providers identically regardless of their underlying API differences.

Provider Implementations

Concrete clients adapter various cloud services:

Each client handles provider-specific authentication, request formatting, and response parsing while exposing standardized token usage statistics that feed into the ModelUsageSummary aggregations.

Utility and Support Modules

Supporting functionality spans parsing, prompt construction, logging, and error handling across rlm/utils/ and rlm/logger/.

Parsing and Prompt Management

rlm/utils/parsing.py extracts structured data from LLM responses, handling various output formats and metadata extraction. rlm/utils/prompts.py provides pre-defined system prompts and utilities for constructing message arrays that conform to each provider's expected format.

Token and Logging Infrastructure

rlm/utils/token_utils.py implements token counting algorithms matching provider-specific tokenizers, while rlm/logger/rlm_logger.py provides structured logging for debugging recursive execution flows. Error handling consolidates in rlm/utils/exceptions.py, defining custom exception types for environment crashes, LLM timeouts, and communication failures.

How the Architecture Fits Together

The data flow through the main architectural components follows a precise sequence:

  1. Initialization – The RLM class instantiates an environment (e.g., LocalREPL or ModalREPL) and registers one or more BaseLM clients
  2. Execution – User code runs inside the environment and calls llm_query() or rlm_query()
  3. Routing – The call routes through rlm/core/comms_utils.py to the lm_handler.py socket server
  4. Completion – The handler selects the appropriate client (e.g., OpenAI from rlm/clients/openai.py) and obtains the LLM response
  5. Return – Results flow back to the executing REPL, populating the answer dictionary that signals turn completion to the Core Engine
  6. Tracking – Clients report usage via _track_cost(), aggregating in ModelUsageSummary instances for final reporting

This pipeline ensures that adding a new provider requires only subclassing BaseLM, while new execution environments extend BaseEnv without touching orchestration logic.

Practical Implementation Examples

Below are runnable patterns demonstrating the architectural interactions.

Local Execution with OpenAI

from rlm.core.rlm import RLM
from rlm.clients.openai import OpenAI
from rlm.environments.local_repl import LocalREPL

# Initialize LLM client

openai_client = OpenAI(api_key="YOUR_OPENAI_KEY", model_name="gpt-4o-mini")

# Select local execution environment

env = LocalREPL()

# Build RLM engine

rlm = RLM(env=env, lm_clients=[openai_client])

# Execute recursive query

result = rlm.rlm_query("Write a Python function that returns the nth Fibonacci number.")
print(result.final_answer)

Cloud-Isolated Environment with Modal

from rlm.core.rlm import RLM
from rlm.clients.portkey import Portkey
from rlm.environments.modal_repl import ModalREPL

client = Portkey(api_key="PORTKEY_KEY", model_name="gpt-4")
env = ModalREPL()  # Spins up Modal sandbox automatically

rlm = RLM(env=env, lm_clients=[client])

prompts = [
    "Summarize the Python GIL in two sentences.",
    "Explain why recursion can cause a stack overflow."
]
batch_result = rlm.rlm_query_batched(prompts)
for answer in batch_result:
    print(answer.final_answer)

Inspecting Usage Statistics

usage = rlm.get_usage_summary()
print(f"Total input tokens: {usage.input_tokens}")
print(f"Total output tokens: {usage.output_tokens}")

Summary

  • Four Pillars: The RLM architecture consists of the Core Engine (rlm/core/), Environment Layer (rlm/environments/), Client Layer (rlm/clients/), and Utility modules (rlm/utils/)
  • Core Engine: rlm/core/rlm.py orchestrates execution while rlm/core/lm_handler.py manages socket-based LLM communication via length-prefixed JSON protocols
  • Environment Abstraction: Seven implementations from LocalREPL to ModalREPL provide flexible sandboxing strategies without changing the query interface
  • Provider Agnostic: The BaseLM interface in rlm/clients/base_lm.py standardizes OpenAI, Anthropic, Gemini, and other providers with unified cost tracking
  • Extensible Design: New providers subclass BaseLM and new environments extend BaseEnv, allowing the Core Engine to remain unchanged when adding functionality

Frequently Asked Questions

How does the RLM repository handle communication between isolated environments and the LLM handler?

Isolated environments like ModalREPL or PrimeREPL communicate with the Core Engine through an HTTP broker pattern implemented in rlm/core/lm_handler.py. The environment-side code sends requests via HTTP to a host-side socket server, which then forwards the request to the appropriate LLM client. This architecture prevents API keys from residing in sandboxed containers while maintaining low-latency communication through the length-prefixed JSON protocol defined in rlm/core/comms_utils.py.

What is the difference between rlm_query() and llm_query() in the RLM architecture?

rlm_query() is the high-level method exposed by the RLM class in rlm/core/rlm.py that manages the full recursive execution loop, including environment setup, REPL monitoring, and multiple turns of interaction. llm_query() is the lower-level function available within the execution environment itself that dispatches single completion requests to the LM handler. While rlm_query() orchestrates the entire session, llm_query() provides direct access to the LLM client from within sandboxed code.

Which environment should I use for secure execution of untrusted code in RLM?

For untrusted code, use isolated environments such as ModalREPL (rlm/environments/modal_repl.py), E2BREPL (rlm/environments/e2b_repl.py), or DockerREPL (rlm/environments/docker_repl.py). These implementations spin up ephemeral sandboxes with restricted network access and filesystem isolation. Avoid LocalREPL for untrusted code as it executes in the host Python process without sandboxing.

How do I add a new LLM provider to the RLM client layer?

Create a new file in rlm/clients/ (e.g., rlm/clients/new_provider.py) and subclass BaseLM from rlm/clients/base_lm.py. Implement the required abstract methods for completion generation, token counting, and cost tracking via _track_cost(). Register the new client class when initializing the RLM engine. The Core Engine will automatically route requests to your new provider using the same communication infrastructure without requiring changes to the environment or handler code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →