How Configuration Works in RLM: A Complete Guide to the Recursive Language Model

Configuration in RLM is centralized through the RLM class constructor, which stores parameters as immutable instance variables, validates them against RLMMetadata, and enforces runtime limits like budget, timeout, and recursion depth during the completion loop.

The Recursive Language Model (RLM) is a highly-configurable wrapper around language model clients that adds REPL-style execution, recursion, and resource management. Understanding how configuration in RLM works is essential for tuning the pipeline's behavior, from selecting backend providers to enforcing strict cost and time budgets. All configuration flows through the RLM class in rlm/core/rlm.py and is captured by the RLMMetadata dataclass for reproducibility.

Configuration Architecture

Constructor and Validation

In rlm/core/rlm.py, the RLM.__init__ method serves as the single entry point for all configuration. It normalizes sampling arguments (lines 23-38), validates the other_backends list (lines 47-55), and stores values as immutable instance attributes. This immutability guarantees that once constructed, the configuration snapshot remains reproducible throughout the recursion tree.

RLMMetadata Snapshot

The RLMMetadata dataclass defined in rlm/core/types.py captures a serializable snapshot of the configuration at initialization time (lines 106-124 of rlm.py). This metadata travels with every completion, enabling perfect reconstruction of the execution environment for debugging and logging purposes.

Core Configuration Parameters

The RLM constructor accepts over twenty parameters that control every aspect of recursive inference:

  • backend and backend_kwargs: Specify the LLM client ("openai", "anthropic", "gemini") and its initialization arguments like model_name. These are passed to get_client at line 34 of rlm.py.
  • environment and environment_kwargs: Define the REPL type ("local", "ipython", "docker") and its specific settings. The default "local" environment is instantiated in _spawn_completion_context (lines 71-74).
  • depth and max_depth: Track current recursion depth and the hard limit. When depth >= max_depth, the system falls back to plain LM calls (checked in completion at lines 47-49 and _subcall at lines 31-33).
  • max_iterations: Caps the number of REPL iterations before forcing a default answer, enforced in the completion loop (lines 59-61).
  • max_budget, max_timeout, and max_tokens: Set the USD cost ceiling, wall-clock seconds, and total token budget. These are validated in _check_iteration_limits (lines 50-57, 66-71) and _check_timeout (lines 90-100).
  • max_errors: Limits consecutive code-execution failures before abort, checked in _check_iteration_limits (lines 31-38).
  • custom_system_prompt: Replaces the default RLM_SYSTEM_PROMPT when assigned to self.system_prompt (lines 74-75).
  • other_backends and other_backend_kwargs: Configure alternative clients for recursive sub-calls (depth > 0), validated at lines 47-55 and registered in _spawn_completion_context (lines 46-55).
  • custom_tools and custom_sub_tools: Inject Python functions into the REPL globals for root and child RLMs, passed via env_kwargs (lines 78-82).
  • compaction and compaction_threshold_pct: Enable history summarization when context reaches 85% of model limits (default 0.85), implemented in _compact_history (lines 100-142) and _get_compaction_status (lines 85-93).
  • max_concurrent_subcalls: Limits parallelism for batched rlm_query_batched operations.
  • verbose: Activates rich-formatted console output via VerbosePrinter (lines 81-82).
  • logger: Accepts an RLMLogger instance to capture full trajectories.

Runtime Configuration Flow

Initialization Phase

During construction, the RLM class performs three critical steps: normalizing arguments, validating persistence requirements (ensuring docker environments don't attempt persistence, which raises ValueError in _validate_persistent_environment_support at lines 71-84), and creating the RLMMetadata snapshot.

Per-Completion Context

Each call to completion spawns a fresh execution context via _spawn_completion_context:

  1. A BaseLM client is built from backend and backend_kwargs.
  2. If other_backends exist, additional clients are built and registered with the LMHandler.
  3. The LMHandler socket server starts at line 56.
  4. An environment instance is created via get_environment(self.environment_type, env_kwargs), where env_kwargs contains the LM handler address, initial prompts, depth, and custom tools.

Iteration Enforcement

The main completion loop consults stored configuration limits before every iteration:

  • Timeout: _check_timeout validates wall-clock time against max_timeout (lines 90-100).
  • Budget and Tokens: _check_iteration_limits enforces max_budget, max_tokens, and max_errors after each iteration (lines 50-71).
  • Compaction: When enabled, the loop checks _get_compaction_status and triggers _compact_history if history exceeds compaction_threshold_pct of the context window (lines 63-74).

Recursive Propagation

When environments invoke rlm_query, the parent spawns a child RLM via _subcall with incremented depth and reduced budget/timeout allowances. This child inherits the same configuration structure, ensuring consistent enforcement throughout the recursion tree.

Practical Configuration Examples

Basic Setup

from rlm.core.rlm import RLM

# Simple RLM using OpenAI's gpt-4 in local REPL

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    environment="local",
    max_depth=2,               # Allow one recursive sub-call

    max_iterations=20,
    max_budget=5.0,            # $5 cap

    max_timeout=120,           # 2-minute wall-clock limit

)

Injecting Custom Tools

def fetch_url(url: str) -> str:
    """Helper injected into the REPL namespace."""
    import urllib.request
    return urllib.request.urlopen(url).read().decode()

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    custom_tools={"fetch_url": fetch_url},
    environment="local",
)

Multi-Backend Recursion

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    other_backends=["anthropic"],          # Child RLMs use Claude

    other_backend_kwargs=[{"model_name": "claude-2"}],
    max_depth=3,
)

Persistent Execution Mode

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    environment="local",
    persistent=True,           # Reuse the same REPL across calls

)

answer1 = rlm.completion("What is the capital of France?").response
answer2 = rlm.completion("What about its population?").response  # Same REPL context

rlm.close()   # Clean up

Strict Resource Limits

try:
    rlm = RLM(
        backend="openai",
        backend_kwargs={"model_name": "gpt-4"},
        max_budget=0.10,   # $0.10 limit

        max_timeout=30,    # 30-second limit

    )
    result = rlm.completion("Run a long simulation.").response
except Exception as e:
    print("Stopped early:", e)  # Raises BudgetExceededError or TimeoutExceededError

Key Source Files

According to the alexzhang13/rlm source code, configuration logic spans these critical files:

  • rlm/core/rlm.py: Central class handling constructor validation, parameter storage, and runtime enforcement via completion, _check_iteration_limits, and _check_timeout.
  • rlm/core/types.py: Defines the RLMMetadata dataclass for configuration snapshots and RLMChatCompletion return types.
  • rlm/environments/base_env.py: Specifies the SupportsPersistence protocol with methods like update_handler_address, add_context, and get_context_count.
  • rlm/environments/local_repl.py: Implements the default local REPL that respects custom_tools and persistence APIs.
  • rlm/core/lm_handler.py: The LMHandler socket server that forwards LM requests from environments and manages other_backends registration.
  • rlm/logger/rlm_logger.py: Captures full execution trajectories when a logger instance is provided to the constructor.
  • rlm/utils/prompts.py: Builds system prompts and compaction metadata based on configuration settings.

Summary

Configuration in RLM follows a strict pipeline that ensures reproducibility and resource safety:

  • Immutable Setup: The RLM constructor in rlm/core/rlm.py normalizes and stores all parameters as instance variables, creating a frozen configuration state.
  • Metadata Capture: The RLMMetadata dataclass serializes the initial configuration for logging and debugging.
  • Context Isolation: Each completion call spawns fresh LMHandler and Environment instances configured with the stored parameters.
  • Runtime Enforcement: The main loop actively checks max_budget, max_timeout, max_tokens, max_errors, and max_depth on every iteration, raising typed exceptions when limits are exceeded.
  • Recursive Inheritance: Child RLMs created via _subcall inherit configuration but track independent depth and resource consumption.

Frequently Asked Questions

What happens if I exceed the max_budget or max_timeout in RLM?

The completion method raises BudgetExceededError or TimeoutExceededError immediately when _check_iteration_limits or _check_timeout detects a violation. These checks run after every REPL iteration and before the next one begins, ensuring no further API calls incur costs once limits are reached.

Can I use different LLM providers for the main RLM and its recursive sub-calls?

Yes. Pass other_backends and other_backend_kwargs to the constructor. According to the validation logic in rlm/core/rlm.py (lines 47-55), the system supports one additional backend for sub-calls. When rlm_query triggers a sub-call, the child RLM uses the alternate client while maintaining the same environment and tool configuration.

How does RLM handle long conversations that exceed the context window?

When compaction=True (enabled by default), RLM monitors context usage via _get_compaction_status. If history exceeds compaction_threshold_pct (default 0.85) of the model's context length, the _compact_history method (lines 100-142) summarizes older messages to free tokens, allowing the conversation to continue indefinitely without truncation.

Why does my RLM configuration fail when enabling persistent mode with Docker?

The _validate_persistent_environment_support method (lines 71-84 of rlm.py) explicitly checks environment compatibility. Persistence requires the SupportsPersistence protocol defined in rlm/environments/base_env.py, which only local and ipython environments implement. Attempting to set persistent=True with environment="docker" raises a ValueError because container state cannot be reliably preserved across calls in the current implementation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →