# Performance Considerations for RLM: Optimizing Recursive Language Model Execution

> Optimize RLM performance with concurrency limits, token budgets, and depth caps. Learn how RLM balances throughput and overhead for efficient recursive language model execution.

- Repository: [az/rlm](https://github.com/alexzhang13/rlm)
- Tags: performance
- Published: 2026-06-18

---

**RLM manages performance through concurrency limits, token budgets, and recursive depth caps, using thread pools and async semaphores to balance throughput against the overhead of process spawning and socket communication.**

The `alexzhang13/rlm` repository implements a Recursive Language Model that executes LLM calls inside a REPL-style environment capable of recursively spawning child RLMs. Because each recursion creates new processes, environments, and socket connections, the primary performance considerations for RLM center on managing parallelism, context size, and execution boundaries.

## Concurrency Architecture and Thread Management

### Thread-Pool Executors for Child RLMs

In [`rlm/environments/local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/local_repl.py), the `_rlm_query_batched` method uses a `ThreadPoolExecutor` to launch parallel child RLMs, constrained by the `max_concurrent_subcalls` parameter (default 4). This defines the upper bound of parallel sub-calls that can execute simultaneously from a single REPL block. Larger values increase throughput but raise CPU and memory pressure, and can trigger rate limits on hosted providers like OpenAI or Anthropic.

### Asyncio Semaphores for Batched Requests

The `LMHandler` class in [`rlm/core/lm_handler.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/lm_handler.py) manages concurrent LM requests through `batch_max_concurrent` (default 16), implemented via asyncio semaphores in `_handle_batched` (lines 88-92). Additionally, the `ThreadingLMServer` (lines 10-12) services each request on its own thread to prevent I/O blocking, enabling the socket server to handle many prompts concurrently.

## Token Limits and Context Compaction

### Hard Token Ceilings

The `_check_iteration_limits` method in [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py) (lines 66-78) enforces the `max_tokens` parameter by calculating total input plus output tokens and aborting execution when the budget is exceeded. This prevents runaway costs from unbounded generation.

### Automatic Context Summarization

When accumulated message history approaches the model’s context window, RLM triggers `_compact_history` (lines 63-77) to summarize the trajectory. The `compaction_threshold_pct` parameter (default 0.85) determines when this compaction occurs, keeping token usage bounded while adding minimal latency overhead. Lowering this threshold to 0.7 reduces prompt size but increases summarization frequency.

## Recursive Depth and Execution Boundaries

### Preventing Runaway Recursion

The `RLM.__init__` method defines `max_depth` and `max_iterations` (lines 55-58) to limit nested RLMs and REPL loops. The `_check_timeout` and `_check_iteration_limits` methods (lines 90-108) enforce wall-clock timeouts and iteration caps, aborting execution immediately when ceilings are hit to prevent infinite loops.

## Environment Persistence and Reuse

### Persistent REPL Processes

When `persistent=True`, `RLM._spawn_completion_context` reuses the `_persistent_env` (lines 58-71), avoiding the cost of reinitializing globals, temporary directories, and socket connections across multiple `completion()` calls. This can shave seconds off repeated calls but may increase memory consumption over long-running sessions as context and history versions accumulate.

## Key Performance Parameters

| Parameter | Source Location | Effect | Typical Range |
|-----------|----------------|--------|---------------|
| `max_concurrent_subcalls` | [`rlm/environments/local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/local_repl.py) | Limits parallel child RLMs via ThreadPoolExecutor | 1-8 (default 4) |
| `batch_max_concurrent` | [`rlm/core/lm_handler.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/lm_handler.py) | Async semaphore size for batched LM calls | 4-16 (default 16) |
| `max_iterations` | [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py) | Caps REPL loops per completion | 10-30 (default 30) |
| `max_tokens` | [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py) | Hard ceiling on input + output tokens | Model-specific (e.g., 4k-128k) |
| `compaction_threshold_pct` | [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py) | Triggers context summarization | 0.7-0.9 (default 0.85) |
| `persistent` | [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py) | Reuses REPL environment across calls | `True`/`False` |

## Common Performance Bottlenecks

- **Socket round-trip overhead**: Every `llm_query` and `rlm_query` incurs length-prefixed JSON exchange overhead. For low-latency local models, this can dominate per-call costs; batching via `_rlm_query_batched` mitigates this.
- **Thread-pool size vs. API limits**: High `max_concurrent_subcalls` values can exhaust rate limits on hosted providers. The framework limits parallelism but requires conservative tuning for specific provider tiers.
- **Compaction latency**: Each summarization step issues an extra LM call. While cheap for large contexts, frequent compaction on small contexts adds unnecessary latency.
- **Memory growth in persistent mode**: Storing every context and history version in `_persistent_env` grows in-process memory, potentially causing GC pressure in long-running applications.

## Code Examples

### Tuning Concurrency for Child RLMs

```python
from rlm import RLM

# Use a local model for low latency; allow up to 6 parallel child RLMs

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    environment="local",
    max_concurrent_subcalls=6,
    max_iterations=20,
    max_tokens=8000,
    compaction=True,
    compaction_threshold_pct=0.80,
)

result = rlm.completion(
    """Write a short Python function that computes the nth Fibonacci number
    and then call it for n=10."""
)
print(result.response)

```

### Batched LM Queries to Reduce Socket Overhead

```python
from rlm import RLM

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-3.5-turbo"},
    environment="local",
)

prompts = [
    "Explain the concept of memoization in 2 sentences.",
    "Give a one‑line Python list‑comprehension that squares numbers 0‑9.",
    "Summarise the differences between TCP and UDP."
]

# Directly use the environment's low‑level API for batched queries

env = rlm._spawn_completion_context("dummy")[1]   # internal – illustrative only

responses = env._llm_query_batched(prompts)
print(responses)

```

### Persistent REPL for Multi-Turn Interaction

```python
from rlm import RLM

# Persistent mode keeps the REPL alive across calls

rlm = RLM(persistent=True)

# First turn – load a data set

rlm.completion("Load the CSV file 'data.csv' as a pandas DataFrame named df.")

# Second turn – perform analysis without re‑loading the file

analysis = rlm.completion(
    """Calculate the mean of the column 'price' and store it in a variable `avg_price`.
    Then print `avg_price`."""
)
print(analysis.response)

```

## Summary

- RLM manages performance through **six key parameters** controlling concurrency, tokens, and recursion depth.
- **ThreadPoolExecutor** in [`local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/local_repl.py) and **asyncio semaphores** in [`lm_handler.py`](https://github.com/alexzhang13/rlm/blob/main/lm_handler.py) provide parallelism while preventing resource exhaustion.
- **Context compaction** via `_compact_history` keeps token usage within model limits but adds computational overhead proportional to summarization frequency.
- **Persistent environments** eliminate setup costs across multiple calls but require memory monitoring to prevent GC pressure.
- **Socket communication** overhead favors batched operations (`_rlm_query_batched`) over individual calls.

## Frequently Asked Questions

### How does RLM prevent infinite recursive loops?

RLM enforces hard limits through `max_depth` and `max_iterations` initialized in `RLM.__init__` (lines 55-58) and monitored by `_check_iteration_limits` (lines 66-78). These parameters cap both nested child RLMs and REPL loop iterations, aborting execution immediately when exceeded.

### What causes high latency in RLM execution?

High latency typically stems from **socket round-trip overhead** for each query, **compaction costs** when summarizing context via `_compact_history`, or excessive **thread-pool contention** against API rate limits. Using `persistent=True` and tuning `compaction_threshold_pct` can reduce repeated setup costs.

### How does the `max_concurrent_subcalls` parameter affect performance?

This parameter controls the `ThreadPoolExecutor` size in [`local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/local_repl.py) (lines 66-71), limiting how many child RLMs run in parallel. Values above 4-8 increase throughput but raise CPU and memory pressure, and may trigger rate limits on hosted LLM providers.

### When should I enable persistent mode?

Enable `persistent=True` when making multiple sequential calls to the same RLM instance, as it reuses the REPL process and socket connection via `_spawn_completion_context` (lines 58-71). This eliminates initialization overhead but requires monitoring for memory growth in long-running applications.