# How to Integrate RLM with Other Systems: A Complete Developer Guide

> Learn how to integrate RLM with other systems. This guide shows you how to instantiate the RLM class, configure your backend, and execute recursive reasoning workflows.

- Repository: [az/rlm](https://github.com/alexzhang13/rlm)
- Tags: how-to-guide
- Published: 2026-06-18

---

**Integrate RLM with other systems by instantiating the `RLM` class from [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py), configuring your backend via `LMHandler`, and calling `completion()` to execute recursive reasoning workflows with custom tools, callbacks, and persistence options.**

The alexzhang13/rlm repository provides a thin orchestration layer for recursive language model execution that you can embed into web services, data pipelines, or CLI tools. Understanding how to integrate RLM with other systems requires familiarity with its three-component architecture and the specific integration points exposed in the core modules.

## Understanding RLM's Architecture

RLM operates as a coordination layer between three distinct components that handle different aspects of recursive execution.

### LM Handler ([`rlm/core/lm_handler.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/lm_handler.py))

The **LM Handler** is a multi-threaded TCP server that forwards language model requests to any registered client backend. Located in [`rlm/core/lm_handler.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/lm_handler.py), this component supports OpenAI, Anthropic, Gemini, Azure, and other providers through a unified interface. It returns structured **`LMResponse`** objects that the core orchestrator uses to track usage and manage the execution loop.

### Environment ([`rlm/environments/local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/local_repl.py))

The **Environment** provides a sandboxed execution context for code generated by the language model. The default implementation is **`LocalREPL`** in [`rlm/environments/local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/local_repl.py), which offers a local Python REPL with persistent namespace and injected query functions (`llm_query`, `rlm_query`). Alternative environments include IPython, Modal (via [`modal_repl.py`](https://github.com/alexzhang13/rlm/blob/main/modal_repl.py)), Prime, and Docker, each implementing the `SupportsPersistence` interface where applicable.

### RLM Core ([`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py))

The **`RLM`** class in [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py) serves as the high-level orchestration API. It manages recursion depth, iteration loops, budgeting, timeouts, compaction, callbacks, and persistence. When you integrate RLM with other systems, you interact primarily with this class and its `completion()` method.

## Integration Points and Methods

### Basic Client Integration

To embed RLM in any Python system, instantiate the `RLM` class and call `completion(prompt)`:

```python
from rlm.core.rlm import RLM

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    max_iterations=5,
    max_depth=2,
)

result = rlm.completion("Write a Python function that returns the nth Fibonacci number.")
print(result.response)

```

The `RLM.__init__` method creates a fresh LM client via `get_client`, initializes a `LMHandler` instance, and spawns an environment through `get_environment`. The `completion` method runs an iterative REPL loop, feeding the model with system prompts, user prompts, and execution results from each code block.

### Custom Tools Injection

Pass a dictionary of callables to the `custom_tools` parameter to make external functions available inside the sandbox:

```python
def fetch_user(user_id: int) -> dict:
    return {"id": user_id, "name": "Alice", "balance": 42.0}

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    custom_tools={"fetch_user": fetch_user},
    max_iterations=4,
)

prompt = "Use the provided `fetch_user` tool to get the balance of user 7."
result = rlm.completion(prompt)

```

The `LocalREPL.setup` method injects each callable into the sandbox globals, making them callable from model-generated code. The environment validates that custom tool names do not overwrite reserved names (`llm_query`, `rlm_query`, etc.) via `validate_custom_tools`.

### Recursive Sub-calls

When the model invokes `rlm_query`, the system supports recursive execution through the `subcall_fn` mechanism. `LocalREPL._rlm_query` forwards the request to the parent `RLM._subcall`, which creates a new `RLM` instance with increased `depth`. If `max_depth` is reached, the call falls back to a plain LM query via `LMHandler`.

### Persistence Across Calls

Set `persistent=True` when creating an `RLM` instance to maintain state across multiple `completion()` calls:

```python
rlm = RLM(
    backend="anthropic",
    backend_kwargs={"model_name": "claude-2"},
    persistent=True,
)

```

The root `RLM` reuses the same `Environment` across calls, updating the handler address and adding new context via `environment.add_context`. This requires the environment to implement `SupportsPersistence`, which both `LocalREPL` and `IPythonREPL` support.

### Resource Limits and Timeouts

Control computational costs through budget, timeout, and token parameters:

- **`max_budget`**: Tracks cumulative cost via `_cumulative_cost` and `LMHandler.get_usage_summary`
- **`max_timeout`**: Monitors execution duration per iteration
- **`max_tokens`**: Enforces token limit boundaries

The `RLM` object raises `BudgetExceededError`, `TimeoutExceededError`, or `TokenLimitExceededError` if limits are crossed after any iteration.

### Callback Hooks for Monitoring

Supply callback functions to hook into the execution lifecycle for UI updates, logging, or external monitoring:

```python
def on_iter_start(depth, it):
    print(f"[Depth {depth}] Starting iteration {it}")

def on_iter_end(depth, it, dur):
    print(f"[Depth {depth}] Finished iteration {it} in {dur:.2f}s")

rlm = RLM(
    backend="openai",
    on_iteration_start=on_iter_start,
    on_iteration_complete=on_iter_end,
)

```

Available callbacks include `on_subcall_start`, `on_subcall_complete`, `on_iteration_start`, and `on_iteration_complete`.

## Integration Examples

### Standalone Script Integration

For simple command-line tools or scripts, create a disposable `RLM` instance:

```python
from rlm.core.rlm import RLM

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    max_iterations=5,
    max_depth=2,
)

prompt = "Write a Python function that returns the nth Fibonacci number."
result = rlm.completion(prompt)

print("Answer:", result.response)
print("Usage:", result.usage_summary.total_cost, "USD")

```

### Web API Integration with Flask

Embed RLM in a web service to expose recursive reasoning via HTTP endpoints:

```python
from flask import Flask, request, jsonify
from rlm.core.rlm import RLM

app = Flask(__name__)

rlm = RLM(
    backend="anthropic",
    backend_kwargs={"model_name": "claude-2"},
    max_iterations=6,
    max_depth=2,
    persistent=True,
)

@app.post("/ask")
def ask():
    data = request.get_json()
    prompt = data.get("prompt", "")
    result = rlm.completion(prompt)
    return jsonify({
        "answer": result.response,
        "usage": result.usage_summary.to_dict(),
    })

if __name__ == "__main__":
    app.run(port=5000)

```

Because `persistent=True`, the same REPL environment is reused for follow-up calls, allowing the model to maintain state between requests.

### Database Integration with Custom Tools

Connect RLM to external databases or APIs by injecting custom functions:

```python
def fetch_user(user_id: int) -> dict:
    # Query your database here

    return {"id": user_id, "name": "Alice", "balance": 42.0}

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    custom_tools={"fetch_user": fetch_user},
    max_iterations=4,
)

prompt = """
Use the provided `fetch_user` tool to get the balance of user 7 and
return a sentence like "User 7 has a balance of $42.00".
"""
result = rlm.completion(prompt)
print(result.response)

```

### Progress Tracking with Callbacks

For long-running tasks, implement callbacks to report progress to external monitoring systems:

```python
def on_iter_start(depth, it):
    print(f"[Depth {depth}] Starting iteration {it}")

def on_iter_end(depth, it, dur):
    print(f"[Depth {depth}] Finished iteration {it} in {dur:.2f}s")

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-4"},
    on_iteration_start=on_iter_start,
    on_iteration_complete=on_iter_end,
)

prompt = "Plan and execute a simple data-processing pipeline."
result = rlm.completion(prompt)
print("Final answer:", result.response)

```

## Key Source Files

Understanding these files helps when debugging or extending integrations:

- **[`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py)**: Main orchestration class (`RLM`) with the `completion()` method and recursion logic
- **[`rlm/core/lm_handler.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/lm_handler.py)**: Socket-based LM request router that registers clients from [`rlm/clients/__init__.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/clients/__init__.py)
- **[`rlm/environments/local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/local_repl.py)**: Default non-isolated Python REPL implementation with tool injection
- **[`rlm/environments/modal_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/modal_repl.py)**: Example of an isolated environment using HTTP brokers for remote execution
- **[`rlm/environments/base_env.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/base_env.py)**: Abstract base defining reserved names and environment interfaces
- **`rlm/clients/`**: Provider-specific implementations (OpenAI, Anthropic, Gemini, Azure, Portkey)

## Summary

- **Instantiate `RLM`** from [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py) to begin integration, configuring the backend through `backend` and `backend_kwargs`
- **Use `custom_tools`** to inject Python functions into the sandboxed environment, making external APIs and databases callable from model-generated code
- **Enable `persistent=True`** for stateful conversations across multiple `completion()` calls, requiring the environment to support `SupportsPersistence`
- **Set resource limits** (`max_budget`, `max_timeout`, `max_tokens`) to prevent runaway execution and control costs
- **Implement callbacks** (`on_iteration_start`, `on_iteration_complete`, etc.) to integrate with monitoring systems, progress bars, or logging frameworks
- **Swap components** independently: change the LM backend via `backend` parameter, replace the environment via `environment` parameter, or extend tools without modifying core logic

## Frequently Asked Questions

### How do I add custom functions to RLM's execution environment?

Pass a dictionary of callables to the `custom_tools` parameter when creating an `RLM` instance. The `LocalREPL.setup` method in [`rlm/environments/local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/local_repl.py) injects these functions into the sandbox globals, making them available to model-generated code. The system validates that your tool names do not conflict with reserved names like `llm_query` or `rlm_query`.

### Can I use RLM in a production web application?

Yes. Instantiate `RLM` with `persistent=True` to maintain state across HTTP requests, or create per-request instances for isolated execution. The Flask example in [`rlm/core/rlm.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/core/rlm.py) documentation shows how to expose the `completion()` method via REST endpoints, returning `RLMChatCompletion` objects containing responses and usage summaries.

### How does RLM handle resource limits and prevent infinite loops?

The `RLM` class tracks cumulative cost through `_cumulative_cost` and `LMHandler.get_usage_summary`, enforcing `max_budget`, `max_timeout`, and `max_tokens` limits. After each iteration, it raises `BudgetExceededError`, `TimeoutExceededError`, or `TokenLimitExceededError` if thresholds are crossed. The `max_iterations` and `max_depth` parameters provide additional guardrails against unbounded recursion.

### What is the difference between LocalREPL and Modal environments?

`LocalREPL` (in [`rlm/environments/local_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/local_repl.py)) executes code in the same Python process, providing low latency and easy debugging but limited isolation. `ModalREPL` (in [`rlm/environments/modal_repl.py`](https://github.com/alexzhang13/rlm/blob/main/rlm/environments/modal_repl.py)) demonstrates an isolated environment that communicates via HTTP brokers, suitable for untrusted code execution or distributed deployments. Both implement the same environment interface, allowing seamless swapping through the `environment` parameter.