How to Integrate RLM with Other Systems: A Complete Developer Guide
Integrate RLM with other systems by instantiating the RLM class from rlm/core/rlm.py, configuring your backend via LMHandler, and calling completion() to execute recursive reasoning workflows with custom tools, callbacks, and persistence options.
The alexzhang13/rlm repository provides a thin orchestration layer for recursive language model execution that you can embed into web services, data pipelines, or CLI tools. Understanding how to integrate RLM with other systems requires familiarity with its three-component architecture and the specific integration points exposed in the core modules.
Understanding RLM's Architecture
RLM operates as a coordination layer between three distinct components that handle different aspects of recursive execution.
LM Handler (rlm/core/lm_handler.py)
The LM Handler is a multi-threaded TCP server that forwards language model requests to any registered client backend. Located in rlm/core/lm_handler.py, this component supports OpenAI, Anthropic, Gemini, Azure, and other providers through a unified interface. It returns structured LMResponse objects that the core orchestrator uses to track usage and manage the execution loop.
Environment (rlm/environments/local_repl.py)
The Environment provides a sandboxed execution context for code generated by the language model. The default implementation is LocalREPL in rlm/environments/local_repl.py, which offers a local Python REPL with persistent namespace and injected query functions (llm_query, rlm_query). Alternative environments include IPython, Modal (via modal_repl.py), Prime, and Docker, each implementing the SupportsPersistence interface where applicable.
RLM Core (rlm/core/rlm.py)
The RLM class in rlm/core/rlm.py serves as the high-level orchestration API. It manages recursion depth, iteration loops, budgeting, timeouts, compaction, callbacks, and persistence. When you integrate RLM with other systems, you interact primarily with this class and its completion() method.
Integration Points and Methods
Basic Client Integration
To embed RLM in any Python system, instantiate the RLM class and call completion(prompt):
from rlm.core.rlm import RLM
rlm = RLM(
backend="openai",
backend_kwargs={"model_name": "gpt-4"},
max_iterations=5,
max_depth=2,
)
result = rlm.completion("Write a Python function that returns the nth Fibonacci number.")
print(result.response)
The RLM.__init__ method creates a fresh LM client via get_client, initializes a LMHandler instance, and spawns an environment through get_environment. The completion method runs an iterative REPL loop, feeding the model with system prompts, user prompts, and execution results from each code block.
Custom Tools Injection
Pass a dictionary of callables to the custom_tools parameter to make external functions available inside the sandbox:
def fetch_user(user_id: int) -> dict:
return {"id": user_id, "name": "Alice", "balance": 42.0}
rlm = RLM(
backend="openai",
backend_kwargs={"model_name": "gpt-4"},
custom_tools={"fetch_user": fetch_user},
max_iterations=4,
)
prompt = "Use the provided `fetch_user` tool to get the balance of user 7."
result = rlm.completion(prompt)
The LocalREPL.setup method injects each callable into the sandbox globals, making them callable from model-generated code. The environment validates that custom tool names do not overwrite reserved names (llm_query, rlm_query, etc.) via validate_custom_tools.
Recursive Sub-calls
When the model invokes rlm_query, the system supports recursive execution through the subcall_fn mechanism. LocalREPL._rlm_query forwards the request to the parent RLM._subcall, which creates a new RLM instance with increased depth. If max_depth is reached, the call falls back to a plain LM query via LMHandler.
Persistence Across Calls
Set persistent=True when creating an RLM instance to maintain state across multiple completion() calls:
rlm = RLM(
backend="anthropic",
backend_kwargs={"model_name": "claude-2"},
persistent=True,
)
The root RLM reuses the same Environment across calls, updating the handler address and adding new context via environment.add_context. This requires the environment to implement SupportsPersistence, which both LocalREPL and IPythonREPL support.
Resource Limits and Timeouts
Control computational costs through budget, timeout, and token parameters:
max_budget: Tracks cumulative cost via_cumulative_costandLMHandler.get_usage_summarymax_timeout: Monitors execution duration per iterationmax_tokens: Enforces token limit boundaries
The RLM object raises BudgetExceededError, TimeoutExceededError, or TokenLimitExceededError if limits are crossed after any iteration.
Callback Hooks for Monitoring
Supply callback functions to hook into the execution lifecycle for UI updates, logging, or external monitoring:
def on_iter_start(depth, it):
print(f"[Depth {depth}] Starting iteration {it}")
def on_iter_end(depth, it, dur):
print(f"[Depth {depth}] Finished iteration {it} in {dur:.2f}s")
rlm = RLM(
backend="openai",
on_iteration_start=on_iter_start,
on_iteration_complete=on_iter_end,
)
Available callbacks include on_subcall_start, on_subcall_complete, on_iteration_start, and on_iteration_complete.
Integration Examples
Standalone Script Integration
For simple command-line tools or scripts, create a disposable RLM instance:
from rlm.core.rlm import RLM
rlm = RLM(
backend="openai",
backend_kwargs={"model_name": "gpt-4"},
max_iterations=5,
max_depth=2,
)
prompt = "Write a Python function that returns the nth Fibonacci number."
result = rlm.completion(prompt)
print("Answer:", result.response)
print("Usage:", result.usage_summary.total_cost, "USD")
Web API Integration with Flask
Embed RLM in a web service to expose recursive reasoning via HTTP endpoints:
from flask import Flask, request, jsonify
from rlm.core.rlm import RLM
app = Flask(__name__)
rlm = RLM(
backend="anthropic",
backend_kwargs={"model_name": "claude-2"},
max_iterations=6,
max_depth=2,
persistent=True,
)
@app.post("/ask")
def ask():
data = request.get_json()
prompt = data.get("prompt", "")
result = rlm.completion(prompt)
return jsonify({
"answer": result.response,
"usage": result.usage_summary.to_dict(),
})
if __name__ == "__main__":
app.run(port=5000)
Because persistent=True, the same REPL environment is reused for follow-up calls, allowing the model to maintain state between requests.
Database Integration with Custom Tools
Connect RLM to external databases or APIs by injecting custom functions:
def fetch_user(user_id: int) -> dict:
# Query your database here
return {"id": user_id, "name": "Alice", "balance": 42.0}
rlm = RLM(
backend="openai",
backend_kwargs={"model_name": "gpt-4"},
custom_tools={"fetch_user": fetch_user},
max_iterations=4,
)
prompt = """
Use the provided `fetch_user` tool to get the balance of user 7 and
return a sentence like "User 7 has a balance of $42.00".
"""
result = rlm.completion(prompt)
print(result.response)
Progress Tracking with Callbacks
For long-running tasks, implement callbacks to report progress to external monitoring systems:
def on_iter_start(depth, it):
print(f"[Depth {depth}] Starting iteration {it}")
def on_iter_end(depth, it, dur):
print(f"[Depth {depth}] Finished iteration {it} in {dur:.2f}s")
rlm = RLM(
backend="openai",
backend_kwargs={"model_name": "gpt-4"},
on_iteration_start=on_iter_start,
on_iteration_complete=on_iter_end,
)
prompt = "Plan and execute a simple data-processing pipeline."
result = rlm.completion(prompt)
print("Final answer:", result.response)
Key Source Files
Understanding these files helps when debugging or extending integrations:
rlm/core/rlm.py: Main orchestration class (RLM) with thecompletion()method and recursion logicrlm/core/lm_handler.py: Socket-based LM request router that registers clients fromrlm/clients/__init__.pyrlm/environments/local_repl.py: Default non-isolated Python REPL implementation with tool injectionrlm/environments/modal_repl.py: Example of an isolated environment using HTTP brokers for remote executionrlm/environments/base_env.py: Abstract base defining reserved names and environment interfacesrlm/clients/: Provider-specific implementations (OpenAI, Anthropic, Gemini, Azure, Portkey)
Summary
- Instantiate
RLMfromrlm/core/rlm.pyto begin integration, configuring the backend throughbackendandbackend_kwargs - Use
custom_toolsto inject Python functions into the sandboxed environment, making external APIs and databases callable from model-generated code - Enable
persistent=Truefor stateful conversations across multiplecompletion()calls, requiring the environment to supportSupportsPersistence - Set resource limits (
max_budget,max_timeout,max_tokens) to prevent runaway execution and control costs - Implement callbacks (
on_iteration_start,on_iteration_complete, etc.) to integrate with monitoring systems, progress bars, or logging frameworks - Swap components independently: change the LM backend via
backendparameter, replace the environment viaenvironmentparameter, or extend tools without modifying core logic
Frequently Asked Questions
How do I add custom functions to RLM's execution environment?
Pass a dictionary of callables to the custom_tools parameter when creating an RLM instance. The LocalREPL.setup method in rlm/environments/local_repl.py injects these functions into the sandbox globals, making them available to model-generated code. The system validates that your tool names do not conflict with reserved names like llm_query or rlm_query.
Can I use RLM in a production web application?
Yes. Instantiate RLM with persistent=True to maintain state across HTTP requests, or create per-request instances for isolated execution. The Flask example in rlm/core/rlm.py documentation shows how to expose the completion() method via REST endpoints, returning RLMChatCompletion objects containing responses and usage summaries.
How does RLM handle resource limits and prevent infinite loops?
The RLM class tracks cumulative cost through _cumulative_cost and LMHandler.get_usage_summary, enforcing max_budget, max_timeout, and max_tokens limits. After each iteration, it raises BudgetExceededError, TimeoutExceededError, or TokenLimitExceededError if thresholds are crossed. The max_iterations and max_depth parameters provide additional guardrails against unbounded recursion.
What is the difference between LocalREPL and Modal environments?
LocalREPL (in rlm/environments/local_repl.py) executes code in the same Python process, providing low latency and easy debugging but limited isolation. ModalREPL (in rlm/environments/modal_repl.py) demonstrates an isolated environment that communicates via HTTP brokers, suitable for untrusted code execution or distributed deployments. Both implement the same environment interface, allowing seamless swapping through the environment parameter.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →