How to Integrate Custom Agentic Environments with Miles: A Complete Guide
You integrate custom agentic environments with Miles by implementing an async generate function that conforms to the GenerateFnInput → GenerateFnOutput contract, then registering it via --custom-generate-function-path or per-sample through dataset configuration.
Miles, an open-source reinforcement learning framework for language models, provides a flexible pluggable generation backend that lets you swap its default SGLang generator for your own custom logic. This integration point lives at miles/rollout/sglang_rollout.py and enables you to connect external APIs, simulation environments, tool-use loops, or proprietary models while retaining Miles' reward modeling and training infrastructure.
How the Custom Generate Function Works
Miles' rollout engine discovers custom functions through two channels:
- Global CLI flag:
--custom-generate-function-path my_module.my_function - Per-sample override:
sample.generate_function_path(typically set in dataset configs)
When either path is present, Miles calls load_generate_function from miles/utils/function_registry.py to dynamically import your callable. The function then executes inside the async generate_and_rm loop, receiving full context about the current generation state and returning structured results.
The Core Data Flow
The integration follows this contract as defined in miles/rollout/base_types.py:
from dataclasses import dataclass
from typing import List
from miles.utils.types import Sample
@dataclass
class GenerateFnInput:
state: "GlobalGenerationState" # Shared state across all samples
sample: Sample # The sample being processed
sampling_params: "SamplingParams" # Temperature, top-p, etc.
evaluation: bool # Whether this is eval or training
@dataclass
class GenerateFnOutput:
samples: List[Sample] # Updated samples (usually length 1)
Your function receives GenerateFnInput, performs arbitrary computation, populates sample.response (and optionally sample.reward), then returns GenerateFnOutput. Miles handles log-probability computation, token tracking, and downstream RM scoring automatically.
Step-by-Step Implementation
Step 1: Create Your Custom Generate Function
Define an async function in a Python module. This example shows a minimal echo environment:
# file: my_agentic_env.py
from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput
async def my_generate_fn(inp: GenerateFnInput) -> GenerateFnOutput:
"""
Custom agentic environment for Miles.
Parameters
----------
inp : GenerateFnInput
Contains the sample, global state, and generation parameters.
Returns
-------
GenerateFnOutput
The sample(s) with populated response field.
"""
sample = inp.sample
# ----- Agentic logic: replace with your environment -----
# Examples:
# - Call OpenAI/Anthropic API with tool use
# - Execute in a Python sandbox (e.g., Jupyter kernel)
# - Query a simulation environment (Minecraft, WebArena, etc.)
# - Route to a local vLLM instance with custom sampling
sample.response = f"[AGENT EXECUTION] {sample.prompt}"
sample.reward = 0.5 # Optional: provide intermediate reward
# --------------------------------------------------------
return GenerateFnOutput(samples=[sample])
Key requirements:
- Use
async def— the rollout loop is fully asynchronous - Modify
sample.response— this field becomes the model's output - Return
GenerateFnOutputwith a list of samples (usually one)
Step 2: Register via Command Line
Pass your function's fully-qualified import path to any Miles rollout script:
python -m miles.main.scripts.run_qwen3_dense \
--custom-generate-function-path my_agentic_env.my_generate_fn \
--model-path Qwen/Qwen3-8B \
--output-dir ./outputs
Miles executes load_generate_function("my_agentic_env.my_generate_fn"), which imports my_agentic_env and extracts the callable.
Step 3: Register Per-Sample via Dataset Config
For evaluation datasets where different samples need different environments, use TOML configuration:
# file: webarena_eval.toml
[dataset]
path = "data/webarena_tasks.jsonl"
input_key = "prompt"
label_key = "gold_answer"
# Custom generate function applies to all samples in this dataset
custom_generate_function_path = "environments.webarena_agent.generate"
[[dataset]]
path = "data/code_execution_tasks.jsonl"
input_key = "coding_problem"
label_key = "test_cases"
custom_generate_function_path = "environments.sandbox_executor.generate"
When EvalDatasetConfig loads this file (in miles/utils/eval_config.py), it stores custom_generate_function_path on each Sample instance. The rollout code at miles/rollout/sglang_rollout.py (lines 317-322) prefers this per-sample attribute:
# From miles/rollout/sglang_rollout.py
custom_func_path = (
getattr(sample, "generate_function_path", None)
or args.custom_generate_function_path
)
generate_fn = load_generate_function(custom_func_path) if custom_func_path else None
if generate_fn is not None:
output = await generate_fn(
GenerateFnInput(
state=state,
sample=sample,
sampling_params=sampling_params,
evaluation=evaluation
)
)
samples = output.samples
else:
# Fall back to default SGLang generation
samples = await generate(args, sample, sampling_params)
This branching logic ensures your custom environment integrates seamlessly with Miles' existing sampling and reward infrastructure.
Advanced Integration Patterns
Pattern 1: Multi-Turn Agent Loops
For environments requiring multiple generation steps (e.g., ReAct-style tool use), maintain state in inp.state:
# file: react_agent.py
from typing import List
from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput
async def react_loop(inp: GenerateFnInput) -> GenerateFnOutput:
"""
ReAct-style agent with tool use.
Uses inp.state to track conversation history across calls.
"""
sample = inp.sample
max_turns = 10
# Access or initialize per-sample state
if not hasattr(inp.state, "react_histories"):
inp.state.react_histories = {}
history: List[dict] = inp.state.react_histories.get(sample.id, [])
history.append({"role": "user", "content": sample.prompt})
for turn in range(max_turns):
# Call your underlying model (external API, local LLM, etc.)
response = await call_agent_model(history)
if "<tool>" in response:
# Execute tool and continue loop
tool_result = await execute_tool(extract_tool_call(response))
history.append({"role": "assistant", "content": response})
history.append({"role": "tool", "content": tool_result})
else:
# Final answer
break
sample.response = response
sample.reward = score_react_trajectory(history) # Custom reward shaping
# Persist history for potential future use
inp.state.react_histories[sample.id] = history
return GenerateFnOutput(samples=[sample])
Pattern 2: External API Integration
Connect to proprietary models or services:
# file: external_api_bridge.py
import aiohttp
from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput
API_ENDPOINT = "https://api.proprietary-model.com/v1/generate"
async def external_api_generate(inp: GenerateFnInput) -> GenerateFnOutput:
"""Bridge to external generation API."""
sample = inp.sample
async with aiohttp.ClientSession() as session:
payload = {
"prompt": sample.prompt,
"temperature": inp.sampling_params.temperature,
"max_tokens": inp.sampling_params.max_new_tokens,
}
async with session.post(API_ENDPOINT, json=payload) as resp:
result = await resp.json()
sample.response = result["text"]
# External APIs may not provide log-probs; Miles handles missing fields gracefully
return GenerateFnOutput(samples=[sample])
Pattern 3: Deterministic Test Environments
For unit testing or reproducible baselines:
# file: deterministic_env.py
from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput
async def fixed_response_generator(inp: GenerateFnInput) -> GenerateFnOutput:
"""Returns canned responses for regression testing."""
sample = inp.sample
# Lookup based on prompt hash or ID
canned_responses = {
"math_problem_001": "The answer is 42.",
"code_task_042": "def solve():\n return 0",
}
sample.response = canned_responses.get(sample.sample_id, "Unknown")
sample.reward = 1.0 if sample.response != "Unknown" else 0.0
return GenerateFnOutput(samples=[sample])
Key Repository Files for Custom Integration
| File | Purpose | Critical Lines |
|---|---|---|
miles/rollout/sglang_rollout.py |
Orchestrates generation; loads and invokes custom functions | 317-322: branching logic for custom vs. default generation |
miles/utils/function_registry.py |
load_generate_function() — dynamic import mechanism |
Entire file: ~80 lines of import utilities |
miles/rollout/base_types.py |
GenerateFnInput, GenerateFnOutput dataclasses |
Class definitions and field specifications |
miles/utils/types.py |
Sample model — the object your function manipulates |
Response, reward, and metadata fields |
miles/utils/eval_config.py |
EvalDatasetConfig — parses custom_generate_function_path from TOML |
Dataset loading and attribute propagation |
tests/fast/rollout/generate_hub/test_multi_turn.py |
Reference tests for custom generate paths | Example fixtures and assertions |
Common Pitfalls and Solutions
- Synchronous blocking: Always use
async def. If your underlying library is synchronous, wrap it withasyncio.to_thread()orloop.run_in_executor(). - Missing
responsefield: Miles requiressample.responseto be set. Unset fields cause downstream failures in reward model scoring. - Import path errors: Ensure your module is in
PYTHONPATH. Usepython -c "import my_module; print(my_module.my_fn)"to verify before running Miles. - State isolation:
inp.stateis global across samples in a batch. Usesample.idas a key to maintain per-sample isolation.
Summary
- Implement an async function with signature
GenerateFnInput → GenerateFnOutputto create a custom agentic environment - Register your function globally via
--custom-generate-function-pathor per-sample viacustom_generate_function_pathin dataset configs - Manipulate the
Sampleobject to setresponseand optionallyreward; Miles handles metadata attachment - Leverage
inp.statefor cross-turn state in multi-step agents, andinp.sampling_paramsfor generation controls - Consult
miles/rollout/sglang_rollout.pylines 317-322 for the exact integration point in the codebase
Frequently Asked Questions
What happens if both the CLI flag and per-sample path are set?
The per-sample generate_function_path takes precedence. This allows you to set a default global environment while overriding specific evaluation datasets with specialized ones. The resolution logic in miles/rollout/sglang_rollout.py explicitly checks the sample attribute first, falling back to args.custom_generate_function_path only when absent.
Can my custom function access GPU resources or distributed state?
Yes. The GenerateFnInput.state parameter contains GlobalGenerationState, which holds references to the model, tokenizer, and distributed training context when available. However, most agentic environments using external APIs or simulators will ignore these fields and manage their own compute.
How do I debug a custom generate function?
Set PYTHONPATH to include your module directory, then test the import chain: python -c "from miles.utils.function_registry import load_generate_function; fn = load_generate_function('my_module.my_fn'); print(fn)". For runtime debugging, add logging inside your async function—output appears in Miles' standard logs. The test file tests/fast/rollout/generate_hub/test_multi_turn.py provides a minimal harness for isolated testing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →