How to Integrate Custom Agentic Environments with Miles: A Complete Guide

You integrate custom agentic environments with Miles by implementing an async generate function that conforms to the GenerateFnInput → GenerateFnOutput contract, then registering it via --custom-generate-function-path or per-sample through dataset configuration.

Miles, an open-source reinforcement learning framework for language models, provides a flexible pluggable generation backend that lets you swap its default SGLang generator for your own custom logic. This integration point lives at miles/rollout/sglang_rollout.py and enables you to connect external APIs, simulation environments, tool-use loops, or proprietary models while retaining Miles' reward modeling and training infrastructure.

How the Custom Generate Function Works

Miles' rollout engine discovers custom functions through two channels:

  1. Global CLI flag: --custom-generate-function-path my_module.my_function
  2. Per-sample override: sample.generate_function_path (typically set in dataset configs)

When either path is present, Miles calls load_generate_function from miles/utils/function_registry.py to dynamically import your callable. The function then executes inside the async generate_and_rm loop, receiving full context about the current generation state and returning structured results.

The Core Data Flow

The integration follows this contract as defined in miles/rollout/base_types.py:

from dataclasses import dataclass
from typing import List
from miles.utils.types import Sample

@dataclass
class GenerateFnInput:
    state: "GlobalGenerationState"  # Shared state across all samples

    sample: Sample                   # The sample being processed

    sampling_params: "SamplingParams"  # Temperature, top-p, etc.

    evaluation: bool                 # Whether this is eval or training

@dataclass
class GenerateFnOutput:
    samples: List[Sample]            # Updated samples (usually length 1)

Your function receives GenerateFnInput, performs arbitrary computation, populates sample.response (and optionally sample.reward), then returns GenerateFnOutput. Miles handles log-probability computation, token tracking, and downstream RM scoring automatically.

Step-by-Step Implementation

Step 1: Create Your Custom Generate Function

Define an async function in a Python module. This example shows a minimal echo environment:


# file: my_agentic_env.py

from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput

async def my_generate_fn(inp: GenerateFnInput) -> GenerateFnOutput:
    """
    Custom agentic environment for Miles.
    
    Parameters
    ----------
    inp : GenerateFnInput
        Contains the sample, global state, and generation parameters.
    
    Returns
    -------
    GenerateFnOutput
        The sample(s) with populated response field.
    """
    sample = inp.sample
    
    # ----- Agentic logic: replace with your environment -----

    # Examples:

    # - Call OpenAI/Anthropic API with tool use

    # - Execute in a Python sandbox (e.g., Jupyter kernel)

    # - Query a simulation environment (Minecraft, WebArena, etc.)

    # - Route to a local vLLM instance with custom sampling

    
    sample.response = f"[AGENT EXECUTION] {sample.prompt}"
    sample.reward = 0.5  # Optional: provide intermediate reward

    # --------------------------------------------------------

    
    return GenerateFnOutput(samples=[sample])

Key requirements:

  • Use async def — the rollout loop is fully asynchronous
  • Modify sample.response — this field becomes the model's output
  • Return GenerateFnOutput with a list of samples (usually one)

Step 2: Register via Command Line

Pass your function's fully-qualified import path to any Miles rollout script:

python -m miles.main.scripts.run_qwen3_dense \
    --custom-generate-function-path my_agentic_env.my_generate_fn \
    --model-path Qwen/Qwen3-8B \
    --output-dir ./outputs

Miles executes load_generate_function("my_agentic_env.my_generate_fn"), which imports my_agentic_env and extracts the callable.

Step 3: Register Per-Sample via Dataset Config

For evaluation datasets where different samples need different environments, use TOML configuration:


# file: webarena_eval.toml

[dataset]
path = "data/webarena_tasks.jsonl"
input_key = "prompt"
label_key = "gold_answer"

# Custom generate function applies to all samples in this dataset

custom_generate_function_path = "environments.webarena_agent.generate"

[[dataset]]
path = "data/code_execution_tasks.jsonl"
input_key = "coding_problem"
label_key = "test_cases"
custom_generate_function_path = "environments.sandbox_executor.generate"

When EvalDatasetConfig loads this file (in miles/utils/eval_config.py), it stores custom_generate_function_path on each Sample instance. The rollout code at miles/rollout/sglang_rollout.py (lines 317-322) prefers this per-sample attribute:


# From miles/rollout/sglang_rollout.py

custom_func_path = (
    getattr(sample, "generate_function_path", None) 
    or args.custom_generate_function_path
)
generate_fn = load_generate_function(custom_func_path) if custom_func_path else None

if generate_fn is not None:
    output = await generate_fn(
        GenerateFnInput(
            state=state,
            sample=sample,
            sampling_params=sampling_params,
            evaluation=evaluation
        )
    )
    samples = output.samples
else:
    # Fall back to default SGLang generation

    samples = await generate(args, sample, sampling_params)

This branching logic ensures your custom environment integrates seamlessly with Miles' existing sampling and reward infrastructure.

Advanced Integration Patterns

Pattern 1: Multi-Turn Agent Loops

For environments requiring multiple generation steps (e.g., ReAct-style tool use), maintain state in inp.state:


# file: react_agent.py

from typing import List
from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput

async def react_loop(inp: GenerateFnInput) -> GenerateFnOutput:
    """
    ReAct-style agent with tool use.
    Uses inp.state to track conversation history across calls.
    """
    sample = inp.sample
    max_turns = 10
    
    # Access or initialize per-sample state

    if not hasattr(inp.state, "react_histories"):
        inp.state.react_histories = {}
    
    history: List[dict] = inp.state.react_histories.get(sample.id, [])
    history.append({"role": "user", "content": sample.prompt})
    
    for turn in range(max_turns):
        # Call your underlying model (external API, local LLM, etc.)

        response = await call_agent_model(history)
        
        if "<tool>" in response:
            # Execute tool and continue loop

            tool_result = await execute_tool(extract_tool_call(response))
            history.append({"role": "assistant", "content": response})
            history.append({"role": "tool", "content": tool_result})
        else:
            # Final answer

            break
    
    sample.response = response
    sample.reward = score_react_trajectory(history)  # Custom reward shaping

    
    # Persist history for potential future use

    inp.state.react_histories[sample.id] = history
    
    return GenerateFnOutput(samples=[sample])

Pattern 2: External API Integration

Connect to proprietary models or services:


# file: external_api_bridge.py

import aiohttp
from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput

API_ENDPOINT = "https://api.proprietary-model.com/v1/generate"

async def external_api_generate(inp: GenerateFnInput) -> GenerateFnOutput:
    """Bridge to external generation API."""
    sample = inp.sample
    
    async with aiohttp.ClientSession() as session:
        payload = {
            "prompt": sample.prompt,
            "temperature": inp.sampling_params.temperature,
            "max_tokens": inp.sampling_params.max_new_tokens,
        }
        async with session.post(API_ENDPOINT, json=payload) as resp:
            result = await resp.json()
    
    sample.response = result["text"]
    # External APIs may not provide log-probs; Miles handles missing fields gracefully

    
    return GenerateFnOutput(samples=[sample])

Pattern 3: Deterministic Test Environments

For unit testing or reproducible baselines:


# file: deterministic_env.py

from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput

async def fixed_response_generator(inp: GenerateFnInput) -> GenerateFnOutput:
    """Returns canned responses for regression testing."""
    sample = inp.sample
    
    # Lookup based on prompt hash or ID

    canned_responses = {
        "math_problem_001": "The answer is 42.",
        "code_task_042": "def solve():\n    return 0",
    }
    
    sample.response = canned_responses.get(sample.sample_id, "Unknown")
    sample.reward = 1.0 if sample.response != "Unknown" else 0.0
    
    return GenerateFnOutput(samples=[sample])

Key Repository Files for Custom Integration

File Purpose Critical Lines
miles/rollout/sglang_rollout.py Orchestrates generation; loads and invokes custom functions 317-322: branching logic for custom vs. default generation
miles/utils/function_registry.py load_generate_function() — dynamic import mechanism Entire file: ~80 lines of import utilities
miles/rollout/base_types.py GenerateFnInput, GenerateFnOutput dataclasses Class definitions and field specifications
miles/utils/types.py Sample model — the object your function manipulates Response, reward, and metadata fields
miles/utils/eval_config.py EvalDatasetConfig — parses custom_generate_function_path from TOML Dataset loading and attribute propagation
tests/fast/rollout/generate_hub/test_multi_turn.py Reference tests for custom generate paths Example fixtures and assertions

Common Pitfalls and Solutions

  • Synchronous blocking: Always use async def. If your underlying library is synchronous, wrap it with asyncio.to_thread() or loop.run_in_executor().
  • Missing response field: Miles requires sample.response to be set. Unset fields cause downstream failures in reward model scoring.
  • Import path errors: Ensure your module is in PYTHONPATH. Use python -c "import my_module; print(my_module.my_fn)" to verify before running Miles.
  • State isolation: inp.state is global across samples in a batch. Use sample.id as a key to maintain per-sample isolation.

Summary

  • Implement an async function with signature GenerateFnInput → GenerateFnOutput to create a custom agentic environment
  • Register your function globally via --custom-generate-function-path or per-sample via custom_generate_function_path in dataset configs
  • Manipulate the Sample object to set response and optionally reward; Miles handles metadata attachment
  • Leverage inp.state for cross-turn state in multi-step agents, and inp.sampling_params for generation controls
  • Consult miles/rollout/sglang_rollout.py lines 317-322 for the exact integration point in the codebase

Frequently Asked Questions

What happens if both the CLI flag and per-sample path are set?

The per-sample generate_function_path takes precedence. This allows you to set a default global environment while overriding specific evaluation datasets with specialized ones. The resolution logic in miles/rollout/sglang_rollout.py explicitly checks the sample attribute first, falling back to args.custom_generate_function_path only when absent.

Can my custom function access GPU resources or distributed state?

Yes. The GenerateFnInput.state parameter contains GlobalGenerationState, which holds references to the model, tokenizer, and distributed training context when available. However, most agentic environments using external APIs or simulators will ignore these fields and manage their own compute.

How do I debug a custom generate function?

Set PYTHONPATH to include your module directory, then test the import chain: python -c "from miles.utils.function_registry import load_generate_function; fn = load_generate_function('my_module.my_fn'); print(fn)". For runtime debugging, add logging inside your async function—output appears in Miles' standard logs. The test file tests/fast/rollout/generate_hub/test_multi_turn.py provides a minimal harness for isolated testing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →