# How to Integrate Custom Agentic Environments with Miles: A Complete Guide

> Integrate custom agentic environments with Miles by implementing a generate function and registering it via command-line or dataset configuration. Unlock Miles's full potential today.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: how-to-guide
- Published: 2026-09-06

---

**You integrate custom agentic environments with Miles by implementing an async generate function that conforms to the `GenerateFnInput → GenerateFnOutput` contract, then registering it via `--custom-generate-function-path` or per-sample through dataset configuration.**

Miles, an open-source reinforcement learning framework for language models, provides a flexible **pluggable generation backend** that lets you swap its default SGLang generator for your own custom logic. This integration point lives at [`miles/rollout/sglang_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/sglang_rollout.py) and enables you to connect external APIs, simulation environments, tool-use loops, or proprietary models while retaining Miles' reward modeling and training infrastructure.

## How the Custom Generate Function Works

Miles' rollout engine discovers custom functions through two channels:

1. **Global CLI flag**: `--custom-generate-function-path my_module.my_function`
2. **Per-sample override**: `sample.generate_function_path` (typically set in dataset configs)

When either path is present, Miles calls `load_generate_function` from [`miles/utils/function_registry.py`](https://github.com/radixark/miles/blob/main/miles/utils/function_registry.py) to dynamically import your callable. The function then executes inside the async `generate_and_rm` loop, receiving full context about the current generation state and returning structured results.

### The Core Data Flow

The integration follows this contract as defined in [`miles/rollout/base_types.py`](https://github.com/radixark/miles/blob/main/miles/rollout/base_types.py):

```python
from dataclasses import dataclass
from typing import List
from miles.utils.types import Sample

@dataclass
class GenerateFnInput:
    state: "GlobalGenerationState"  # Shared state across all samples

    sample: Sample                   # The sample being processed

    sampling_params: "SamplingParams"  # Temperature, top-p, etc.

    evaluation: bool                 # Whether this is eval or training

@dataclass
class GenerateFnOutput:
    samples: List[Sample]            # Updated samples (usually length 1)

```

Your function receives `GenerateFnInput`, performs arbitrary computation, populates `sample.response` (and optionally `sample.reward`), then returns `GenerateFnOutput`. Miles handles log-probability computation, token tracking, and downstream RM scoring automatically.

## Step-by-Step Implementation

### Step 1: Create Your Custom Generate Function

Define an async function in a Python module. This example shows a minimal echo environment:

```python

# file: my_agentic_env.py

from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput

async def my_generate_fn(inp: GenerateFnInput) -> GenerateFnOutput:
    """
    Custom agentic environment for Miles.
    
    Parameters
    ----------
    inp : GenerateFnInput
        Contains the sample, global state, and generation parameters.
    
    Returns
    -------
    GenerateFnOutput
        The sample(s) with populated response field.
    """
    sample = inp.sample
    
    # ----- Agentic logic: replace with your environment -----

    # Examples:

    # - Call OpenAI/Anthropic API with tool use

    # - Execute in a Python sandbox (e.g., Jupyter kernel)

    # - Query a simulation environment (Minecraft, WebArena, etc.)

    # - Route to a local vLLM instance with custom sampling

    
    sample.response = f"[AGENT EXECUTION] {sample.prompt}"
    sample.reward = 0.5  # Optional: provide intermediate reward

    # --------------------------------------------------------

    
    return GenerateFnOutput(samples=[sample])

```

Key requirements:

- Use `async def` — the rollout loop is fully asynchronous
- Modify `sample.response` — this field becomes the model's output
- Return `GenerateFnOutput` with a list of samples (usually one)

### Step 2: Register via Command Line

Pass your function's fully-qualified import path to any Miles rollout script:

```bash
python -m miles.main.scripts.run_qwen3_dense \
    --custom-generate-function-path my_agentic_env.my_generate_fn \
    --model-path Qwen/Qwen3-8B \
    --output-dir ./outputs

```

Miles executes `load_generate_function("my_agentic_env.my_generate_fn")`, which imports `my_agentic_env` and extracts the callable.

### Step 3: Register Per-Sample via Dataset Config

For evaluation datasets where different samples need different environments, use TOML configuration:

```toml

# file: webarena_eval.toml

[dataset]
path = "data/webarena_tasks.jsonl"
input_key = "prompt"
label_key = "gold_answer"

# Custom generate function applies to all samples in this dataset

custom_generate_function_path = "environments.webarena_agent.generate"

[[dataset]]
path = "data/code_execution_tasks.jsonl"
input_key = "coding_problem"
label_key = "test_cases"
custom_generate_function_path = "environments.sandbox_executor.generate"

```

When `EvalDatasetConfig` loads this file (in [`miles/utils/eval_config.py`](https://github.com/radixark/miles/blob/main/miles/utils/eval_config.py)), it stores `custom_generate_function_path` on each `Sample` instance. The rollout code at [`miles/rollout/sglang_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/sglang_rollout.py) (lines 317-322) prefers this per-sample attribute:

```python

# From miles/rollout/sglang_rollout.py

custom_func_path = (
    getattr(sample, "generate_function_path", None) 
    or args.custom_generate_function_path
)
generate_fn = load_generate_function(custom_func_path) if custom_func_path else None

if generate_fn is not None:
    output = await generate_fn(
        GenerateFnInput(
            state=state,
            sample=sample,
            sampling_params=sampling_params,
            evaluation=evaluation
        )
    )
    samples = output.samples
else:
    # Fall back to default SGLang generation

    samples = await generate(args, sample, sampling_params)

```

This branching logic ensures your custom environment integrates seamlessly with Miles' existing sampling and reward infrastructure.

## Advanced Integration Patterns

### Pattern 1: Multi-Turn Agent Loops

For environments requiring multiple generation steps (e.g., ReAct-style tool use), maintain state in `inp.state`:

```python

# file: react_agent.py

from typing import List
from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput

async def react_loop(inp: GenerateFnInput) -> GenerateFnOutput:
    """
    ReAct-style agent with tool use.
    Uses inp.state to track conversation history across calls.
    """
    sample = inp.sample
    max_turns = 10
    
    # Access or initialize per-sample state

    if not hasattr(inp.state, "react_histories"):
        inp.state.react_histories = {}
    
    history: List[dict] = inp.state.react_histories.get(sample.id, [])
    history.append({"role": "user", "content": sample.prompt})
    
    for turn in range(max_turns):
        # Call your underlying model (external API, local LLM, etc.)

        response = await call_agent_model(history)
        
        if "<tool>" in response:
            # Execute tool and continue loop

            tool_result = await execute_tool(extract_tool_call(response))
            history.append({"role": "assistant", "content": response})
            history.append({"role": "tool", "content": tool_result})
        else:
            # Final answer

            break
    
    sample.response = response
    sample.reward = score_react_trajectory(history)  # Custom reward shaping

    
    # Persist history for potential future use

    inp.state.react_histories[sample.id] = history
    
    return GenerateFnOutput(samples=[sample])

```

### Pattern 2: External API Integration

Connect to proprietary models or services:

```python

# file: external_api_bridge.py

import aiohttp
from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput

API_ENDPOINT = "https://api.proprietary-model.com/v1/generate"

async def external_api_generate(inp: GenerateFnInput) -> GenerateFnOutput:
    """Bridge to external generation API."""
    sample = inp.sample
    
    async with aiohttp.ClientSession() as session:
        payload = {
            "prompt": sample.prompt,
            "temperature": inp.sampling_params.temperature,
            "max_tokens": inp.sampling_params.max_new_tokens,
        }
        async with session.post(API_ENDPOINT, json=payload) as resp:
            result = await resp.json()
    
    sample.response = result["text"]
    # External APIs may not provide log-probs; Miles handles missing fields gracefully

    
    return GenerateFnOutput(samples=[sample])

```

### Pattern 3: Deterministic Test Environments

For unit testing or reproducible baselines:

```python

# file: deterministic_env.py

from miles.rollout.base_types import GenerateFnInput, GenerateFnOutput

async def fixed_response_generator(inp: GenerateFnInput) -> GenerateFnOutput:
    """Returns canned responses for regression testing."""
    sample = inp.sample
    
    # Lookup based on prompt hash or ID

    canned_responses = {
        "math_problem_001": "The answer is 42.",
        "code_task_042": "def solve():\n    return 0",
    }
    
    sample.response = canned_responses.get(sample.sample_id, "Unknown")
    sample.reward = 1.0 if sample.response != "Unknown" else 0.0
    
    return GenerateFnOutput(samples=[sample])

```

## Key Repository Files for Custom Integration

| File | Purpose | Critical Lines |
|------|---------|--------------|
| [`miles/rollout/sglang_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/sglang_rollout.py) | Orchestrates generation; loads and invokes custom functions | 317-322: branching logic for custom vs. default generation |
| [`miles/utils/function_registry.py`](https://github.com/radixark/miles/blob/main/miles/utils/function_registry.py) | `load_generate_function()` — dynamic import mechanism | Entire file: ~80 lines of import utilities |
| [`miles/rollout/base_types.py`](https://github.com/radixark/miles/blob/main/miles/rollout/base_types.py) | `GenerateFnInput`, `GenerateFnOutput` dataclasses | Class definitions and field specifications |
| [`miles/utils/types.py`](https://github.com/radixark/miles/blob/main/miles/utils/types.py) | `Sample` model — the object your function manipulates | Response, reward, and metadata fields |
| [`miles/utils/eval_config.py`](https://github.com/radixark/miles/blob/main/miles/utils/eval_config.py) | `EvalDatasetConfig` — parses `custom_generate_function_path` from TOML | Dataset loading and attribute propagation |
| [`tests/fast/rollout/generate_hub/test_multi_turn.py`](https://github.com/radixark/miles/blob/main/tests/fast/rollout/generate_hub/test_multi_turn.py) | Reference tests for custom generate paths | Example fixtures and assertions |

## Common Pitfalls and Solutions

- **Synchronous blocking**: Always use `async def`. If your underlying library is synchronous, wrap it with `asyncio.to_thread()` or `loop.run_in_executor()`.
- **Missing `response` field**: Miles requires `sample.response` to be set. Unset fields cause downstream failures in reward model scoring.
- **Import path errors**: Ensure your module is in `PYTHONPATH`. Use `python -c "import my_module; print(my_module.my_fn)"` to verify before running Miles.
- **State isolation**: `inp.state` is global across samples in a batch. Use `sample.id` as a key to maintain per-sample isolation.

## Summary

- **Implement** an async function with signature `GenerateFnInput → GenerateFnOutput` to create a custom agentic environment
- **Register** your function globally via `--custom-generate-function-path` or per-sample via `custom_generate_function_path` in dataset configs
- **Manipulate** the `Sample` object to set `response` and optionally `reward`; Miles handles metadata attachment
- **Leverage** `inp.state` for cross-turn state in multi-step agents, and `inp.sampling_params` for generation controls
- **Consult** [`miles/rollout/sglang_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/sglang_rollout.py) lines 317-322 for the exact integration point in the codebase

## Frequently Asked Questions

### What happens if both the CLI flag and per-sample path are set?

The per-sample `generate_function_path` takes precedence. This allows you to set a default global environment while overriding specific evaluation datasets with specialized ones. The resolution logic in [`miles/rollout/sglang_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/sglang_rollout.py) explicitly checks the sample attribute first, falling back to `args.custom_generate_function_path` only when absent.

### Can my custom function access GPU resources or distributed state?

Yes. The `GenerateFnInput.state` parameter contains `GlobalGenerationState`, which holds references to the model, tokenizer, and distributed training context when available. However, most agentic environments using external APIs or simulators will ignore these fields and manage their own compute.

### How do I debug a custom generate function?

Set `PYTHONPATH` to include your module directory, then test the import chain: `python -c "from miles.utils.function_registry import load_generate_function; fn = load_generate_function('my_module.my_fn'); print(fn)"`. For runtime debugging, add logging inside your async function—output appears in Miles' standard logs. The test file [`tests/fast/rollout/generate_hub/test_multi_turn.py`](https://github.com/radixark/miles/blob/main/tests/fast/rollout/generate_hub/test_multi_turn.py) provides a minimal harness for isolated testing.