# How Marin Configures and Runs lm-evaluation-harness Evaluations

> Learn how Marin configures and runs lm-evaluation-harness evaluations using a uv environment and CLI arguments for OpenAI-compatible endpoints.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-28

---

**Marin wraps lm-evaluation-harness in an isolated uv environment, converting the `LmEvalRun` configuration dataclass into CLI arguments that query OpenAI-compatible endpoints.**

Marin integrates the EleutherAI lm-evaluation-harness through a lightweight Python wrapper that abstracts CLI complexity into type-safe configuration objects. The core integration lives in [`lib/marin/src/marin/evaluation/lm_eval.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/evaluation/lm_eval.py), where `LmEvalRun` captures evaluation parameters and `run_lm_eval` constructs the exact subprocess command needed to benchmark models served via vLLM, Levanter, or similar APIs.

## Configuration Dataclass: LmEvalRun

The `LmEvalRun` dataclass in [`lib/marin/src/marin/evaluation/lm_eval.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/evaluation/lm_eval.py) serves as the single source of truth for evaluation parameters. It exposes eight configurable fields that map directly to lm-evaluation-harness CLI flags:

- **`tasks`** – List of task names (e.g., `["mmlu", "gsm8k"]`) required for the benchmark suite.
- **`adapter`** – Enum specifying the OpenAI-compatible endpoint style, either `LOCAL_COMPLETIONS` or `LOCAL_CHAT_COMPLETIONS`. Defaults to `LOCAL_COMPLETIONS`.
- **`apply_chat_template`** – Boolean flag that tells LM Eval to apply the model’s chat template before generation. Defaults to `False`.
- **`limit`** – Hard integer cap on evaluation instances per task.
- **`num_fewshot`** – Integer specifying how many few-shot examples to prepend.
- **`batch_size`** – Integer or `"auto"` string controlling request batching.
- **`confirm_run_unsafe_code`** – Safety boolean required by LM Eval for tasks that execute generated code.
- **`extra_model_args`** – Dictionary of arbitrary key-value pairs merged into the model arguments string.

This structure allows researchers to define reproducible evaluation protocols programmatically rather than manipulating raw shell commands.

## Constructing Model Arguments

Before execution, Marin builds the comma-separated `model_args` string that lm-evaluation-harness expects via the `build_lm_eval_model_args` function. This helper constructs a dictionary containing:

```python
model_args = {
    "model": model.endpoint.model,
    "base_url": model.endpoint.url(run.adapter.endpoint_path),
    "tokenizer_backend": "huggingface",
    "tokenized_requests": False,
}

```

If the served model requires authentication, the function injects `api_key`; if a specific tokenizer path is provided, it adds `tokenizer`. The function then validates that no keys or values contain commas or equals signs—characters that would break the CLI parsing—and joins entries into the `key=value,key=value` format required by LM Eval. Finally, it merges `run.extra_model_args` to allow custom configuration overrides.

## Isolated Execution Environment

The `run_lm_eval` function constructs a hermetic execution environment using `uv run` to prevent dependency conflicts. The implementation issues the following command structure:

```python
def run_lm_eval(model: RunningModel, run: LmEvalRun, output_path: str) -> None:
    command = ["uv", "run", "--isolated", "--no-project"]
    for package in LM_EVAL_UV_PACKAGES:
        command.extend(["--with", package])
    command.extend([
        "lm_eval",
        "--model", run.adapter.value,
        "--model_args", build_lm_eval_model_args(model, run),
        "--tasks", ",".join(run.tasks),
        "--output_path", output_path,
        "--log_samples",
    ])
    # ... conditional flags ...

    subprocess.run(command, check=True)

```

**Pinned dependencies** are defined in `LM_EVAL_UV_PACKAGES`, which specifies the exact lm-evaluation-harness commit and a compatible `transformers` version. The `--isolated --no-project` flags ensure the evaluation runs in a fresh virtual environment, completely decoupled from Marin’s own dependency tree.

## Complete Evaluation Workflow

Marin orchestrates lm-evaluation-harness evaluations through a five-stage pipeline:

1. **Model Serving** – Marin launches an OpenAI-compatible server (vLLM, Levanter, etc.) exposing `/v1/completions` or `/v1/chat/completions` endpoints.
2. **Configuration** – Instantiate `LmEvalRun` with desired tasks, adapter type, and constraints.
3. **Execution** – `run_lm_eval` builds the `uv` command, installs LM Eval on-the-fly, and spawns the subprocess.
4. **Artifact Generation** – LM Eval writes JSON results to `--output_path` alongside `samples.parquet` containing per-instance generations.
5. **Result Wrapping** – Marin encapsulates outputs in an `LmEvalResults` artifact for downstream analysis or reporting pipelines.

The high-level launcher available at [`experiments/evaluation/cli.py`](https://github.com/marin-community/marin/blob/main/experiments/evaluation/cli.py) exposes this functionality via command-line flags, while the programmatic API allows direct integration into custom training loops.

## Usage Examples

Run MMLU and GSM8K benchmarks programmatically against an already-served model:

```python
from marin.evaluation.lm_eval import LmEvalRun, run_lm_eval, LmEvalAdapter

lm_run = LmEvalRun(
    tasks=["mmlu", "gsm8k"],
    adapter=LmEvalAdapter.LOCAL_COMPLETIONS,
    limit=10,
    num_fewshot=5,
    apply_chat_template=False,
)

run_lm_eval(model, lm_run, output_path="/tmp/eval-output")

```

Launch the same evaluation via the shared CLI launcher:

```bash
uv run python -m experiments.evaluation.cli launch \
  --model qwen3-0.6b \
  --evals mmlu,gsm8k \
  --limit 10 \
  --no-wait

```

## Summary

- **Type-safe configuration** – The `LmEvalRun` dataclass in [`lib/marin/src/marin/evaluation/lm_eval.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/evaluation/lm_eval.py) centralizes all lm-evaluation-harness parameters.
- **Dynamic CLI construction** – `build_lm_eval_model_args` validates and formats the comma-separated argument string required by LM Eval.
- **Hermetic execution** – `run_lm_eval` uses `uv run --isolated` to install pinned versions of `lm-eval[api]` without contaminating the host environment.
- **Adapter flexibility** – Supports both completion-style (`local-completions`) and chat-style (`local-chat-completions`) OpenAI-compatible endpoints.
- **Dual interfaces** – Available through both the Python API (`run_lm_eval`) and the CLI entry point at [`experiments/evaluation/cli.py`](https://github.com/marin-community/marin/blob/main/experiments/evaluation/cli.py).

## Frequently Asked Questions

### How does Marin isolate lm-evaluation-harness dependencies from the main project?

Marin uses `uv run` with the `--isolated` and `--no-project` flags to create a temporary virtual environment for each evaluation. This environment installs specific package versions defined in `LM_EVAL_UV_PACKAGES` (located in [`lib/marin/src/marin/evaluation/lm_eval.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/evaluation/lm_eval.py)), ensuring that LM Eval’s dependencies never conflict with Marin’s core libraries.

### What endpoint types does Marin support for model serving?

Marin supports two adapter types through the `LmEvalAdapter` enum: `LOCAL_COMPLETIONS` for legacy completion endpoints and `LOCAL_CHAT_COMPLETIONS` for chat-formatted APIs. The adapter selection determines the `base_url` path construction and the `--model` flag passed to lm-evaluation-harness.

### How are custom model arguments passed to lm-evaluation-harness?

Users provide arbitrary key-value pairs via the `extra_model_args` field in `LmEvalRun`. These dictionary entries merge with the automatically generated model arguments (such as `base_url` and `tokenizer_backend`) in `build_lm_eval_model_args`, allowing customization of parameters like temperature or specific model revision hashes.

### Where does Marin store evaluation results?

The `output_path` parameter in `run_lm_eval` specifies the directory where lm-evaluation-harness writes its JSON results file and `samples.parquet`. Marin then wraps these artifacts into an `LmEvalResults` object, making them available for downstream processing or export to experiment tracking systems.