How Marin Configures and Runs lm-evaluation-harness Evaluations
Marin wraps lm-evaluation-harness in an isolated uv environment, converting the LmEvalRun configuration dataclass into CLI arguments that query OpenAI-compatible endpoints.
Marin integrates the EleutherAI lm-evaluation-harness through a lightweight Python wrapper that abstracts CLI complexity into type-safe configuration objects. The core integration lives in lib/marin/src/marin/evaluation/lm_eval.py, where LmEvalRun captures evaluation parameters and run_lm_eval constructs the exact subprocess command needed to benchmark models served via vLLM, Levanter, or similar APIs.
Configuration Dataclass: LmEvalRun
The LmEvalRun dataclass in lib/marin/src/marin/evaluation/lm_eval.py serves as the single source of truth for evaluation parameters. It exposes eight configurable fields that map directly to lm-evaluation-harness CLI flags:
tasks– List of task names (e.g.,["mmlu", "gsm8k"]) required for the benchmark suite.adapter– Enum specifying the OpenAI-compatible endpoint style, eitherLOCAL_COMPLETIONSorLOCAL_CHAT_COMPLETIONS. Defaults toLOCAL_COMPLETIONS.apply_chat_template– Boolean flag that tells LM Eval to apply the model’s chat template before generation. Defaults toFalse.limit– Hard integer cap on evaluation instances per task.num_fewshot– Integer specifying how many few-shot examples to prepend.batch_size– Integer or"auto"string controlling request batching.confirm_run_unsafe_code– Safety boolean required by LM Eval for tasks that execute generated code.extra_model_args– Dictionary of arbitrary key-value pairs merged into the model arguments string.
This structure allows researchers to define reproducible evaluation protocols programmatically rather than manipulating raw shell commands.
Constructing Model Arguments
Before execution, Marin builds the comma-separated model_args string that lm-evaluation-harness expects via the build_lm_eval_model_args function. This helper constructs a dictionary containing:
model_args = {
"model": model.endpoint.model,
"base_url": model.endpoint.url(run.adapter.endpoint_path),
"tokenizer_backend": "huggingface",
"tokenized_requests": False,
}
If the served model requires authentication, the function injects api_key; if a specific tokenizer path is provided, it adds tokenizer. The function then validates that no keys or values contain commas or equals signs—characters that would break the CLI parsing—and joins entries into the key=value,key=value format required by LM Eval. Finally, it merges run.extra_model_args to allow custom configuration overrides.
Isolated Execution Environment
The run_lm_eval function constructs a hermetic execution environment using uv run to prevent dependency conflicts. The implementation issues the following command structure:
def run_lm_eval(model: RunningModel, run: LmEvalRun, output_path: str) -> None:
command = ["uv", "run", "--isolated", "--no-project"]
for package in LM_EVAL_UV_PACKAGES:
command.extend(["--with", package])
command.extend([
"lm_eval",
"--model", run.adapter.value,
"--model_args", build_lm_eval_model_args(model, run),
"--tasks", ",".join(run.tasks),
"--output_path", output_path,
"--log_samples",
])
# ... conditional flags ...
subprocess.run(command, check=True)
Pinned dependencies are defined in LM_EVAL_UV_PACKAGES, which specifies the exact lm-evaluation-harness commit and a compatible transformers version. The --isolated --no-project flags ensure the evaluation runs in a fresh virtual environment, completely decoupled from Marin’s own dependency tree.
Complete Evaluation Workflow
Marin orchestrates lm-evaluation-harness evaluations through a five-stage pipeline:
- Model Serving – Marin launches an OpenAI-compatible server (vLLM, Levanter, etc.) exposing
/v1/completionsor/v1/chat/completionsendpoints. - Configuration – Instantiate
LmEvalRunwith desired tasks, adapter type, and constraints. - Execution –
run_lm_evalbuilds theuvcommand, installs LM Eval on-the-fly, and spawns the subprocess. - Artifact Generation – LM Eval writes JSON results to
--output_pathalongsidesamples.parquetcontaining per-instance generations. - Result Wrapping – Marin encapsulates outputs in an
LmEvalResultsartifact for downstream analysis or reporting pipelines.
The high-level launcher available at experiments/evaluation/cli.py exposes this functionality via command-line flags, while the programmatic API allows direct integration into custom training loops.
Usage Examples
Run MMLU and GSM8K benchmarks programmatically against an already-served model:
from marin.evaluation.lm_eval import LmEvalRun, run_lm_eval, LmEvalAdapter
lm_run = LmEvalRun(
tasks=["mmlu", "gsm8k"],
adapter=LmEvalAdapter.LOCAL_COMPLETIONS,
limit=10,
num_fewshot=5,
apply_chat_template=False,
)
run_lm_eval(model, lm_run, output_path="/tmp/eval-output")
Launch the same evaluation via the shared CLI launcher:
uv run python -m experiments.evaluation.cli launch \
--model qwen3-0.6b \
--evals mmlu,gsm8k \
--limit 10 \
--no-wait
Summary
- Type-safe configuration – The
LmEvalRundataclass inlib/marin/src/marin/evaluation/lm_eval.pycentralizes all lm-evaluation-harness parameters. - Dynamic CLI construction –
build_lm_eval_model_argsvalidates and formats the comma-separated argument string required by LM Eval. - Hermetic execution –
run_lm_evalusesuv run --isolatedto install pinned versions oflm-eval[api]without contaminating the host environment. - Adapter flexibility – Supports both completion-style (
local-completions) and chat-style (local-chat-completions) OpenAI-compatible endpoints. - Dual interfaces – Available through both the Python API (
run_lm_eval) and the CLI entry point atexperiments/evaluation/cli.py.
Frequently Asked Questions
How does Marin isolate lm-evaluation-harness dependencies from the main project?
Marin uses uv run with the --isolated and --no-project flags to create a temporary virtual environment for each evaluation. This environment installs specific package versions defined in LM_EVAL_UV_PACKAGES (located in lib/marin/src/marin/evaluation/lm_eval.py), ensuring that LM Eval’s dependencies never conflict with Marin’s core libraries.
What endpoint types does Marin support for model serving?
Marin supports two adapter types through the LmEvalAdapter enum: LOCAL_COMPLETIONS for legacy completion endpoints and LOCAL_CHAT_COMPLETIONS for chat-formatted APIs. The adapter selection determines the base_url path construction and the --model flag passed to lm-evaluation-harness.
How are custom model arguments passed to lm-evaluation-harness?
Users provide arbitrary key-value pairs via the extra_model_args field in LmEvalRun. These dictionary entries merge with the automatically generated model arguments (such as base_url and tokenizer_backend) in build_lm_eval_model_args, allowing customization of parameters like temperature or specific model revision hashes.
Where does Marin store evaluation results?
The output_path parameter in run_lm_eval specifies the directory where lm-evaluation-harness writes its JSON results file and samples.parquet. Marin then wraps these artifacts into an LmEvalResults object, making them available for downstream processing or export to experiment tracking systems.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →