# How to Run Benchmarks to Evaluate and Optimize Your Configuration in Local Deep Research

> Learn to run benchmarks in local deep research using our CLI or Python API. Evaluate accuracy and speed, then optimize your configuration automatically for peak performance.

- Repository: [learningcircuit/local-deep-research](https://github.com/learningcircuit/local-deep-research)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Run benchmarks using the CLI wrapper at [`examples/run_benchmark.py`](https://github.com/learningcircuit/local-deep-research/blob/main/examples/run_benchmark.py) or the Python API in [`src/local_deep_research/api/benchmark_functions.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/api/benchmark_functions.py) to measure accuracy and speed, then use `optimize_parameters()` for automated Bayesian-style tuning.**

The **learningcircuit/local-deep-research** repository ships with a comprehensive benchmarking framework that lets you systematically run benchmarks to evaluate and optimize configuration parameters. Whether you need to compare search strategies or fine-tune iteration counts, the toolkit provides predefined evaluations like SimpleQA and BrowseComp alongside automated optimization routines.

## Architecture Overview

The benchmarking system follows a layered architecture that separates user interfaces from execution logic. At the top level, [`examples/run_benchmark.py`](https://github.com/learningcircuit/local-deep-research/blob/main/examples/run_benchmark.py) provides a **CLI wrapper** that parses arguments and dispatches to evaluation functions. The middle layer lives in [`src/local_deep_research/api/benchmark_functions.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/api/benchmark_functions.py), exposing high-level Python functions such as `evaluate_simpleqa()`, `evaluate_browsecomp()`, and `compare_configurations()`. These functions construct search and evaluation configs before invoking the low-level **benchmark runners** in `src/local_deep_research/benchmarks/`, which handle data loading, query execution via your selected tool (e.g., `searxng`), and metric calculation through `calculate_metrics()`. Finally, the **optimization driver** in [`examples/optimization/example_optimization.py`](https://github.com/learningcircuit/local-deep-research/blob/main/examples/optimization/example_optimization.py) demonstrates how to invoke `optimize_parameters()` from [`src/local_deep_research/benchmarks/optimization.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/benchmarks/optimization.py) to perform automated parameter search.

The execution flow proceeds as follows:

1. User invokes CLI or Python API to build `search_config` and `evaluation_config`
2. Benchmark runner executes searches using the configured tool
3. System collects answers and evaluates via LLM or human judgment
4. `calculate_metrics()` computes accuracy and average time
5. `generate_report()` writes detailed markdown reports to the output directory

## Running Benchmarks from the Command Line

The fastest way to run benchmarks to evaluate and optimize configuration settings is through the provided CLI wrapper. This script supports running individual benchmarks or comparing multiple configurations side-by-side.

To run a SimpleQA benchmark with specific search parameters:

```bash
python examples/run_benchmark.py --benchmark simpleqa --examples 10 \
    --iterations 3 --questions 3 --search-tool searxng \
    --output-dir my_benchmark_results

```

Switch the `--benchmark` argument to `browsecomp` for web-browsing evaluations or `compare` to test multiple configurations in a single run. All CLI arguments map directly to the parameters of the underlying Python API functions defined in [`benchmark_functions.py`](https://github.com/learningcircuit/local-deep-research/blob/main/benchmark_functions.py).

## Benchmarking via the Python API

For programmatic integration, import functions from [`src/local_deep_research/api/benchmark_functions.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/api/benchmark_functions.py) to embed benchmarking into your training or CI pipelines.

### Running SimpleQA and BrowseComp

Use `evaluate_simpleqa()` or `evaluate_browsecomp()` to run a single benchmark configuration and retrieve detailed metrics:

```python
from local_deep_research.api.benchmark_functions import evaluate_simpleqa

# Run SimpleQA with 2 iterations and 3 questions per iteration

results = evaluate_simpleqa(
    num_examples=20,
    search_iterations=2,
    questions_per_iteration=3,
    search_tool="searxng",
    human_evaluation=False,
    output_dir="benchmark_output"
)

print(f"Accuracy: {results['metrics']['accuracy']}")
print(f"Report saved to: {results['report_path']}")

```

Both functions return a dictionary containing `metrics`, `total_examples`, and the path to the generated markdown report.

### Comparing Multiple Configurations

When you need to run benchmarks to evaluate and optimize configuration alternatives, use `compare_configurations()` to test multiple parameter sets against the same dataset:

```python
from local_deep_research.api.benchmark_functions import compare_configurations

comparison = compare_configurations(
    dataset_type="simpleqa",
    num_examples=20,
    configurations=[
        {
            "name": "Base",
            "search_tool": "searxng",
            "iterations": 1,
            "questions_per_iteration": 3
        },
        {
            "name": "DeepSearch",
            "search_tool": "searxng",
            "iterations": 3,
            "questions_per_iteration": 3
        },
        {
            "name": "BroadQuery",
            "search_tool": "searxng",
            "iterations": 1,
            "questions_per_iteration": 5
        }
    ],
    output_dir="comparison_results"
)

print(f"Comparison report: {comparison['report_path']}")

```

This routine iterates over each configuration dictionary, runs the selected benchmark, and aggregates results into a single markdown comparison report.

## Automated Parameter Optimization

For hands-free tuning, the framework provides Bayesian-style optimization through `optimize_parameters()` in [`src/local_deep_research/benchmarks/optimization.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/benchmarks/optimization.py). This function searches your defined parameter space to find the optimal balance between quality and speed.

The example script at [`examples/optimization/example_optimization.py`](https://github.com/learningcircuit/local-deep-research/blob/main/examples/optimization/example_optimization.py) demonstrates the complete workflow:

```python
from local_deep_research.benchmarks.optimization import optimize_parameters
from pathlib import Path
import json

# Define the parameter space to search

param_space = {
    "iterations": {"type": "int", "low": 1, "high": 3, "step": 1},
    "questions_per_iteration": {"type": "int", "low": 1, "high": 5, "step": 1},
    "search_strategy": {"type": "categorical", "choices": ["rapid", "thorough"]},
}

# Run optimization with 10 trials

best_params, best_score = optimize_parameters(
    query="SimpleQA optimization run",
    search_tool="searxng",
    n_trials=10,
    output_dir=str(Path("optimization_results") / "run1"),
    param_space=param_space,
    metric_weights={"quality": 0.6, "speed": 0.4}
)

print(f"Best configuration: {best_params}")
print(f"Best weighted score: {best_score}")

# Persist results

summary = {"best_params": best_params, "best_score": best_score}
Path("optimization_results/run1/summary.json").write_text(
    json.dumps(summary, indent=2)
)

```

The optimizer internally repeats benchmarks for sampled parameter combinations, scoring each using your specified `metric_weights`, and returns the highest-scoring configuration.

## Summary

- **CLI Entry Point**: Use [`examples/run_benchmark.py`](https://github.com/learningcircuit/local-deep-research/blob/main/examples/run_benchmark.py) for quick command-line evaluations of SimpleQA, BrowseComp, or configuration comparisons.
- **Python API**: Import `evaluate_simpleqa()`, `evaluate_browsecomp()`, and `compare_configurations()` from [`src/local_deep_research/api/benchmark_functions.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/api/benchmark_functions.py) for programmatic control.
- **Optimization**: Leverage `optimize_parameters()` from [`src/local_deep_research/benchmarks/optimization.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/benchmarks/optimization.py) to automatically search parameter spaces using weighted quality and speed metrics.
- **Output**: All methods generate detailed markdown reports and JSON summaries containing accuracy, timing metrics, and configuration details.
- **Integration**: The architecture supports `searxng` and other search tools, allowing you to evaluate different retrieval backends against the same benchmark datasets.

## Frequently Asked Questions

### What benchmarks are available in Local Deep Research?

The framework currently supports **SimpleQA** (factual question answering), **BrowseComp** (web browsing comprehension), and **DeepSearch** evaluations. These are implemented in the `src/local_deep_research/benchmarks/` directory and accessible via the CLI or Python API.

### How do I choose between running benchmarks via CLI versus Python API?

Use the **CLI** ([`examples/run_benchmark.py`](https://github.com/learningcircuit/local-deep-research/blob/main/examples/run_benchmark.py)) for one-off evaluations or shell-scripted automation. Use the **Python API** when you need to integrate benchmarking into larger applications, customize evaluation logic, or process results programmatically within a Python environment.

### What metrics does the benchmarking framework calculate?

The system calculates **accuracy** (percentage of correct answers) and **average execution time** via the `calculate_metrics()` function. When using `optimize_parameters()`, you can specify `metric_weights` to balance these factors according to your quality-versus-speed priorities.

### How does the automated optimization determine the best configuration?

The `optimize_parameters()` function performs a **Bayesian-style search** over your defined `param_space`, running the benchmark for each sampled configuration. It scores trials using the weighted combination of metrics you specify (typically accuracy and speed), then returns the parameter set that achieves the highest composite score after the specified number of trials.