How to Run Benchmarks to Evaluate and Optimize Your Configuration in Local Deep Research

Run benchmarks using the CLI wrapper at examples/run_benchmark.py or the Python API in src/local_deep_research/api/benchmark_functions.py to measure accuracy and speed, then use optimize_parameters() for automated Bayesian-style tuning.

The learningcircuit/local-deep-research repository ships with a comprehensive benchmarking framework that lets you systematically run benchmarks to evaluate and optimize configuration parameters. Whether you need to compare search strategies or fine-tune iteration counts, the toolkit provides predefined evaluations like SimpleQA and BrowseComp alongside automated optimization routines.

Architecture Overview

The benchmarking system follows a layered architecture that separates user interfaces from execution logic. At the top level, examples/run_benchmark.py provides a CLI wrapper that parses arguments and dispatches to evaluation functions. The middle layer lives in src/local_deep_research/api/benchmark_functions.py, exposing high-level Python functions such as evaluate_simpleqa(), evaluate_browsecomp(), and compare_configurations(). These functions construct search and evaluation configs before invoking the low-level benchmark runners in src/local_deep_research/benchmarks/, which handle data loading, query execution via your selected tool (e.g., searxng), and metric calculation through calculate_metrics(). Finally, the optimization driver in examples/optimization/example_optimization.py demonstrates how to invoke optimize_parameters() from src/local_deep_research/benchmarks/optimization.py to perform automated parameter search.

The execution flow proceeds as follows:

  1. User invokes CLI or Python API to build search_config and evaluation_config
  2. Benchmark runner executes searches using the configured tool
  3. System collects answers and evaluates via LLM or human judgment
  4. calculate_metrics() computes accuracy and average time
  5. generate_report() writes detailed markdown reports to the output directory

Running Benchmarks from the Command Line

The fastest way to run benchmarks to evaluate and optimize configuration settings is through the provided CLI wrapper. This script supports running individual benchmarks or comparing multiple configurations side-by-side.

To run a SimpleQA benchmark with specific search parameters:

python examples/run_benchmark.py --benchmark simpleqa --examples 10 \
    --iterations 3 --questions 3 --search-tool searxng \
    --output-dir my_benchmark_results

Switch the --benchmark argument to browsecomp for web-browsing evaluations or compare to test multiple configurations in a single run. All CLI arguments map directly to the parameters of the underlying Python API functions defined in benchmark_functions.py.

Benchmarking via the Python API

For programmatic integration, import functions from src/local_deep_research/api/benchmark_functions.py to embed benchmarking into your training or CI pipelines.

Running SimpleQA and BrowseComp

Use evaluate_simpleqa() or evaluate_browsecomp() to run a single benchmark configuration and retrieve detailed metrics:

from local_deep_research.api.benchmark_functions import evaluate_simpleqa

# Run SimpleQA with 2 iterations and 3 questions per iteration

results = evaluate_simpleqa(
    num_examples=20,
    search_iterations=2,
    questions_per_iteration=3,
    search_tool="searxng",
    human_evaluation=False,
    output_dir="benchmark_output"
)

print(f"Accuracy: {results['metrics']['accuracy']}")
print(f"Report saved to: {results['report_path']}")

Both functions return a dictionary containing metrics, total_examples, and the path to the generated markdown report.

Comparing Multiple Configurations

When you need to run benchmarks to evaluate and optimize configuration alternatives, use compare_configurations() to test multiple parameter sets against the same dataset:

from local_deep_research.api.benchmark_functions import compare_configurations

comparison = compare_configurations(
    dataset_type="simpleqa",
    num_examples=20,
    configurations=[
        {
            "name": "Base",
            "search_tool": "searxng",
            "iterations": 1,
            "questions_per_iteration": 3
        },
        {
            "name": "DeepSearch",
            "search_tool": "searxng",
            "iterations": 3,
            "questions_per_iteration": 3
        },
        {
            "name": "BroadQuery",
            "search_tool": "searxng",
            "iterations": 1,
            "questions_per_iteration": 5
        }
    ],
    output_dir="comparison_results"
)

print(f"Comparison report: {comparison['report_path']}")

This routine iterates over each configuration dictionary, runs the selected benchmark, and aggregates results into a single markdown comparison report.

Automated Parameter Optimization

For hands-free tuning, the framework provides Bayesian-style optimization through optimize_parameters() in src/local_deep_research/benchmarks/optimization.py. This function searches your defined parameter space to find the optimal balance between quality and speed.

The example script at examples/optimization/example_optimization.py demonstrates the complete workflow:

from local_deep_research.benchmarks.optimization import optimize_parameters
from pathlib import Path
import json

# Define the parameter space to search

param_space = {
    "iterations": {"type": "int", "low": 1, "high": 3, "step": 1},
    "questions_per_iteration": {"type": "int", "low": 1, "high": 5, "step": 1},
    "search_strategy": {"type": "categorical", "choices": ["rapid", "thorough"]},
}

# Run optimization with 10 trials

best_params, best_score = optimize_parameters(
    query="SimpleQA optimization run",
    search_tool="searxng",
    n_trials=10,
    output_dir=str(Path("optimization_results") / "run1"),
    param_space=param_space,
    metric_weights={"quality": 0.6, "speed": 0.4}
)

print(f"Best configuration: {best_params}")
print(f"Best weighted score: {best_score}")

# Persist results

summary = {"best_params": best_params, "best_score": best_score}
Path("optimization_results/run1/summary.json").write_text(
    json.dumps(summary, indent=2)
)

The optimizer internally repeats benchmarks for sampled parameter combinations, scoring each using your specified metric_weights, and returns the highest-scoring configuration.

Summary

  • CLI Entry Point: Use examples/run_benchmark.py for quick command-line evaluations of SimpleQA, BrowseComp, or configuration comparisons.
  • Python API: Import evaluate_simpleqa(), evaluate_browsecomp(), and compare_configurations() from src/local_deep_research/api/benchmark_functions.py for programmatic control.
  • Optimization: Leverage optimize_parameters() from src/local_deep_research/benchmarks/optimization.py to automatically search parameter spaces using weighted quality and speed metrics.
  • Output: All methods generate detailed markdown reports and JSON summaries containing accuracy, timing metrics, and configuration details.
  • Integration: The architecture supports searxng and other search tools, allowing you to evaluate different retrieval backends against the same benchmark datasets.

Frequently Asked Questions

What benchmarks are available in Local Deep Research?

The framework currently supports SimpleQA (factual question answering), BrowseComp (web browsing comprehension), and DeepSearch evaluations. These are implemented in the src/local_deep_research/benchmarks/ directory and accessible via the CLI or Python API.

How do I choose between running benchmarks via CLI versus Python API?

Use the CLI (examples/run_benchmark.py) for one-off evaluations or shell-scripted automation. Use the Python API when you need to integrate benchmarking into larger applications, customize evaluation logic, or process results programmatically within a Python environment.

What metrics does the benchmarking framework calculate?

The system calculates accuracy (percentage of correct answers) and average execution time via the calculate_metrics() function. When using optimize_parameters(), you can specify metric_weights to balance these factors according to your quality-versus-speed priorities.

How does the automated optimization determine the best configuration?

The optimize_parameters() function performs a Bayesian-style search over your defined param_space, running the benchmark for each sampled configuration. It scores trials using the weighted combination of metrics you specify (typically accuracy and speed), then returns the parameter set that achieves the highest composite score after the specified number of trials.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →