How to Run Benchmarks to Evaluate and Optimize Your Configuration in Local Deep Research
Run benchmarks using the CLI wrapper at examples/run_benchmark.py or the Python API in src/local_deep_research/api/benchmark_functions.py to measure accuracy and speed, then use optimize_parameters() for automated Bayesian-style tuning.
The learningcircuit/local-deep-research repository ships with a comprehensive benchmarking framework that lets you systematically run benchmarks to evaluate and optimize configuration parameters. Whether you need to compare search strategies or fine-tune iteration counts, the toolkit provides predefined evaluations like SimpleQA and BrowseComp alongside automated optimization routines.
Architecture Overview
The benchmarking system follows a layered architecture that separates user interfaces from execution logic. At the top level, examples/run_benchmark.py provides a CLI wrapper that parses arguments and dispatches to evaluation functions. The middle layer lives in src/local_deep_research/api/benchmark_functions.py, exposing high-level Python functions such as evaluate_simpleqa(), evaluate_browsecomp(), and compare_configurations(). These functions construct search and evaluation configs before invoking the low-level benchmark runners in src/local_deep_research/benchmarks/, which handle data loading, query execution via your selected tool (e.g., searxng), and metric calculation through calculate_metrics(). Finally, the optimization driver in examples/optimization/example_optimization.py demonstrates how to invoke optimize_parameters() from src/local_deep_research/benchmarks/optimization.py to perform automated parameter search.
The execution flow proceeds as follows:
- User invokes CLI or Python API to build
search_configandevaluation_config - Benchmark runner executes searches using the configured tool
- System collects answers and evaluates via LLM or human judgment
calculate_metrics()computes accuracy and average timegenerate_report()writes detailed markdown reports to the output directory
Running Benchmarks from the Command Line
The fastest way to run benchmarks to evaluate and optimize configuration settings is through the provided CLI wrapper. This script supports running individual benchmarks or comparing multiple configurations side-by-side.
To run a SimpleQA benchmark with specific search parameters:
python examples/run_benchmark.py --benchmark simpleqa --examples 10 \
--iterations 3 --questions 3 --search-tool searxng \
--output-dir my_benchmark_results
Switch the --benchmark argument to browsecomp for web-browsing evaluations or compare to test multiple configurations in a single run. All CLI arguments map directly to the parameters of the underlying Python API functions defined in benchmark_functions.py.
Benchmarking via the Python API
For programmatic integration, import functions from src/local_deep_research/api/benchmark_functions.py to embed benchmarking into your training or CI pipelines.
Running SimpleQA and BrowseComp
Use evaluate_simpleqa() or evaluate_browsecomp() to run a single benchmark configuration and retrieve detailed metrics:
from local_deep_research.api.benchmark_functions import evaluate_simpleqa
# Run SimpleQA with 2 iterations and 3 questions per iteration
results = evaluate_simpleqa(
num_examples=20,
search_iterations=2,
questions_per_iteration=3,
search_tool="searxng",
human_evaluation=False,
output_dir="benchmark_output"
)
print(f"Accuracy: {results['metrics']['accuracy']}")
print(f"Report saved to: {results['report_path']}")
Both functions return a dictionary containing metrics, total_examples, and the path to the generated markdown report.
Comparing Multiple Configurations
When you need to run benchmarks to evaluate and optimize configuration alternatives, use compare_configurations() to test multiple parameter sets against the same dataset:
from local_deep_research.api.benchmark_functions import compare_configurations
comparison = compare_configurations(
dataset_type="simpleqa",
num_examples=20,
configurations=[
{
"name": "Base",
"search_tool": "searxng",
"iterations": 1,
"questions_per_iteration": 3
},
{
"name": "DeepSearch",
"search_tool": "searxng",
"iterations": 3,
"questions_per_iteration": 3
},
{
"name": "BroadQuery",
"search_tool": "searxng",
"iterations": 1,
"questions_per_iteration": 5
}
],
output_dir="comparison_results"
)
print(f"Comparison report: {comparison['report_path']}")
This routine iterates over each configuration dictionary, runs the selected benchmark, and aggregates results into a single markdown comparison report.
Automated Parameter Optimization
For hands-free tuning, the framework provides Bayesian-style optimization through optimize_parameters() in src/local_deep_research/benchmarks/optimization.py. This function searches your defined parameter space to find the optimal balance between quality and speed.
The example script at examples/optimization/example_optimization.py demonstrates the complete workflow:
from local_deep_research.benchmarks.optimization import optimize_parameters
from pathlib import Path
import json
# Define the parameter space to search
param_space = {
"iterations": {"type": "int", "low": 1, "high": 3, "step": 1},
"questions_per_iteration": {"type": "int", "low": 1, "high": 5, "step": 1},
"search_strategy": {"type": "categorical", "choices": ["rapid", "thorough"]},
}
# Run optimization with 10 trials
best_params, best_score = optimize_parameters(
query="SimpleQA optimization run",
search_tool="searxng",
n_trials=10,
output_dir=str(Path("optimization_results") / "run1"),
param_space=param_space,
metric_weights={"quality": 0.6, "speed": 0.4}
)
print(f"Best configuration: {best_params}")
print(f"Best weighted score: {best_score}")
# Persist results
summary = {"best_params": best_params, "best_score": best_score}
Path("optimization_results/run1/summary.json").write_text(
json.dumps(summary, indent=2)
)
The optimizer internally repeats benchmarks for sampled parameter combinations, scoring each using your specified metric_weights, and returns the highest-scoring configuration.
Summary
- CLI Entry Point: Use
examples/run_benchmark.pyfor quick command-line evaluations of SimpleQA, BrowseComp, or configuration comparisons. - Python API: Import
evaluate_simpleqa(),evaluate_browsecomp(), andcompare_configurations()fromsrc/local_deep_research/api/benchmark_functions.pyfor programmatic control. - Optimization: Leverage
optimize_parameters()fromsrc/local_deep_research/benchmarks/optimization.pyto automatically search parameter spaces using weighted quality and speed metrics. - Output: All methods generate detailed markdown reports and JSON summaries containing accuracy, timing metrics, and configuration details.
- Integration: The architecture supports
searxngand other search tools, allowing you to evaluate different retrieval backends against the same benchmark datasets.
Frequently Asked Questions
What benchmarks are available in Local Deep Research?
The framework currently supports SimpleQA (factual question answering), BrowseComp (web browsing comprehension), and DeepSearch evaluations. These are implemented in the src/local_deep_research/benchmarks/ directory and accessible via the CLI or Python API.
How do I choose between running benchmarks via CLI versus Python API?
Use the CLI (examples/run_benchmark.py) for one-off evaluations or shell-scripted automation. Use the Python API when you need to integrate benchmarking into larger applications, customize evaluation logic, or process results programmatically within a Python environment.
What metrics does the benchmarking framework calculate?
The system calculates accuracy (percentage of correct answers) and average execution time via the calculate_metrics() function. When using optimize_parameters(), you can specify metric_weights to balance these factors according to your quality-versus-speed priorities.
How does the automated optimization determine the best configuration?
The optimize_parameters() function performs a Bayesian-style search over your defined param_space, running the benchmark for each sampled configuration. It scores trials using the weighted combination of metrics you specify (typically accuracy and speed), then returns the parameter set that achieves the highest composite score after the specified number of trials.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →