Research Strategies for Benchmark Optimization with BrowseComp and SimpleQA
Local Deep Research provides five categorical search strategies—iterdrag, standard, rapid, parallel, and source_based—that can be optimized across BrowseComp and SimpleQA benchmarks using Optuna-based hyperparameter tuning.
The learningcircuit/local-deep-research repository ships a fully-featured benchmark-optimization subsystem built on Optuna that treats research strategies as configurable hyperparameters. This system enables automated tuning of search behaviors against BrowseComp (multi-step web-search and synthesis tasks) and SimpleQA (classic question-answering tests) through weighted benchmark combinations and metric objectives.
Available Research Strategies for Optimization
The optimization engine defines search strategies as categorical hyperparameters in the default parameter space of OptunaOptimizer. According to src/local_deep_research/benchmarks/optimization/optuna_optimizer.py (lines 61-65), the system accepts five distinct strategies:
"search_strategy": {
"type": "categorical",
"choices": [
"iterdrag",
"standard",
"rapid",
"parallel",
"source_based",
],
},
Each strategy exhibits different trade-offs between quality, speed, and resource efficiency:
- iterdrag: Iterative drag-and-drop style search that prioritizes high-quality results at the cost of slower execution.
- standard: Baseline multi-turn search providing balanced behavior without aggressive optimizations.
- rapid: Aggressive early-stopping strategy that favors speed over comprehensive coverage.
- parallel: Fires several search workers simultaneously to balance speed and quality through concurrent execution.
- source_based: Uses source-ranking heuristics to deliver medium quality with improved speed.
These strategy names are accepted wherever a param_space dictionary defines a "search_strategy" entry.
Configuring BrowseComp and SimpleQA Weights
The system evaluates performance against both benchmarks using the benchmark_weights argument, allowing you to tune optimization toward specific task types. BrowseComp targets complex "browse-and-compare" tasks requiring multi-step web search, while SimpleQA covers straightforward question-answering scenarios.
To bias the optimizer toward SimpleQA with a secondary BrowseComp component, specify weights such as:
benchmark_weights={"simpleqa": 0.6, "browsecomp": 0.4}
This weighting directly influences the composite score calculation in CompositeBenchmarkEvaluator (src/local_deep_research/benchmarks/evaluators.py), determining how much each benchmark contributes to the final optimization objective.
Optimization Goals and Metrics
OptunaOptimizer can be instructed to optimize any combination of three metrics via the metric_weights parameter:
- quality: Answer accuracy measured against benchmark reference answers.
- speed: Wall-clock time converted to a normalized 0-1 score.
- resource: Computational cost proxy (primarily used in efficiency-focused presets).
For quality-centric research, use {"quality": 0.9, "speed": 0.1}. For latency-sensitive applications, invert these weights to favor speed.
Implementation Examples
Basic Multi-Benchmark Optimization
The high-level API in src/local_deep_research/benchmarks/optimization/api.py provides the optimize_parameters function for custom parameter spaces:
from local_deep_research.benchmarks.optimization import optimize_parameters
param_space = {
"search_strategy": {
"type": "categorical",
"choices": ["iterdrag", "source_based"],
},
"iterations": {"type": "int", "low": 1, "high": 3, "step": 1},
"questions_per_iteration": {"type": "int", "low": 1, "high": 3, "step": 1},
}
best_params, best_score = optimize_parameters(
query="What are the latest developments in fusion energy research?",
param_space=param_space,
n_trials=20,
benchmark_weights={"simpleqa": 0.6, "browsecomp": 0.4},
metric_weights={"quality": 0.7, "speed": 0.3},
num_examples=200,
)
This example varies the search strategy between iterdrag and source_based while testing 1-3 iterations and questions per iteration across 200 examples.
Speed-Focused Optimization Helper
For rapid prototyping, use the predefined optimize_for_speed wrapper (defined in api.py lines 78-91):
from local_deep_research.benchmarks.optimization import optimize_for_speed
best_params, best_score = optimize_for_speed(
query="Explain the health benefits of the Mediterranean diet.",
benchmark_weights={"simpleqa": 1.0},
n_trials=15,
num_examples=100,
)
This helper automatically applies speed-biased metric weights ({"speed": 0.8, "quality": 0.2}) and a compact numeric parameter space optimized for latency reduction.
Full-Scale Strategy Benchmarking
The repository includes ready-to-run scripts for comprehensive strategy comparison. Execute the BrowseComp optimization demo to compare iterdrag versus source_based across quality-focused, speed-focused, balanced, and multi-benchmark experiments:
python -m examples.optimization.browsecomp_optimization
This script (located at examples/optimization/browsecomp_optimization.py) demonstrates large-scale evaluation with num_examples=500 for statistically significant results. For deeper strategy analysis, refer to examples/optimization/strategy_benchmark_plan.py, which implements head-to-head comparisons between specific research strategies.
Summary
- Five research strategies (iterdrag, standard, rapid, parallel, source_based) are available as categorical hyperparameters in
OptunaOptimizer. - BrowseComp and SimpleQA benchmarks can be weighted via
benchmark_weightsto customize optimization objectives toward multi-step search or direct question-answering. - Three metric dimensions (quality, speed, resource) allow fine-grained control over optimization goals through
metric_weights. - Core implementation resides in
src/local_deep_research/benchmarks/optimization/optuna_optimizer.pywith convenience wrappers inapi.py.
Frequently Asked Questions
What are the five research strategies available for benchmark optimization?
The five strategies defined in optuna_optimizer.py are iterdrag (high quality, slower), standard (baseline multi-turn), rapid (speed-focused with early stopping), parallel (concurrent workers), and source_based (heuristic ranking with medium quality). These are passed as categorical choices in the search_strategy parameter space.
How do I weight BrowseComp versus SimpleQA in an optimization run?
Supply the benchmark_weights dictionary when calling optimize_parameters or related helpers. For example, {"simpleqa": 0.6, "browsecomp": 0.4} assigns 60% influence to SimpleQA scores and 40% to BrowseComp. The CompositeBenchmarkEvaluator uses these weights to calculate the final composite score during the Optuna trial loop.
Can I optimize for speed instead of answer quality?
Yes. Pass metric_weights favoring speed, such as {"speed": 0.8, "quality": 0.2}, or use the optimize_for_speed convenience function which bundles these weights automatically. The speed metric converts wall-clock time to a normalized 0-1 score where higher values indicate faster execution.
Where is the core optimization logic implemented?
The core Optuna integration lives in src/local_deep_research/benchmarks/optimization/optuna_optimizer.py, which handles parameter space sampling, mini-benchmark execution (defaulting to 5 examples, configurable via num_examples), and visualization generation. High-level wrappers like optimize_parameters and optimize_for_speed are implemented in src/local_deep_research/benchmarks/optimization/api.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →