ADR Benchmark summary.json and result.json File Structure Guide
The ADR benchmark generates two JSON files: summary.json containing run-level aggregate statistics, and result.json containing per-task performance metrics with fields like run_id, success_rate, actual_speedup, task_id, mcp_ratio, and speedup.
The Uber ADR (Application Detection and Response) repository includes a comprehensive benchmark suite that measures performance characteristics of instrumented versus non-instrumented code paths. After execution, the benchmark writes structured JSON output that enables systematic analysis of MCP (Method Call Profiling) overhead and concurrency efficiency. Understanding these file formats is essential for parsing benchmark results and integrating ADR performance data into CI/CD pipelines.
Where summary.json and result.json Are Generated
The benchmark output files are produced by two core components in the ADR detection system.
Detection/main_benchmark.py— Orchestrates the full benchmark run and writessummary.jsonvia the_generate_summary()method at line 1152Detection/benchmark/benchmark_pack.py— Executes individual benchmark tasks and writes per-taskresult.jsonfiles
The _generate_summary() method aggregates statistics across all completed tasks to produce the run-level summary, while each task's results are serialized independently for granular analysis.
summary.json Structure
The summary.json file captures aggregate statistics for the entire benchmark execution. It serves as the primary artifact for comparing benchmark runs and tracking performance trends over time.
Core Fields in summary.json
| Field | Type | Description |
|---|---|---|
run_id |
string | Unique identifier for the benchmark run |
benchmark_type |
string | Category such as "baseline" or "replication" |
wall_clock_time |
float | Total elapsed seconds for the complete run |
total_tasks |
int | Number of benchmark tasks executed |
successful |
int | Tasks that completed without error |
failed |
int | Tasks that raised an exception |
success_rate |
float | Ratio of successful / total_tasks (range 0–1) |
actual_speedup |
float | Measured speed-up compared to baseline |
concurrency_efficiency |
float | Calculated as actual_speedup / number_of_workers |
overall_mcp_ratio |
float | Proportion of MCP-instrumented calls across all tasks |
total_mcp_calls |
int | Total MCP-instrumented calls across all tasks |
total_non_mcp_calls |
int | Calls that bypassed MCP instrumentation |
timestamp |
string | ISO-8601 timestamp when the run completed |
The file is written with json.dump(..., indent=2) for human readability.
Loading and Parsing summary.json
import json
from pathlib import Path
summary_path = Path("benchmark_output") / "summary.json"
with summary_path.open() as f:
summary = json.load(f)
print(f"Run ID: {summary['run_id']}")
print(f"Type: {summary['benchmark_type']}")
print(f"Tasks: {summary['successful']}/{summary['total_tasks']} succeeded")
print(f"Success rate: {summary['success_rate']:.1%}")
print(f"Actual speedup: {summary['actual_speedup']:.2f}×")
print(f"MCP coverage: {summary['overall_mcp_ratio']:.1%}")
result.json Structure
Each benchmark task produces its own result.json file, typically located in a task-specific subdirectory. These files enable drill-down analysis of individual performance characteristics.
Core Fields in result.json
| Field | Type | Description |
|---|---|---|
task_id |
string | Unique identifier for the specific task |
task_name |
string | Human-readable task description |
status |
string | "success" or "error" |
error_message |
string | Present only when status is "error" |
wall_time |
float | Wall-clock seconds the task consumed |
cpu_time |
float | CPU seconds consumed by the task |
mcp_calls |
int | Number of MCP-instrumented method calls |
non_mcp_calls |
int | Number of non-instrumented method calls |
mcp_ratio |
float | Calculated as mcp_calls / (mcp_calls + non_mcp_calls) |
speedup |
float | Task-level speed-up factor versus baseline |
metrics |
object | Additional per-task numeric data (latency, throughput, etc.) |
The metrics object provides extensibility for task-specific measurements without requiring schema changes.
Loading and Parsing result.json
import json
from pathlib import Path
# Example: loading a specific task result
task_path = Path("benchmark_output") / "task_001" / "result.json"
with task_path.open() as f:
result = json.load(f)
print(f"Task: {result['task_name']} ({result['task_id']})")
print(f"Status: {result['status']}")
if result['status'] == 'error':
print(f"Failed with: {result['error_message']}")
else:
print(f"Wall time: {result['wall_time']:.3f}s")
print(f"CPU time: {result['cpu_time']:.3f}s")
print(f"Speedup: {result['speedup']:.2f}×")
print(f"MCP ratio: {result['mcp_ratio']:.2%}")
# Access extended metrics if present
if result.get('metrics'):
print(f"Custom metrics: {result['metrics']}")
Key Relationships Between the Files
The summary.json and result.json files share derived values that must remain consistent.
actual_speedupinsummary.jsonrepresents the aggregate speed-up computed from individualspeedupvalues in eachresult.jsonoverall_mcp_ratioinsummary.jsonis the weighted average of per-taskmcp_ratiovalues, scaled by total call volumetotal_mcp_callsandtotal_non_mcp_callsinsummary.jsonare simple sums of corresponding fields across allresult.jsonfiles
Validation tests in Detection/tests/test_benchmark_pack.py verify that these aggregation relationships hold and that all required fields are present in generated output.
Working with Benchmark Output Programmatically
For automation and reporting pipelines, process both file types together:
import json
from pathlib import Path
from dataclasses import dataclass
from typing import List, Optional
@dataclass
class BenchmarkRun:
summary: dict
task_results: List[dict]
@property
def failed_tasks(self) -> List[dict]:
return [t for t in self.task_results if t['status'] == 'error']
@property
def average_task_speedup(self) -> float:
successful = [t['speedup'] for t in self.task_results
if t['status'] == 'success']
return sum(successful) / len(successful) if successful else 0.0
def load_benchmark_run(output_dir: Path) -> BenchmarkRun:
"""Load complete benchmark results from output directory."""
# Load summary
with (output_dir / "summary.json").open() as f:
summary = json.load(f)
# Load all task results
task_results = []
for task_dir in output_dir.glob("task_*/"):
result_file = task_dir / "result.json"
if result_file.exists():
with result_file.open() as f:
task_results.append(json.load(f))
return BenchmarkRun(summary=summary, task_results=task_results)
# Usage
run = load_benchmark_run(Path("benchmark_output/2024-01-15"))
print(f"Failed tasks: {len(run.failed_tasks)}")
print(f"Avg speedup: {run.average_task_speedup:.2f}×")
Summary
summary.jsoncontains run-level aggregates includingrun_id,success_rate,actual_speedup, andoverall_mcp_ratioresult.jsonfiles capture per-task details withtask_id,status,mcp_ratio, andspeedupfor individual analysis- Both files are generated in
Detection/main_benchmark.pyandDetection/benchmark/benchmark_pack.pyrespectively - The schema enables both human inspection (pretty-printed with 2-space indentation) and automated processing
- Validation tests ensure field presence and aggregation consistency between summary and task-level data
Frequently Asked Questions
How do I find all result.json files for a benchmark run?
Search recursively from the benchmark output directory. Each task typically writes to its own subdirectory following the pattern task_*/result.json. Use Path(output_dir).glob('**/result.json') in Python or find . -name 'result.json' in shell.
What does mcp_ratio measure in ADR benchmark results?
The mcp_ratio field indicates what proportion of method calls were captured by MCP (Method Call Profiling) instrumentation versus executed natively. A ratio near 1.0 means nearly all calls were profiled, while lower values indicate significant uninstrumented code paths that may affect accuracy.
Can I add custom fields to the metrics object in result.json?
Yes. The metrics object in result.json is designed for extensibility. Benchmark tasks can inject arbitrary key-value pairs for additional measurements like latency percentiles, memory usage, or throughput. These custom fields pass through aggregation unchanged and appear in programmatic access without modifying the core schema.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →