ADR Benchmark summary.json and result.json File Structure Guide

The ADR benchmark generates two JSON files: summary.json containing run-level aggregate statistics, and result.json containing per-task performance metrics with fields like run_id, success_rate, actual_speedup, task_id, mcp_ratio, and speedup.

The Uber ADR (Application Detection and Response) repository includes a comprehensive benchmark suite that measures performance characteristics of instrumented versus non-instrumented code paths. After execution, the benchmark writes structured JSON output that enables systematic analysis of MCP (Method Call Profiling) overhead and concurrency efficiency. Understanding these file formats is essential for parsing benchmark results and integrating ADR performance data into CI/CD pipelines.

Where summary.json and result.json Are Generated

The benchmark output files are produced by two core components in the ADR detection system.

The _generate_summary() method aggregates statistics across all completed tasks to produce the run-level summary, while each task's results are serialized independently for granular analysis.

summary.json Structure

The summary.json file captures aggregate statistics for the entire benchmark execution. It serves as the primary artifact for comparing benchmark runs and tracking performance trends over time.

Core Fields in summary.json

Field Type Description
run_id string Unique identifier for the benchmark run
benchmark_type string Category such as "baseline" or "replication"
wall_clock_time float Total elapsed seconds for the complete run
total_tasks int Number of benchmark tasks executed
successful int Tasks that completed without error
failed int Tasks that raised an exception
success_rate float Ratio of successful / total_tasks (range 0–1)
actual_speedup float Measured speed-up compared to baseline
concurrency_efficiency float Calculated as actual_speedup / number_of_workers
overall_mcp_ratio float Proportion of MCP-instrumented calls across all tasks
total_mcp_calls int Total MCP-instrumented calls across all tasks
total_non_mcp_calls int Calls that bypassed MCP instrumentation
timestamp string ISO-8601 timestamp when the run completed

The file is written with json.dump(..., indent=2) for human readability.

Loading and Parsing summary.json

import json
from pathlib import Path

summary_path = Path("benchmark_output") / "summary.json"

with summary_path.open() as f:
    summary = json.load(f)

print(f"Run ID: {summary['run_id']}")
print(f"Type: {summary['benchmark_type']}")
print(f"Tasks: {summary['successful']}/{summary['total_tasks']} succeeded")
print(f"Success rate: {summary['success_rate']:.1%}")
print(f"Actual speedup: {summary['actual_speedup']:.2f}×")
print(f"MCP coverage: {summary['overall_mcp_ratio']:.1%}")

result.json Structure

Each benchmark task produces its own result.json file, typically located in a task-specific subdirectory. These files enable drill-down analysis of individual performance characteristics.

Core Fields in result.json

Field Type Description
task_id string Unique identifier for the specific task
task_name string Human-readable task description
status string "success" or "error"
error_message string Present only when status is "error"
wall_time float Wall-clock seconds the task consumed
cpu_time float CPU seconds consumed by the task
mcp_calls int Number of MCP-instrumented method calls
non_mcp_calls int Number of non-instrumented method calls
mcp_ratio float Calculated as mcp_calls / (mcp_calls + non_mcp_calls)
speedup float Task-level speed-up factor versus baseline
metrics object Additional per-task numeric data (latency, throughput, etc.)

The metrics object provides extensibility for task-specific measurements without requiring schema changes.

Loading and Parsing result.json

import json
from pathlib import Path

# Example: loading a specific task result

task_path = Path("benchmark_output") / "task_001" / "result.json"

with task_path.open() as f:
    result = json.load(f)

print(f"Task: {result['task_name']} ({result['task_id']})")
print(f"Status: {result['status']}")

if result['status'] == 'error':
    print(f"Failed with: {result['error_message']}")
else:
    print(f"Wall time: {result['wall_time']:.3f}s")
    print(f"CPU time: {result['cpu_time']:.3f}s")
    print(f"Speedup: {result['speedup']:.2f}×")
    print(f"MCP ratio: {result['mcp_ratio']:.2%}")
    
    # Access extended metrics if present

    if result.get('metrics'):
        print(f"Custom metrics: {result['metrics']}")

Key Relationships Between the Files

The summary.json and result.json files share derived values that must remain consistent.

  • actual_speedup in summary.json represents the aggregate speed-up computed from individual speedup values in each result.json
  • overall_mcp_ratio in summary.json is the weighted average of per-task mcp_ratio values, scaled by total call volume
  • total_mcp_calls and total_non_mcp_calls in summary.json are simple sums of corresponding fields across all result.json files

Validation tests in Detection/tests/test_benchmark_pack.py verify that these aggregation relationships hold and that all required fields are present in generated output.

Working with Benchmark Output Programmatically

For automation and reporting pipelines, process both file types together:

import json
from pathlib import Path
from dataclasses import dataclass
from typing import List, Optional

@dataclass
class BenchmarkRun:
    summary: dict
    task_results: List[dict]
    
    @property
    def failed_tasks(self) -> List[dict]:
        return [t for t in self.task_results if t['status'] == 'error']
    
    @property
    def average_task_speedup(self) -> float:
        successful = [t['speedup'] for t in self.task_results 
                     if t['status'] == 'success']
        return sum(successful) / len(successful) if successful else 0.0

def load_benchmark_run(output_dir: Path) -> BenchmarkRun:
    """Load complete benchmark results from output directory."""
    
    # Load summary

    with (output_dir / "summary.json").open() as f:
        summary = json.load(f)
    
    # Load all task results

    task_results = []
    for task_dir in output_dir.glob("task_*/"):
        result_file = task_dir / "result.json"
        if result_file.exists():
            with result_file.open() as f:
                task_results.append(json.load(f))
    
    return BenchmarkRun(summary=summary, task_results=task_results)

# Usage

run = load_benchmark_run(Path("benchmark_output/2024-01-15"))
print(f"Failed tasks: {len(run.failed_tasks)}")
print(f"Avg speedup: {run.average_task_speedup:.2f}×")

Summary

  • summary.json contains run-level aggregates including run_id, success_rate, actual_speedup, and overall_mcp_ratio
  • result.json files capture per-task details with task_id, status, mcp_ratio, and speedup for individual analysis
  • Both files are generated in Detection/main_benchmark.py and Detection/benchmark/benchmark_pack.py respectively
  • The schema enables both human inspection (pretty-printed with 2-space indentation) and automated processing
  • Validation tests ensure field presence and aggregation consistency between summary and task-level data

Frequently Asked Questions

How do I find all result.json files for a benchmark run?

Search recursively from the benchmark output directory. Each task typically writes to its own subdirectory following the pattern task_*/result.json. Use Path(output_dir).glob('**/result.json') in Python or find . -name 'result.json' in shell.

What does mcp_ratio measure in ADR benchmark results?

The mcp_ratio field indicates what proportion of method calls were captured by MCP (Method Call Profiling) instrumentation versus executed natively. A ratio near 1.0 means nearly all calls were profiled, while lower values indicate significant uninstrumented code paths that may affect accuracy.

Can I add custom fields to the metrics object in result.json?

Yes. The metrics object in result.json is designed for extensibility. Benchmark tasks can inject arbitrary key-value pairs for additional measurements like latency percentiles, memory usage, or throughput. These custom fields pass through aggregation unchanged and appear in programmatic access without modifying the core schema.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →