# ADR Benchmark summary.json and result.json File Structure Guide

> Understand the ADR benchmark summary.json and result.json file structures. Learn about run-level aggregates and per-task performance metrics including success rate and speedup.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: how-to-guide
- Published: 2026-08-06

---

**The ADR benchmark generates two JSON files: [`summary.json`](https://github.com/uber/ADR/blob/main/summary.json) containing run-level aggregate statistics, and [`result.json`](https://github.com/uber/ADR/blob/main/result.json) containing per-task performance metrics with fields like `run_id`, `success_rate`, `actual_speedup`, `task_id`, `mcp_ratio`, and `speedup`.**

The Uber ADR (Application Detection and Response) repository includes a comprehensive benchmark suite that measures performance characteristics of instrumented versus non-instrumented code paths. After execution, the benchmark writes structured JSON output that enables systematic analysis of MCP (Method Call Profiling) overhead and concurrency efficiency. Understanding these file formats is essential for parsing benchmark results and integrating ADR performance data into CI/CD pipelines.

## Where summary.json and result.json Are Generated

The benchmark output files are produced by two core components in the ADR detection system.

- **[`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py)** — Orchestrates the full benchmark run and writes [`summary.json`](https://github.com/uber/ADR/blob/main/summary.json) via the `_generate_summary()` method at line 1152
- **[`Detection/benchmark/benchmark_pack.py`](https://github.com/uber/ADR/blob/main/Detection/benchmark/benchmark_pack.py)** — Executes individual benchmark tasks and writes per-task [`result.json`](https://github.com/uber/ADR/blob/main/result.json) files

The `_generate_summary()` method aggregates statistics across all completed tasks to produce the run-level summary, while each task's results are serialized independently for granular analysis.

## summary.json Structure

The [`summary.json`](https://github.com/uber/ADR/blob/main/summary.json) file captures aggregate statistics for the entire benchmark execution. It serves as the primary artifact for comparing benchmark runs and tracking performance trends over time.

### Core Fields in summary.json

| Field | Type | Description |
|-------|------|-------------|
| `run_id` | string | Unique identifier for the benchmark run |
| `benchmark_type` | string | Category such as `"baseline"` or `"replication"` |
| `wall_clock_time` | float | Total elapsed seconds for the complete run |
| `total_tasks` | int | Number of benchmark tasks executed |
| `successful` | int | Tasks that completed without error |
| `failed` | int | Tasks that raised an exception |
| `success_rate` | float | Ratio of `successful / total_tasks` (range 0–1) |
| `actual_speedup` | float | Measured speed-up compared to baseline |
| `concurrency_efficiency` | float | Calculated as `actual_speedup / number_of_workers` |
| `overall_mcp_ratio` | float | Proportion of MCP-instrumented calls across all tasks |
| `total_mcp_calls` | int | Total MCP-instrumented calls across all tasks |
| `total_non_mcp_calls` | int | Calls that bypassed MCP instrumentation |
| `timestamp` | string | ISO-8601 timestamp when the run completed |

The file is written with `json.dump(..., indent=2)` for human readability.

### Loading and Parsing summary.json

```python
import json
from pathlib import Path

summary_path = Path("benchmark_output") / "summary.json"

with summary_path.open() as f:
    summary = json.load(f)

print(f"Run ID: {summary['run_id']}")
print(f"Type: {summary['benchmark_type']}")
print(f"Tasks: {summary['successful']}/{summary['total_tasks']} succeeded")
print(f"Success rate: {summary['success_rate']:.1%}")
print(f"Actual speedup: {summary['actual_speedup']:.2f}×")
print(f"MCP coverage: {summary['overall_mcp_ratio']:.1%}")

```

## result.json Structure

Each benchmark task produces its own [`result.json`](https://github.com/uber/ADR/blob/main/result.json) file, typically located in a task-specific subdirectory. These files enable drill-down analysis of individual performance characteristics.

### Core Fields in result.json

| Field | Type | Description |
|-------|------|-------------|
| `task_id` | string | Unique identifier for the specific task |
| `task_name` | string | Human-readable task description |
| `status` | string | `"success"` or `"error"` |
| `error_message` | string | Present only when `status` is `"error"` |
| `wall_time` | float | Wall-clock seconds the task consumed |
| `cpu_time` | float | CPU seconds consumed by the task |
| `mcp_calls` | int | Number of MCP-instrumented method calls |
| `non_mcp_calls` | int | Number of non-instrumented method calls |
| `mcp_ratio` | float | Calculated as `mcp_calls / (mcp_calls + non_mcp_calls)` |
| `speedup` | float | Task-level speed-up factor versus baseline |
| `metrics` | object | Additional per-task numeric data (latency, throughput, etc.) |

The `metrics` object provides extensibility for task-specific measurements without requiring schema changes.

### Loading and Parsing result.json

```python
import json
from pathlib import Path

# Example: loading a specific task result

task_path = Path("benchmark_output") / "task_001" / "result.json"

with task_path.open() as f:
    result = json.load(f)

print(f"Task: {result['task_name']} ({result['task_id']})")
print(f"Status: {result['status']}")

if result['status'] == 'error':
    print(f"Failed with: {result['error_message']}")
else:
    print(f"Wall time: {result['wall_time']:.3f}s")
    print(f"CPU time: {result['cpu_time']:.3f}s")
    print(f"Speedup: {result['speedup']:.2f}×")
    print(f"MCP ratio: {result['mcp_ratio']:.2%}")
    
    # Access extended metrics if present

    if result.get('metrics'):
        print(f"Custom metrics: {result['metrics']}")

```

## Key Relationships Between the Files

The [`summary.json`](https://github.com/uber/ADR/blob/main/summary.json) and [`result.json`](https://github.com/uber/ADR/blob/main/result.json) files share derived values that must remain consistent.

- **`actual_speedup`** in [`summary.json`](https://github.com/uber/ADR/blob/main/summary.json) represents the aggregate speed-up computed from individual `speedup` values in each [`result.json`](https://github.com/uber/ADR/blob/main/result.json)
- **`overall_mcp_ratio`** in [`summary.json`](https://github.com/uber/ADR/blob/main/summary.json) is the weighted average of per-task `mcp_ratio` values, scaled by total call volume
- **`total_mcp_calls`** and **`total_non_mcp_calls`** in [`summary.json`](https://github.com/uber/ADR/blob/main/summary.json) are simple sums of corresponding fields across all [`result.json`](https://github.com/uber/ADR/blob/main/result.json) files

Validation tests in [`Detection/tests/test_benchmark_pack.py`](https://github.com/uber/ADR/blob/main/Detection/tests/test_benchmark_pack.py) verify that these aggregation relationships hold and that all required fields are present in generated output.

## Working with Benchmark Output Programmatically

For automation and reporting pipelines, process both file types together:

```python
import json
from pathlib import Path
from dataclasses import dataclass
from typing import List, Optional

@dataclass
class BenchmarkRun:
    summary: dict
    task_results: List[dict]
    
    @property
    def failed_tasks(self) -> List[dict]:
        return [t for t in self.task_results if t['status'] == 'error']
    
    @property
    def average_task_speedup(self) -> float:
        successful = [t['speedup'] for t in self.task_results 
                     if t['status'] == 'success']
        return sum(successful) / len(successful) if successful else 0.0

def load_benchmark_run(output_dir: Path) -> BenchmarkRun:
    """Load complete benchmark results from output directory."""
    
    # Load summary

    with (output_dir / "summary.json").open() as f:
        summary = json.load(f)
    
    # Load all task results

    task_results = []
    for task_dir in output_dir.glob("task_*/"):
        result_file = task_dir / "result.json"
        if result_file.exists():
            with result_file.open() as f:
                task_results.append(json.load(f))
    
    return BenchmarkRun(summary=summary, task_results=task_results)

# Usage

run = load_benchmark_run(Path("benchmark_output/2024-01-15"))
print(f"Failed tasks: {len(run.failed_tasks)}")
print(f"Avg speedup: {run.average_task_speedup:.2f}×")

```

## Summary

- **[`summary.json`](https://github.com/uber/ADR/blob/main/summary.json)** contains run-level aggregates including `run_id`, `success_rate`, `actual_speedup`, and `overall_mcp_ratio`
- **[`result.json`](https://github.com/uber/ADR/blob/main/result.json)** files capture per-task details with `task_id`, `status`, `mcp_ratio`, and `speedup` for individual analysis
- Both files are generated in [`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py) and [`Detection/benchmark/benchmark_pack.py`](https://github.com/uber/ADR/blob/main/Detection/benchmark/benchmark_pack.py) respectively
- The schema enables both human inspection (pretty-printed with 2-space indentation) and automated processing
- Validation tests ensure field presence and aggregation consistency between summary and task-level data

## Frequently Asked Questions

### How do I find all result.json files for a benchmark run?

Search recursively from the benchmark output directory. Each task typically writes to its own subdirectory following the pattern `task_*/result.json`. Use `Path(output_dir).glob('**/result.json')` in Python or `find . -name 'result.json'` in shell.

### What does mcp_ratio measure in ADR benchmark results?

The `mcp_ratio` field indicates what proportion of method calls were captured by MCP (Method Call Profiling) instrumentation versus executed natively. A ratio near 1.0 means nearly all calls were profiled, while lower values indicate significant uninstrumented code paths that may affect accuracy.

### Can I add custom fields to the metrics object in result.json?

Yes. The `metrics` object in [`result.json`](https://github.com/uber/ADR/blob/main/result.json) is designed for extensibility. Benchmark tasks can inject arbitrary key-value pairs for additional measurements like latency percentiles, memory usage, or throughput. These custom fields pass through aggregation unchanged and appear in programmatic access without modifying the core schema.