# How the `plot_paper_figures.py` Script Generates Figures for the ADR Benchmark

> Discover how the plot_paper_figures.py script in the uber/ADR repo generates publication-ready figures from benchmark results. Learn its data processing and plotting pipeline.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: how-to-guide
- Published: 2026-08-06

---

**[`plot_paper_figures.py`](https://github.com/uber/ADR/blob/main/plot_paper_figures.py) is a self-contained Python utility that transforms raw JSON benchmark results into publication-ready figures for the ADR (Attack Detection and Response) MLSys paper through a structured pipeline of data loading, cost recalculation, metric extraction, and modular plotting functions.**

The [`plot_paper_figures.py`](https://github.com/uber/ADR/blob/main/plot_paper_figures.py) script in the [uber/ADR](https://github.com/uber/ADR) repository serves as the main visualization engine for converting detector benchmark outputs into the figures presented in the research paper. Located at [`Detection/plot_paper_figures.py`](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py), this script handles everything from parsing command-line arguments to producing high-resolution PNG outputs at **300 DPI**.

## Command-Line Interface and Workflow Entry Point

The script begins execution in `main()` ([lines 15-25](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L15-L25)), where it establishes a flexible CLI using Python's `argparse` module. Users can specify:

- `--benchmark-dir`: Path to the directory containing `*_baseline_analysis.json` files produced by each detector run
- `--output-dir`: Destination folder for generated figures (customizable without code changes)
- `--detectors`: List of detector names to include in visualizations
- `--mcp-logs`: Custom location for ADR reasoning debug logs

This design separates data locations from plotting logic, enabling the same script to process different experimental setups without modification.

## Loading and Cost Recalculation

The `load_and_recalculate()` function ([lines 87-106](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L87-L106)) implements a critical data enrichment step. It reads each detector's JSON analysis file and **recalculates cost per token** using an up-to-date pricing table rather than relying on potentially stale values from the original benchmark run.

The pricing model is defined in the `PRICING_MODELS` dictionary at the top of the file ([lines 36-69](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L36-L69)). This centralized table specifies per-million-token costs for every LLM evaluated in the benchmark (e.g., Claude variants, GPT-4, Llama models). Authors can adjust pricing to reflect current API rates without rerunning the entire detection pipeline—the script simply updates the computed costs at visualization time.

## Metric Extraction and Normalization

Before any plotting occurs, `extract_metrics()` ([lines 19-27](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L19-L27)) transforms disparate JSON schemas into uniform NumPy arrays. This function collects:

- **Predictions** and **ground-truth labels** for classification accuracy
- **Confidence scores** for threshold analysis
- **Per-task cost** in dollars
- **Latency** values converted to milliseconds

By consolidating these metrics into canonical arrays, all downstream visualization functions operate on consistent data structures regardless of detector-specific output formats.

## Figure 1: Core Benchmark Visualizations

The script generates four complementary figures that characterize detector performance across key dimensions.

### Threat-Technique Detection Bar Chart

`plot_threat_technique_bar()` ([lines 93-140](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L93-L140)) produces **Figure 1a**, mapping each threat technique to its high-level MITRE ATT&CK tactic. The function tallies **true positives** and **false negatives** per detector, then renders a grouped bar chart showing detection rates across tactics. This visualization reveals which attack categories each detector handles effectively.

### Latency Cumulative Distribution

`plot_latency_cdf()` ([lines 45-100](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L45-L100)) generates **Figure 1b**, displaying the cumulative distribution of inference latencies for every detector. The implementation uses a **log-scale x-axis** for readability across the wide latency range (milliseconds to tens of seconds). ADR's curve appears as a **thicker line** for emphasis, making its real-time performance immediately apparent.

### Cost-Recall Scatter with Pareto Frontier

`plot_cost_recall()` ([lines 4-31](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L4-L31)) creates **Figure 1c**, a bubble-size-scaled scatter plot where each detector's position reflects its **precision**, **recall**, **F1-score**, and **average cost per task**. Larger bubbles indicate higher F1 scores. The function also annotates the **Pareto frontier**—detectors that cannot improve one metric without worsening another—clarifying which solutions dominate in the cost-performance tradeoff.

### Confusion Matrix Comparison

`plot_confusion_matrix_comparison()` ([lines 43-84](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L43-L84)) draws **Figure 1d**, summarizing **TP/FP/FN/TN counts** as side-by-side stacked bars. ADR's **zero false positives** receive distinct coloring, highlighting its precision advantage for security-critical deployments where false alarms carry operational costs.

## Figure 2: MCP Provider Usage Analysis

`plot_mcp_usage_cdf()` ([lines 27-64](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L27-L64)) generates **Figure 2** by inspecting ADR's debug logs stored in the MCP-log directory. The function parses `*_claude_output.json` and `*_triage_only.json` files to count invocations of each **Model Context Protocol (MCP)** server: source-code analysis, threat-intelligence lookup, and policy verification.

The resulting grouped bar chart separates **"triage-only"** tasks (fast rejection of benign inputs) from **"reasoning-agent"** tasks (full analysis of suspicious events), illuminating how ADR's modular architecture allocates computational resources.

## Figure 3: Ablation Study Visualizations

The script produces three figures analyzing ADR's design decisions through controlled ablations—detector variants with specific components disabled.

### Triage Advantage Analysis

`plot_triage_advantage()` ([lines 59-84](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L59-L84)) creates **Figure 3a**, a bar chart comparing **precision**, **cost**, and **latency** across configurations: full ADR, ADR without triage, and baseline detectors. This quantifies the benefit of ADR's two-stage architecture.

### MCP Necessity Heatmap

`plot_mcp_necessity()` ([lines 45-63](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L45-L63)) generates **Figure 3b**, a heatmap of **precision/recall/F1** for each MCP-server ablation (e.g., ADR without source-code access, without threat-intel, without policy checking). Color intensity immediately reveals which external tools contribute most to detection accuracy.

### Efficiency Dual-Axis Chart

`plot_efficiency_analysis()` ([lines 23-71](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L23-L71)) produces **Figure 3c**, combining **cost-per-true-positive** (bars) and **average latency** (line) on dual axes. This figure directly addresses deployment concerns: how much must be spent to catch each real attack, and how quickly?

## Execution Examples

### Basic invocation with defaults

```bash
python3 Detection/plot_paper_figures.py \
    --benchmark-dir benchmark/adr_bench_20251017_151604 \
    --output-dir figs

```

This uses the default detector list (`adr`, `llamafirewall`) and the standard MCP-log location.

### Custom detector comparison

```bash
python3 Detection/plot_paper_figures.py \
    --benchmark-dir my_bench \
    --detectors adr llamafirewall my-custom-detector \
    --output-dir paper/figs

```

Provided [`my-custom-detector_baseline_analysis.json`](https://github.com/uber/ADR/blob/main/my-custom-detector_baseline_analysis.json) exists in the benchmark directory, the script will include it in all applicable figures with consistent styling.

### Relocated reasoning logs

```bash
python3 Detection/plot_paper_figures.py \
    --benchmark-dir benchmark/adr_bench_20251017_151604 \
    --mcp-logs /path/to/archived/debug_logs \
    --output-dir figs

```

Useful when analyzing historical ADR runs stored outside the repository structure.

### Updating pricing without rerunning benchmarks

```python

# Edit Detection/plot_paper_figures.py, lines 36-69

PRICING_MODELS = {
    "claude-3-opus-20240229": {
        "input": 15.00,   # $ per million tokens

        "output": 75.00,
    },
    # ... update values as needed

}

```

Then rerun the script. The `load_and_recalculate()` function applies new pricing immediately since it reads `PRICING_MODELS` at runtime.

## Reproducibility Mechanisms

Two global configurations ensure consistent figure aesthetics across executions:

- **`ORDERED_DETECTORS`** ([lines 74-78](https://github.com/uber/ADR/blob/main/Detection/plot_paper_figures.py#L74-L78)): Enforces canonical ordering (ADR first, then LlamaFirewall, etc.) for legends and subplot arrangements
- **`DETECTOR_LABELS`**: Maps internal detector names to display-ready strings with proper capitalization

These lists guarantee that regenerated figures match the paper's visual identity even as new benchmarks are processed.

## Summary

- **[`plot_paper_figures.py`](https://github.com/uber/ADR/blob/main/plot_paper_figures.py)** transforms raw benchmark JSON into publication-ready figures through a **pure-data pipeline**: load → recalculate costs → extract metrics → plot
- **Modular architecture**: Each figure has a dedicated function handling layout, color palette, axis styling, and 300 DPI PNG export, enabling testing and notebook reuse
- **Runtime flexibility**: Command-line arguments control input paths, detector selection, and MCP-log locations without code modification
- **Cost independence**: The `PRICING_MODELS` dictionary decouples economic analysis from benchmark execution, allowing price updates without rerunning detectors
- **Comprehensive coverage**: Figures 1a-d characterize core performance; Figure 2 analyzes resource usage; Figures 3a-c validate architectural decisions through ablation

## Frequently Asked Questions

### What input files does [`plot_paper_figures.py`](https://github.com/uber/ADR/blob/main/plot_paper_figures.py) require?

The script requires `*_baseline_analysis.json` files in the specified benchmark directory—one per detector containing predictions, labels, token counts, and timestamps. For Figure 2, it additionally needs ADR reasoning logs (`*_claude_output.json`, `*_triage_only.json`) from the MCP-log directory. All paths are configurable via command-line flags.

### How does the script handle different LLM pricing over time?

Rather than embedding costs in benchmark outputs, the script recalculates expenses at visualization time using the `PRICING_MODELS` dictionary. Update the per-million-token rates in lines 36-69, rerun the script, and figures reflect current pricing immediately. This design accommodates API rate changes without expensive benchmark re-execution.

### Can I add a new detector to the visualizations?

Yes. Place a properly formatted `<detector_name>_baseline_analysis.json` file in the benchmark directory, then include the detector name in the `--detectors` argument. The script will incorporate it into all applicable figures using `extract_metrics()`'s uniform data representation. For consistent styling, consider adding the detector to `ORDERED_DETECTORS` and `DETECTOR_LABELS`.

### Why does the latency CDF use a log scale?

The benchmark captures detectors with dramatically different latency profiles—some operating in milliseconds, others requiring tens of seconds. `plot_latency_cdf()` employs a log-scale x-axis (implemented via `matplotlib`) to visualize this 1000×+ range without compressing fast detectors into an unreadable left edge. ADR's curve receives thicker line weight for emphasis.