How the `plot_paper_figures.py` Script Generates Figures for the ADR Benchmark
plot_paper_figures.py is a self-contained Python utility that transforms raw JSON benchmark results into publication-ready figures for the ADR (Attack Detection and Response) MLSys paper through a structured pipeline of data loading, cost recalculation, metric extraction, and modular plotting functions.
The plot_paper_figures.py script in the uber/ADR repository serves as the main visualization engine for converting detector benchmark outputs into the figures presented in the research paper. Located at Detection/plot_paper_figures.py, this script handles everything from parsing command-line arguments to producing high-resolution PNG outputs at 300 DPI.
Command-Line Interface and Workflow Entry Point
The script begins execution in main() (lines 15-25), where it establishes a flexible CLI using Python's argparse module. Users can specify:
--benchmark-dir: Path to the directory containing*_baseline_analysis.jsonfiles produced by each detector run--output-dir: Destination folder for generated figures (customizable without code changes)--detectors: List of detector names to include in visualizations--mcp-logs: Custom location for ADR reasoning debug logs
This design separates data locations from plotting logic, enabling the same script to process different experimental setups without modification.
Loading and Cost Recalculation
The load_and_recalculate() function (lines 87-106) implements a critical data enrichment step. It reads each detector's JSON analysis file and recalculates cost per token using an up-to-date pricing table rather than relying on potentially stale values from the original benchmark run.
The pricing model is defined in the PRICING_MODELS dictionary at the top of the file (lines 36-69). This centralized table specifies per-million-token costs for every LLM evaluated in the benchmark (e.g., Claude variants, GPT-4, Llama models). Authors can adjust pricing to reflect current API rates without rerunning the entire detection pipeline—the script simply updates the computed costs at visualization time.
Metric Extraction and Normalization
Before any plotting occurs, extract_metrics() (lines 19-27) transforms disparate JSON schemas into uniform NumPy arrays. This function collects:
- Predictions and ground-truth labels for classification accuracy
- Confidence scores for threshold analysis
- Per-task cost in dollars
- Latency values converted to milliseconds
By consolidating these metrics into canonical arrays, all downstream visualization functions operate on consistent data structures regardless of detector-specific output formats.
Figure 1: Core Benchmark Visualizations
The script generates four complementary figures that characterize detector performance across key dimensions.
Threat-Technique Detection Bar Chart
plot_threat_technique_bar() (lines 93-140) produces Figure 1a, mapping each threat technique to its high-level MITRE ATT&CK tactic. The function tallies true positives and false negatives per detector, then renders a grouped bar chart showing detection rates across tactics. This visualization reveals which attack categories each detector handles effectively.
Latency Cumulative Distribution
plot_latency_cdf() (lines 45-100) generates Figure 1b, displaying the cumulative distribution of inference latencies for every detector. The implementation uses a log-scale x-axis for readability across the wide latency range (milliseconds to tens of seconds). ADR's curve appears as a thicker line for emphasis, making its real-time performance immediately apparent.
Cost-Recall Scatter with Pareto Frontier
plot_cost_recall() (lines 4-31) creates Figure 1c, a bubble-size-scaled scatter plot where each detector's position reflects its precision, recall, F1-score, and average cost per task. Larger bubbles indicate higher F1 scores. The function also annotates the Pareto frontier—detectors that cannot improve one metric without worsening another—clarifying which solutions dominate in the cost-performance tradeoff.
Confusion Matrix Comparison
plot_confusion_matrix_comparison() (lines 43-84) draws Figure 1d, summarizing TP/FP/FN/TN counts as side-by-side stacked bars. ADR's zero false positives receive distinct coloring, highlighting its precision advantage for security-critical deployments where false alarms carry operational costs.
Figure 2: MCP Provider Usage Analysis
plot_mcp_usage_cdf() (lines 27-64) generates Figure 2 by inspecting ADR's debug logs stored in the MCP-log directory. The function parses *_claude_output.json and *_triage_only.json files to count invocations of each Model Context Protocol (MCP) server: source-code analysis, threat-intelligence lookup, and policy verification.
The resulting grouped bar chart separates "triage-only" tasks (fast rejection of benign inputs) from "reasoning-agent" tasks (full analysis of suspicious events), illuminating how ADR's modular architecture allocates computational resources.
Figure 3: Ablation Study Visualizations
The script produces three figures analyzing ADR's design decisions through controlled ablations—detector variants with specific components disabled.
Triage Advantage Analysis
plot_triage_advantage() (lines 59-84) creates Figure 3a, a bar chart comparing precision, cost, and latency across configurations: full ADR, ADR without triage, and baseline detectors. This quantifies the benefit of ADR's two-stage architecture.
MCP Necessity Heatmap
plot_mcp_necessity() (lines 45-63) generates Figure 3b, a heatmap of precision/recall/F1 for each MCP-server ablation (e.g., ADR without source-code access, without threat-intel, without policy checking). Color intensity immediately reveals which external tools contribute most to detection accuracy.
Efficiency Dual-Axis Chart
plot_efficiency_analysis() (lines 23-71) produces Figure 3c, combining cost-per-true-positive (bars) and average latency (line) on dual axes. This figure directly addresses deployment concerns: how much must be spent to catch each real attack, and how quickly?
Execution Examples
Basic invocation with defaults
python3 Detection/plot_paper_figures.py \
--benchmark-dir benchmark/adr_bench_20251017_151604 \
--output-dir figs
This uses the default detector list (adr, llamafirewall) and the standard MCP-log location.
Custom detector comparison
python3 Detection/plot_paper_figures.py \
--benchmark-dir my_bench \
--detectors adr llamafirewall my-custom-detector \
--output-dir paper/figs
Provided my-custom-detector_baseline_analysis.json exists in the benchmark directory, the script will include it in all applicable figures with consistent styling.
Relocated reasoning logs
python3 Detection/plot_paper_figures.py \
--benchmark-dir benchmark/adr_bench_20251017_151604 \
--mcp-logs /path/to/archived/debug_logs \
--output-dir figs
Useful when analyzing historical ADR runs stored outside the repository structure.
Updating pricing without rerunning benchmarks
# Edit Detection/plot_paper_figures.py, lines 36-69
PRICING_MODELS = {
"claude-3-opus-20240229": {
"input": 15.00, # $ per million tokens
"output": 75.00,
},
# ... update values as needed
}
Then rerun the script. The load_and_recalculate() function applies new pricing immediately since it reads PRICING_MODELS at runtime.
Reproducibility Mechanisms
Two global configurations ensure consistent figure aesthetics across executions:
ORDERED_DETECTORS(lines 74-78): Enforces canonical ordering (ADR first, then LlamaFirewall, etc.) for legends and subplot arrangementsDETECTOR_LABELS: Maps internal detector names to display-ready strings with proper capitalization
These lists guarantee that regenerated figures match the paper's visual identity even as new benchmarks are processed.
Summary
plot_paper_figures.pytransforms raw benchmark JSON into publication-ready figures through a pure-data pipeline: load → recalculate costs → extract metrics → plot- Modular architecture: Each figure has a dedicated function handling layout, color palette, axis styling, and 300 DPI PNG export, enabling testing and notebook reuse
- Runtime flexibility: Command-line arguments control input paths, detector selection, and MCP-log locations without code modification
- Cost independence: The
PRICING_MODELSdictionary decouples economic analysis from benchmark execution, allowing price updates without rerunning detectors - Comprehensive coverage: Figures 1a-d characterize core performance; Figure 2 analyzes resource usage; Figures 3a-c validate architectural decisions through ablation
Frequently Asked Questions
What input files does plot_paper_figures.py require?
The script requires *_baseline_analysis.json files in the specified benchmark directory—one per detector containing predictions, labels, token counts, and timestamps. For Figure 2, it additionally needs ADR reasoning logs (*_claude_output.json, *_triage_only.json) from the MCP-log directory. All paths are configurable via command-line flags.
How does the script handle different LLM pricing over time?
Rather than embedding costs in benchmark outputs, the script recalculates expenses at visualization time using the PRICING_MODELS dictionary. Update the per-million-token rates in lines 36-69, rerun the script, and figures reflect current pricing immediately. This design accommodates API rate changes without expensive benchmark re-execution.
Can I add a new detector to the visualizations?
Yes. Place a properly formatted <detector_name>_baseline_analysis.json file in the benchmark directory, then include the detector name in the --detectors argument. The script will incorporate it into all applicable figures using extract_metrics()'s uniform data representation. For consistent styling, consider adding the detector to ORDERED_DETECTORS and DETECTOR_LABELS.
Why does the latency CDF use a log scale?
The benchmark captures detectors with dramatically different latency profiles—some operating in milliseconds, others requiring tens of seconds. plot_latency_cdf() employs a log-scale x-axis (implemented via matplotlib) to visualize this 1000×+ range without compressing fast detectors into an unreadable left edge. ADR's curve receives thicker line weight for emphasis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →