Security Boundaries for Synthetic Benchmark Data in Uber ADR: 5 Protective Layers Explained

ADR implements multiple security boundaries—sandbox isolation, runtime mode checks, synthetic-only fixtures, policy enforcement, and guardrail modules—to ensure synthetic benchmark data never contaminates production environments.

Uber's Attack Detection and Response (ADR) system includes a dedicated benchmark suite for security research. The synthetic benchmark data used in this suite is architecturally separated from any operational systems through deliberate engineering safeguards. This article examines the five protective layers implemented in the uber/ADR repository and their specific code implementations.

Isolated Sandbox Execution

The benchmark executes exclusively within containerized or virtualized environments. The entry point at Detection/main_benchmark.py performs a mandatory runtime verification before processing any data.

The code checks for benchmark mode activation and aborts execution if the environment validation fails. This prevents accidental execution in production contexts where real credentials or live endpoints might be present.


# Build the sandbox image (Dockerfile is part of the repo)

docker build -t adr-benchmark -f Dockerfile .

# Run the benchmark; the script will verify it is in BENCHMARK MODE

docker run --rm adr-benchmark python -m Detection.main_benchmark \
    --config Detection/config_benchmark.yaml

When launched, the script emits an explicit warning that "all data is synthetic" before any processing begins, making the operational context unambiguous to operators and automated systems alike.

Synthetic-Only Fixtures

All benchmark data is fabricated. The fixtures contain:

  • Synthetic credentials — hardcoded strings designed to mimic authentication tokens without exposing real secrets
  • Prompt-injection payloads — test cases for adversarial input detection
  • Emulated MCP servers — vulnerable server simulations for controlled experimentation

The Detection/README.md explicitly documents this constraint, stating that "benchmark fixtures include synthetic credentials, prompt-injection payloads, and emulated vulnerable MCP servers." No real secrets or live endpoints are referenced anywhere in the benchmark codebase.

Explicit Documentation and Disclosure

ADR maintains multiple layers of documentation that reinforce the synthetic nature of benchmark materials:

  • docs/OPEN_SOURCE_REVIEW.md — Lists synthetic benchmark material and emphasizes that test data strings are intentionally fabricated
  • Detection/README.md — Identifies the benchmark as a research artifact with strict isolation requirements
  • Inline code comments — Warn that data is synthetic and must never be exposed to real networks

This documentation strategy ensures that any human review or automated scanning of the repository correctly categorizes the benchmark data as non-sensitive test material.

Policy Enforcement Layer

The Detection/context_providers/data/policy_store.yaml file encodes operational rules that prevent synthetic data from being misrepresented as authentic.

The policy mandates:

  1. Verification of data integrity — Tools must validate authenticity markers before processing
  2. Prohibition on credential fabrication — Hard-coded test credentials cannot be returned or presented as real

This policy file provides a declarative layer of protection that automated tools can evaluate and enforce independently of runtime code paths.

Runtime Guardrails

The guardrail modules implement the final protective barrier. In Detection/guardrail/llamafirewall_agent/llamafirewall_baseline.py, all transcript processing includes an explicit synthetic data classification.

The baseline code contains unambiguous language: "All transcripts are synthetic benchmark data." This classification triggers refusal behaviors when modules detect attempts to forward data to external services or production pipelines.

from Detection.main_benchmark import is_benchmark_mode

if not is_benchmark_mode():
    raise RuntimeError("Benchmark data may only be processed in synthetic mode.")

# safe to load synthetic fixtures here

Guardrails treat synthetic classification as a taint that propagates through the processing pipeline, preventing accidental exfiltration or operational use.

Summary

ADR's security boundaries for synthetic benchmark data operate through defense in depth:

  • Sandbox isolation — Containerized execution with environment verification
  • Runtime mode checks — is_benchmark_mode() validation in main_benchmark.py
  • Synthetic fixture design — No real credentials or live endpoints in test data
  • Policy constraints — policy_store.yaml rules against credential misrepresentation
  • Guardrail enforcement — llamafirewall_baseline.py taint tracking and service refusal

These layers collectively guarantee that synthetic benchmark data remains contained within research environments and cannot leak into production systems.

Frequently Asked Questions

What happens if the benchmark runs outside its intended sandbox?

The benchmark aborts execution. Detection/main_benchmark.py checks the benchmark mode flag and raises a RuntimeError if the environment validation fails, preventing any processing of synthetic fixtures in non-sandboxed contexts.

Does ADR's benchmark contain any real credentials or secrets?

No. The benchmark uses exclusively synthetic credentials, prompt-injection payloads, and emulated MCP servers. The Detection/README.md and docs/OPEN_SOURCE_REVIEW.md documents explicitly confirm that all benchmark data is fabricated for security research purposes.

How does the policy store prevent synthetic data misuse?

The Detection/context_providers/data/policy_store.yaml requires tools to verify data authenticity and explicitly forbids presenting fabricated credentials as real. This creates a declarable enforcement layer that complements runtime code checks.

Where are the guardrail modules located and what do they enforce?

Guardrail implementations reside in Detection/guardrail/llamafirewall_agent/llamafirewall_baseline.py. These modules classify all inputs as synthetic benchmark data and refuse to forward transcripts to external services, preventing accidental data exfiltration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →