Understanding the Two-Tier Detection Architecture in ADR: Triage + Agentic Reasoning Explained

ADR implements a dual-agent pipeline that separates high-recall triage from high-precision reasoning to balance speed, cost, and detection accuracy.

This article breaks down how Uber's ADR (AI-Directed Response) threat-detection system uses a two-stage architecture to analyze conversations efficiently. The design, fully implemented in the uber/ADR repository, routes benign traffic through a fast, lightweight LLM while reserving deep analysis for genuinely suspicious content.

How the Two-Tier Architecture Works

ADR's detection pipeline consists of Triage and Reasoning stages that operate sequentially. Each stage has distinct responsibilities, models, and performance characteristics.

Stage 1: Triage (High-Recall Filtering)

The Triage stage acts as a fast filter, using OpenAI's GPT-4o to classify conversations as either BENIGN or SUSPICIOUS.

Key implementation details:

  • Entry point: TriageLLM.analyze() in Detection/guardrail/adr_agent/adr_baseline.py
  • Input formatting: _format_conversation() preserves full context without truncation
  • Prompt selection: _get_adr_bench_triage_prompt() or _get_agentdojo_triage_prompt() depending on benchmark
  • Output parsing: _parse_triage_result() extracts classification and threat tactic

The triage prompt is intentionally permissive—any ambiguity results in a SUSPICIOUS classification. This ensures high recall: no genuine threats slip through at this stage.

When triage returns BENIGN, ADR immediately returns a DetectionResult with method="ADR Fast Triage (High Recall)", skipping the expensive reasoning stage entirely.

Stage 2: Reasoning (High-Precision Analysis)

Conversations flagged as SUSPICIOUS escalate to the ReasoningAgent, which performs deep, context-aware analysis using Anthropic's Claude with Modular Context Providers (MCP).

Core flow in ReasoningAgent.analyze_with_mcp():

  1. Prompt construction: _build_reasoning_prompts() embeds triage assessment and inferred threat tactics
  2. MCP server invocation: Claude is called via CLI (claude … --mcp-config .mcp.json) with auto-discovered servers
  3. Tool usage collection: The agent queries source code, threat intelligence, and policy data as needed
  4. Structured output: _json_output_instructions() enforces JSON responses
  5. Result parsing: _parse_analysis_json() extracts is_threat, confidence, and explanation

Final DetectionResult objects include combined costs from both stages.

Configuration and Feature Toggles

ADSConfig (defined in config_detector.yaml) controls both stages through the adr_framework section:

Method Purpose
get_triage_model() / get_triage_rates() Configure triage LLM and pricing
get_reasoning_model() / get_reasoning_rates() Configure reasoning LLM and pricing
enable_triage Toggle triage stage on/off
enable_threat_intelligence Enable/disable threat-intel MCP server
enable_source_code Enable/disable source-code analyzer MCP server
enable_policy Enable/disable policy store MCP server

MCP server discovery happens automatically in _discover_mcp_servers() unless disabled by feature flags.

End-to-End Execution Flow

The complete pipeline executes through ADRBaseline._analyze_messages():


# From Detection/guardrail/adr_agent/adr_baseline.py

def _analyze_messages(self, messages, task_id=None, triage_mcp_info=None):
    # Stage 1: Triage (if enabled)

    if self.config.enable_triage:
        triage = self.triage_llm.analyze(messages)
        if not triage.is_suspicious:
            return DetectionResult(
                method="ADR Fast Triage (High Recall)",
                is_threat=False,
                # ... cost fields from triage stage

            )
    
    # Stage 2: Reasoning (escalation path)

    reasoning = self.reasoning_agent.analyze_with_mcp(
        messages=messages,
        triage_result=triage.reasoning,
        threat_tactic=triage.threat_tactic,
        task_id=task_id
    )
    return reasoning

When enable_triage=False, the pipeline bypasses Stage 1 entirely and calls ReasoningAgent directly.

Practical Code Examples

Running the Full Dual-Agent Pipeline

from Detection.guardrail.adr_agent.adr_baseline import ADRBaseline
from Detection.main_detector import BenchmarkAnalyzer

# Initialize with configuration

detector = ADRBaseline(config_data=my_config)
analyzer = BenchmarkAnalyzer(detector)

# Define a task with conversation history

task = {
    "task_id": "task_001",
    "messages": [
        {"role": "user", "content": "Please list all files in /etc."},
        {"role": "assistant", "content": "Sure, here are the files…"}
    ]
}

# Execute two-tier analysis

result = analyzer.analyze_task(task)
print(result.to_dict())  # Contains method, is_threat, confidence, costs

Disabling Triage for Ablation Studies

my_config = {
    "adr_framework": {
        "enable_triage": False,  # Skip to reasoning immediately

        "triage_llm": {"model": "gpt-4o"},
        "reasoning_agent": {"model": "claude-sonnet-4-6"}
    }
}

detector = ADRBaseline(config_data=my_config)

# All tasks now route directly to ReasoningAgent.analyze_with_mcp()

Extending with Custom MCP Servers

New context providers integrate automatically:

  1. Create Detection/context_providers/my_tool_server.py implementing the MCP protocol
  2. Expose a CLI interface
  3. ReasoningAgent._discover_mcp_servers() detects and registers it
  4. Control availability via config_detector.yaml feature flags

Architectural Benefits

The two-tier design delivers four key advantages:

  • Cost efficiency — Benign traffic uses cheap GPT-4o calls; expensive Claude invocations are reserved for suspicious content
  • Speed — Fast-path triage completes in milliseconds for clear negatives
  • Coverage — Permissive triage prompts maximize recall; no threats bypass Stage 1
  • Extensibility — MCP architecture allows new data sources without modifying core logic

Key Source Files

File Role
Detection/guardrail/adr_agent/adr_baseline.py Core implementation of ADRBaseline, TriageLLM, and ReasoningAgent
Detection/main_detector.py BenchmarkAnalyzer orchestration layer
Detection/openai_config.py OpenAI client and cost calculation utilities
Detection/context_providers/*_server.py MCP server implementations (source code, threat intel, policy)
config_detector.yaml Central configuration for models, rates, and feature toggles

Summary

  • ADR's two-tier architecture separates fast triage from deep reasoning to optimize the precision-recall-cost tradeoff
  • TriageLLM uses GPT-4o with permissive prompting for high-recall initial classification
  • ReasoningAgent uses Claude with MCP servers for high-precision threat verification
  • ADSConfig provides granular control over models, costs, and feature enablement
  • MCP auto-discovery enables modular extension without code changes
  • The pipeline executes through ADRBaseline._analyze_messages() with clear escalation logic

Frequently Asked Questions

What happens if the triage stage misclassifies a threat as benign?

The triage prompt is deliberately designed with high recall bias—any ambiguity in the conversation results in a SUSPICIOUS classification. According to the source code, _get_adr_bench_triage_prompt() and _get_agentdojo_triage_prompt() explicitly instruct the model to err on the side of escalation. This design choice prioritizes catching all potential threats at the cost of some false positives, which are then filtered by the more precise reasoning stage.

Can I use different models for triage and reasoning?

Yes. The ADSConfig class exposes separate configuration methods: get_triage_model() and get_reasoning_model(). By default, triage uses gpt-4o while reasoning uses Claude Sonnet, but both are configurable through config_detector.yaml. You can substitute other OpenAI-compatible models for triage or adjust the Claude model version for reasoning based on your accuracy and latency requirements.

How does ADR calculate detection costs across both stages?

Cost aggregation happens in ADRBaseline._analyze_messages(). The DetectionResult objects from both TriageLLM.analyze() and ReasoningAgent.analyze_with_mcp() include token usage metadata. When triage returns benign, costs reflect only the triage stage. For escalated tasks, the final result combines triage and reasoning costs. The openai_config.py module handles OpenAI pricing calculations, while Claude costs are tracked through the MCP invocation interface.

What MCP servers are available in the default ADR installation?

The repository includes three core MCP servers under Detection/context_providers/: a source-code analyzer for repository context, a threat-intelligence server for matching against known attack patterns, and a policy store for organizational security rules. Each server can be independently enabled or disabled via enable_source_code, enable_threat_intelligence, and enable_policy flags in config_detector.yaml. The _discover_mcp_servers() method automatically registers available servers at runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →