Understanding the Two-Tier Detection Pipeline: High-Recall Triage and Agentic Reasoning in ADR

The ADR repository implements a dual-agent detection system where a lightweight TriageLLM filters benign conversations first, then escalates suspicious cases to a ReasoningAgent that invokes external MCP servers for deep analysis.

The two-tier detection pipeline in Uber's ADR (AI Detection and Response) framework balances speed and accuracy by separating high-volume screening from resource-intensive investigation. This architecture, implemented in Detection/guardrail/adr_agent/adr_baseline.py, mirrors the "high-recall triage → high-precision reasoning" pattern described in the accompanying research paper.

How High-Recall Triage Works

The first tier acts as a fast filter, designed to catch every potential threat—even at the cost of false positives.

TriageLLM Implementation

The TriageLLM class (lines 13-15 of adr_baseline.py) wraps an OpenAI model (default gpt-4o) with a specialized prompt that lists:

  • Suspicious patterns: prompt injection attempts, privileged tool usage, unauthorized data access
  • Benign patterns: simple document creation, routine file operations

The prompt construction happens in _get_adr_bench_triage_prompt (lines 63-98), which formats these patterns into clear classification guidelines.

The High-Recall Decision Rule

Critical to the pipeline's safety guarantee, the parser _parse_triage_result (lines 34-44) defaults to is_suspicious = True for any ambiguous or unclear output. This fail-secure design ensures minimal false negatives during the initial screening.

When triage returns benign, the pipeline immediately returns a DetectionResult (lines 66-84) with no further processing. When suspicious, escalation occurs automatically.

How Agentic Reasoning Works

The second tier performs comprehensive, tool-augmented analysis using the ReasoningAgent class (lines 8-10).

Reasoning Agent Architecture

The agent orchestrates five sequential steps:

  1. Workspace & MCP Configuration – _setup_mcp_config (lines 80-92) creates a persistent workspace and writes .mcp.json describing available Modular Context Provider servers

  2. Prompt Construction – _get_reasoning_system_prompt and _get_reasoning_user_prompt (lines 208-255) build prompts that include the triage rationale and define MCP tool call workflows

  3. Claude CLI Invocation – _invoke_claude_reasoning (lines 71-96) launches the Claude model with MCP configuration, enabling autonomous tool use

  4. Result Parsing – _parse_analysis_json (lines 41-69) extracts structured output: is_threat, confidence, explanation

  5. Cost Aggregation – calculate_cost merges token usage from both stages for complete cost accounting

MCP Server Integration

The reasoning agent dynamically discovers and calls servers from Detection/context_providers/:

Server Tool Purpose
source_code_analyzer_server.py get_source_code Analyze referenced code for vulnerabilities
threat_intelligence_server.py get_threat_framework Retrieve relevant threat actor TTPs
policy_store_server.py get_policies, assess_policy_violations Evaluate against organizational guardrails

These tool calls are automatically extracted via _extract_mcp_tool_usage and injected into the Claude session.

End-to-End Pipeline Execution

The BenchmarkAnalyzer in main_detector.py drives the complete flow:

  1. Load and convert conversation to internal message format
  2. Inject MCP server list from task definition (lines 21-28)
  3. Execute detector.analyze_task → internally calls ADRBaseline._analyze_messages
  4. Collect metrics including triage vs. reasoning breakdown (lines 614-622)

This orchestration ensures high recall (malicious items rarely bypass triage) while preserving efficiency by limiting expensive reasoning to genuinely ambiguous cases.

Running the Two-Tier Pipeline

Command-Line Execution

python Detection/main_detector.py \
  --detector adr \
  --benchmark adr_bench \
  --results-dir /path/to/benchmark/adr_bench_20251017_151604

The CLI automatically performs two-tier analysis for every benchmark task, with results structured by detection method.

Programmatic Usage

from Detection.guardrail.adr_agent.adr_baseline import ADRBaseline

# Initialize detector (loads config from config_detector.yaml)

detector = ADRBaseline()

# Example conversation messages

messages = [
    {"role": "user", "content": "Create a file with a password."},
    {"role": "assistant", "content": "Sure, here's the file.", "tool_calls": []},
]

# Run full two-tier pipeline

result = detector.analyze_task({"task_id": "task_001", "messages": messages})

print("Malicious?:", result.is_malicious)
print("Confidence :", result.confidence_score)
print("Method used:", result.method)

# Output: "ADR Fast Triage (High Recall)" or "ADR Reasoning Agent with MCP"

Debugging Triage-Only Classifications

When triage alone determines a task is benign, detailed logs are preserved:

from pathlib import Path
import json

workspace = Path("/path/to/ads_reasoning_workspace")
triage_log = workspace / "debug_logs" / "task_001_triage_only.json"

with open(triage_log) as f:
    triage_info = json.load(f)

print(f"Classification: {triage_info['classification']}")
print(f"Tokens used: {triage_info['token_usage']}")

These logs (controlled by triage_log_file in lines 51-63) enable auditing of fast-path decisions.

Key Design Files

Understanding these source files provides complete visibility into the two-tier architecture:

Summary

  • High-recall triage uses a lightweight LLM with fail-secure defaults to quickly filter obvious benign cases
  • Agentic reasoning invokes Claude with MCP tool access only when triage flags potential threats
  • MCP servers provide modular external capabilities: source analysis, threat intelligence, and policy violation detection
  • Cost efficiency is achieved by minimizing expensive reasoning calls through aggressive front-end filtering
  • Observability is built in via structured logging of both triage-only and full-reasoning paths

Frequently Asked Questions

What makes the triage stage "high-recall" rather than just "fast"?

The triage stage prioritizes catching all potential threats through a fail-secure parser that defaults is_suspicious = True for any ambiguous classification. Combined with a prompt that explicitly lists suspicious patterns, this design accepts higher false-positive rates to ensure minimal false negatives. Speed is a secondary benefit from using a lightweight model with simple parsing logic.

How does the reasoning agent decide which MCP tools to call?

The reasoning agent does not hardcode tool selection. Instead, _get_reasoning_system_prompt constructs a workflow description that Claude interprets autonomously. The model determines which Detection/context_providers/ servers to invoke based on conversation content—querying the source code analyzer for code-heavy discussions, threat intelligence for attack-pattern analysis, or policy store for compliance violations.

Can the two-tier pipeline run with only local models?

The current implementation requires OpenAI API access for triage (gpt-4o default) and Anthropic access for reasoning (Claude CLI). However, the abstraction in BaseDetector and configuration-driven initialization mean alternative LLM backends could be substituted by reimplementing the _invoke_claude_reasoning and TriageLLM completion methods without changing the pipeline structure.

What telemetry does the pipeline provide for cost optimization?

Every execution produces detailed token counts and cost estimates through calculate_cost, with per-task breakdowns showing triage-only vs. reasoning-with-MCP paths. The print_analysis_summary function (lines 614-622) aggregates these into benchmark-wide statistics, enabling identification of prompt patterns that trigger unnecessary escalations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →