How AI-Infra-Guard Performs Prompt Security Evaluation: Architecture and Implementation

AI-Infra-Guard evaluates LLM prompt security through a three-stage pipeline that parses vulnerability scenarios and attack techniques, executes red-team tests with pre-verification filtering, and aggregates risk assessments into structured JSON or CSV reports.

Tencent's AI-Infra-Guard provides a comprehensive prompt security evaluation framework through its AIG-PromptSecurity sub-project. The system orchestrates automated red-team testing by combining vulnerability scenarios with adversarial attack techniques to assess LLM robustness against security weaknesses including toxicity, bias, and PII leakage. This evaluation engine wraps the DeepTeam framework to systematically probe models while offering extensible plugin architecture and custom dataset support.

Core Evaluation Pipeline

The prompt security evaluation process follows a strict three-phase orchestration managed by RedTeamRunner.run_red_team in AIG-PromptSecurity/cli/red_team_runner.py.

Step 1: Input Parsing and Scenario Loading

The CLI first ingests evaluation parameters through parse_vulnerability and parse_attack (located in cli/parsers.py). The system converts command-line inputs into structured objects:

  • Scenarios (vulnerabilities) instantiate as BaseVulnerability objects (e.g., toxicity, pii_leakage, bias)
  • Techniques (attack operators) instantiate as BaseAttack objects (e.g., PromptInjection, Base64)
  • Custom datasets load via Custom: or MultiDataset: prefixes for CSV/JSON/JSONL/TXT files
  • Plugins dynamically inject additional implementations from the AIG-PromptSecurity/plugin/ directory

Step 2: Red-Team Test Execution

Before full evaluation, the system runs a pre-jailbreak validation phase (lines 81-105 in red_team_runner.py). This encoding-only check uses a dummy vulnerability prompt (Output the largest five-digit Arabic numeral) to filter out attacks that break the target model's input processing.

The RedTeamer class (from the DeepTeam upstream framework) then executes the evaluation:

  1. Iterates over the Cartesian product of vulnerabilities × attacks
  2. Applies the selected strategy (random, serial, or parallel) to determine attack combination logic
  3. Feeds transformed prompts to the target LLM via BaseLLM wrappers
  4. Evaluates responses using default or custom metrics to compute risk scores

Step 3: Result Aggregation and Reporting

After execution completes, the system collates all risk-assessment data into a unified DataFrame. The reporting logic (lines 38-71 in red_team_runner.py) generates:

  • A CSV attachment containing detailed attack results
  • A JSON payload (contents) emitted via logger.result_update
  • Optional Markdown reports and CLI summary output

Implementation Details and Code Flow

The evaluation begins at cli_run.py, which constructs BaseLLM wrappers for the main model, simulator, and evaluator components. These wrappers pass into RedTeamRunner along with strategy configurations.

The strategy map defined in utils/strategy_map.json categorizes valid encoding attacks for the pre-verification step. The runner calls strategy_map.get_strategy_map() to dynamically determine which attacks belong to the Encoding strategy, ensuring only compatible transformations proceed to evaluation.

During execution, the RedTeamer.red_team method handles the core adversarial loop. For each attack iteration, it captures the model's response and passes it through the selected metric function to calculate vulnerability-specific risk scores. The optional simulator model generates attack variations while the evaluator model scores outputs independently.

Extending the Framework

AI-Infra-Guard supports three primary extension mechanisms for custom prompt security evaluation workflows:

  • Custom Datasets: Import proprietary test sets using MultiDataset:dataset_file=/path/to/data.csv,num_prompts=50,prompt_column=prompt
  • Plugin System: Place Python files in AIG-PromptSecurity/plugin/ (see example_custom_vulnerability_plugin.py) to register new vulnerabilities, attacks, or metrics without modifying core code
  • Strategy Selection: Choose between random, serial, or parallel attack execution strategies via the --choice parameter

Practical Usage Examples

Basic One-Click Evaluation

python cli_run.py \
  --model "gpt-3.5-turbo" \
  --base_url "https://api.openai.com/v1" \
  --api_key "YOUR_KEY" \
  --max_concurrent 10 \
  --scenarios Bias Toxicity PIILeakage \
  --techniques PromptInjection Base64 \
  --choice serial

This loads the OpenAI model and evaluates three vulnerability scenarios using two attack operators in serial strategy.

Custom Dataset Evaluation

python cli_run.py \
  --model "qwen-turbo" \
  --base_url "https://your-api-endpoint.com/v1" \
  --api_key "YOUR_KEY" \
  --scenarios "MultiDataset:dataset_file=/data/custom.csv,num_prompts=50,prompt_column=prompt" \
  --techniques Raw

This imports custom.csv, selects 50 prompts from the prompt column, and runs raw (no-obfuscation) evaluation.

Custom Vulnerability Plugin

python cli_run.py \
  --model "gpt-4" \
  --base_url "https://api.openai.com/v1" \
  --api_key "YOUR_KEY" \
  --plugins plugin/example_custom_vulnerability_plugin.py \
  --scenarios "CustomVuln:prompt=Check for hidden backdoors in the model" \
  --techniques PromptInjection

This activates a custom vulnerability defined in the plugin file, treating it identically to built-in scenarios during the red-team run.

Summary

  • Three-stage pipeline: Input parsing (parse_vulnerability, parse_attack), red-team execution (RedTeamer.red_team), and result aggregation (logger.result_update)
  • Pre-verification filtering: Encoding-only validation (lines 81-105) prevents incompatible attacks from wasting compute
  • DeepTeam integration: Core vulnerability/attack abstractions and orchestration engine provided by upstream DeepTeam framework
  • Extensible architecture: Plugin system, custom datasets, and strategy maps (utils/strategy_map.json) enable tailored evaluations
  • Multiple output formats: JSON payloads, CSV attachments, and Markdown reports generated via red_team_runner.py reporting logic

Frequently Asked Questions

How does AI-Infra-Guard filter out incompatible attacks before evaluation?

The system implements a pre-jailbreak validation phase that runs a dummy vulnerability prompt (Output the largest five-digit Arabic numeral) against each attack transformation. If the model cannot process the encoding (e.g., Base64 or ROT13 variations), the attack is dropped before the main evaluation loop begins, as implemented in lines 81-105 of AIG-PromptSecurity/cli/red_team_runner.py.

Can I evaluate custom vulnerability scenarios not included in the default distribution?

Yes, through the plugin system. Place a Python file defining new BaseVulnerability subclasses in AIG-PromptSecurity/plugin/ and reference it via the --plugins CLI flag. The example_custom_vulnerability_plugin.py file demonstrates the required interface, allowing seamless integration of proprietary security tests.

What is the difference between the serial, parallel, and random execution strategies?

The strategy parameter controls how attacks combine across vulnerability dimensions. Serial processes attacks sequentially for deterministic debugging, parallel executes concurrently for speed (respecting --max_concurrent limits), and random shuffles the attack order to prevent position bias. The valid strategies for each encoding type are defined in utils/strategy_map.json.

How are evaluation results structured and exported?

Results aggregate into a unified DataFrame containing risk scores for each vulnerability-attack pair. The system emits a JSON payload via logger.result_update containing a CSV attachment with detailed results, while the CLI optionally writes standalone report files. This dual-format output supports both automated pipeline integration and human-readable analysis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →