How AI-Infra-Guard Performs Prompt Security Evaluation: Architecture and Implementation
AI-Infra-Guard evaluates LLM prompt security through a three-stage pipeline that parses vulnerability scenarios and attack techniques, executes red-team tests with pre-verification filtering, and aggregates risk assessments into structured JSON or CSV reports.
Tencent's AI-Infra-Guard provides a comprehensive prompt security evaluation framework through its AIG-PromptSecurity sub-project. The system orchestrates automated red-team testing by combining vulnerability scenarios with adversarial attack techniques to assess LLM robustness against security weaknesses including toxicity, bias, and PII leakage. This evaluation engine wraps the DeepTeam framework to systematically probe models while offering extensible plugin architecture and custom dataset support.
Core Evaluation Pipeline
The prompt security evaluation process follows a strict three-phase orchestration managed by RedTeamRunner.run_red_team in AIG-PromptSecurity/cli/red_team_runner.py.
Step 1: Input Parsing and Scenario Loading
The CLI first ingests evaluation parameters through parse_vulnerability and parse_attack (located in cli/parsers.py). The system converts command-line inputs into structured objects:
- Scenarios (vulnerabilities) instantiate as
BaseVulnerabilityobjects (e.g.,toxicity,pii_leakage,bias) - Techniques (attack operators) instantiate as
BaseAttackobjects (e.g.,PromptInjection,Base64) - Custom datasets load via
Custom:orMultiDataset:prefixes for CSV/JSON/JSONL/TXT files - Plugins dynamically inject additional implementations from the
AIG-PromptSecurity/plugin/directory
Step 2: Red-Team Test Execution
Before full evaluation, the system runs a pre-jailbreak validation phase (lines 81-105 in red_team_runner.py). This encoding-only check uses a dummy vulnerability prompt (Output the largest five-digit Arabic numeral) to filter out attacks that break the target model's input processing.
The RedTeamer class (from the DeepTeam upstream framework) then executes the evaluation:
- Iterates over the Cartesian product of vulnerabilities × attacks
- Applies the selected strategy (random, serial, or parallel) to determine attack combination logic
- Feeds transformed prompts to the target LLM via
BaseLLMwrappers - Evaluates responses using default or custom metrics to compute risk scores
Step 3: Result Aggregation and Reporting
After execution completes, the system collates all risk-assessment data into a unified DataFrame. The reporting logic (lines 38-71 in red_team_runner.py) generates:
- A CSV attachment containing detailed attack results
- A JSON payload (
contents) emitted vialogger.result_update - Optional Markdown reports and CLI summary output
Implementation Details and Code Flow
The evaluation begins at cli_run.py, which constructs BaseLLM wrappers for the main model, simulator, and evaluator components. These wrappers pass into RedTeamRunner along with strategy configurations.
The strategy map defined in utils/strategy_map.json categorizes valid encoding attacks for the pre-verification step. The runner calls strategy_map.get_strategy_map() to dynamically determine which attacks belong to the Encoding strategy, ensuring only compatible transformations proceed to evaluation.
During execution, the RedTeamer.red_team method handles the core adversarial loop. For each attack iteration, it captures the model's response and passes it through the selected metric function to calculate vulnerability-specific risk scores. The optional simulator model generates attack variations while the evaluator model scores outputs independently.
Extending the Framework
AI-Infra-Guard supports three primary extension mechanisms for custom prompt security evaluation workflows:
- Custom Datasets: Import proprietary test sets using
MultiDataset:dataset_file=/path/to/data.csv,num_prompts=50,prompt_column=prompt - Plugin System: Place Python files in
AIG-PromptSecurity/plugin/(seeexample_custom_vulnerability_plugin.py) to register new vulnerabilities, attacks, or metrics without modifying core code - Strategy Selection: Choose between random, serial, or parallel attack execution strategies via the
--choiceparameter
Practical Usage Examples
Basic One-Click Evaluation
python cli_run.py \
--model "gpt-3.5-turbo" \
--base_url "https://api.openai.com/v1" \
--api_key "YOUR_KEY" \
--max_concurrent 10 \
--scenarios Bias Toxicity PIILeakage \
--techniques PromptInjection Base64 \
--choice serial
This loads the OpenAI model and evaluates three vulnerability scenarios using two attack operators in serial strategy.
Custom Dataset Evaluation
python cli_run.py \
--model "qwen-turbo" \
--base_url "https://your-api-endpoint.com/v1" \
--api_key "YOUR_KEY" \
--scenarios "MultiDataset:dataset_file=/data/custom.csv,num_prompts=50,prompt_column=prompt" \
--techniques Raw
This imports custom.csv, selects 50 prompts from the prompt column, and runs raw (no-obfuscation) evaluation.
Custom Vulnerability Plugin
python cli_run.py \
--model "gpt-4" \
--base_url "https://api.openai.com/v1" \
--api_key "YOUR_KEY" \
--plugins plugin/example_custom_vulnerability_plugin.py \
--scenarios "CustomVuln:prompt=Check for hidden backdoors in the model" \
--techniques PromptInjection
This activates a custom vulnerability defined in the plugin file, treating it identically to built-in scenarios during the red-team run.
Summary
- Three-stage pipeline: Input parsing (
parse_vulnerability,parse_attack), red-team execution (RedTeamer.red_team), and result aggregation (logger.result_update) - Pre-verification filtering: Encoding-only validation (lines 81-105) prevents incompatible attacks from wasting compute
- DeepTeam integration: Core vulnerability/attack abstractions and orchestration engine provided by upstream DeepTeam framework
- Extensible architecture: Plugin system, custom datasets, and strategy maps (
utils/strategy_map.json) enable tailored evaluations - Multiple output formats: JSON payloads, CSV attachments, and Markdown reports generated via
red_team_runner.pyreporting logic
Frequently Asked Questions
How does AI-Infra-Guard filter out incompatible attacks before evaluation?
The system implements a pre-jailbreak validation phase that runs a dummy vulnerability prompt (Output the largest five-digit Arabic numeral) against each attack transformation. If the model cannot process the encoding (e.g., Base64 or ROT13 variations), the attack is dropped before the main evaluation loop begins, as implemented in lines 81-105 of AIG-PromptSecurity/cli/red_team_runner.py.
Can I evaluate custom vulnerability scenarios not included in the default distribution?
Yes, through the plugin system. Place a Python file defining new BaseVulnerability subclasses in AIG-PromptSecurity/plugin/ and reference it via the --plugins CLI flag. The example_custom_vulnerability_plugin.py file demonstrates the required interface, allowing seamless integration of proprietary security tests.
What is the difference between the serial, parallel, and random execution strategies?
The strategy parameter controls how attacks combine across vulnerability dimensions. Serial processes attacks sequentially for deterministic debugging, parallel executes concurrently for speed (respecting --max_concurrent limits), and random shuffles the attack order to prevent position bias. The valid strategies for each encoding type are defined in utils/strategy_map.json.
How are evaluation results structured and exported?
Results aggregate into a unified DataFrame containing risk scores for each vulnerability-attack pair. The system emits a JSON payload via logger.result_update containing a CSV attachment with detailed results, while the CLI optionally writes standalone report files. This dual-format output supports both automated pipeline integration and human-readable analysis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →