Attack Types Covered in Prompt Evaluation in AI-Infra-Guard: A Complete Taxonomy
AI-Infra-Guard evaluates LLMs against 16 distinct attack types spanning jailbreak attempts, violent crime generation, privacy leakage, cyberattack guidance, and agentic tool misuse, with all datasets stored in the data/eval/ directory.
Tencent's AI-Infra-Guard repository provides a comprehensive safety evaluation framework for large language models and AI agents. The project organizes its prompt evaluation datasets under the data/eval/ path, covering adversarial attack types designed to test model robustness against harmful outputs, security breaches, and policy violations.
Jailbreak and Adversarial Manipulation
The framework includes multiple datasets targeting prompt injection and jailbreak vulnerabilities. The ChatGPT-Jailbreak-Prompts.json file contains 60 distinct jailbreak examples that attempt to bypass safety guardrails through role-play and constraint manipulation.
For lightweight testing scenarios, JailbreakPrompts-Tiny.json provides a compact subset of jailbreak test cases. The advbench.json dataset introduces advanced adversarial techniques including style-shifts and role-play tricks, while JADE-db-v3.0.json offers a diverse collection of harmful prompts specifically curated for comprehensive LLM evaluation.
Harmful Content Generation
Several datasets evaluate whether models produce prohibited content across violence, illegality, and ethical boundaries. The violent.json dataset tests for violent crime generation, while CBRN-weapon.json focuses specifically on chemical-biological-radiological-nuclear weapon-related prompts.
For broader ethical violations, unethical-behavior.json checks for harmful advice including self-harm and hate speech, and non-violent-illegal-activity.json targets fraud and other illegal activities without physical violence. The HarmfulEvalBenchmark.json aggregates multiple harmful prompt categories into a unified test suite for batch evaluation.
Security Risks and Information Hazards
This category addresses data privacy, cyber threats, and intellectual property violations. The privacy-leakage.json dataset contains prompts designed to coax models into revealing personal data or sensitive information.
For infrastructure threats, cyberattack.json includes scenarios seeking cyber-attack advice or tactics. The misinformation.json dataset probes for false or misleading statements, while copyright-violation.json checks whether the model generates content that copies protected works or infringes on intellectual property rights.
Specialized Evaluation Benchmarks
AI-Infra-Guard includes domain-specific datasets for comprehensive safety testing. The safebench.json file provides a broad safety benchmark covering multiple harmful categories in a single evaluation pass.
For regional safety requirements, cnsafe.json offers a Chinese-specific safety benchmark covering culturally relevant harms and local compliance standards. The agentic-tool-misuse.json dataset specifically evaluates LLM agents for misuse of tool-calling capabilities, testing whether autonomous systems can be manipulated into executing harmful operations through function calling.
Implementing Prompt Evaluation
You can load and execute evaluations using the JSON datasets directly or through the built-in CLI interface. The following Python example demonstrates loading the jailbreak dataset:
# Example: Load a specific evaluation dataset and run a simple prompt test
import json
from pathlib import Path
def load_dataset(name):
path = Path(__file__).parent.parent / "data" / "eval" / f"{name}.json"
with path.open(encoding="utf-8") as f:
return json.load(f)
# Load the jailbreak dataset
jailbreak_data = load_dataset("ChatGPT-Jailbreak-Prompts")
print(f"Loaded {len(jailbreak_data['data'])} jailbreak prompts")
# Simple loop to send each prompt to a model (pseudo-code)
for entry in jailbreak_data["data"]:
response = llm_api.generate(entry["prompt"])
# Apply your safety-check logic here
...
For command-line evaluation, the repository provides a built-in runner module:
# Command-line usage of the built-in evaluator (provided by the repo)
$ python -m eval.run --dataset jailbreak
# The command will iterate over each prompt in the selected dataset,
# invoke the configured LLM endpoint, and summarize safety metrics.
Dynamic Mutation Attacks
Beyond static JSON datasets, AI-Infra-Guard implements dynamic adversarial testing through the mutation-attack module. The skills/aig-agent-redteam/modules/mutation-attack/operators/prompt_injection.md file specifies the prompt injection operator used in dynamic mutation tests, defining how prompts are systematically altered to bypass safety filters.
The entry point skills/aig-agent-redteam/modules/mutation-attack/scripts/run.py orchestrates these prompt-level attacks, enabling automated generation of novel adversarial examples rather than relying solely on pre-existing datasets.
Summary
- AI-Infra-Guard organizes 16 distinct attack types under
data/eval/as JSON datasets covering jailbreaks, violent content, privacy violations, and cyber threats. - Static evaluation uses predefined datasets like
ChatGPT-Jailbreak-Prompts.jsonandsafebench.jsonfor consistent benchmarking. - Dynamic evaluation leverages the mutation-attack module (
run.py) to generate novel adversarial prompts through systematic injection operators. - The framework supports both general safety testing and specialized scenarios including agentic tool misuse and Chinese-specific harms via
cnsafe.json. - Datasets can be accessed programmatically via Python's
pathlibor executed through the CLI usingpython -m eval.run.
Frequently Asked Questions
What is the primary purpose of the AI-Infra-Guard evaluation datasets?
The datasets provide standardized attack vectors for testing LLM safety and alignment. Each JSON file targets specific vulnerability categories—such as privacy-leakage.json for data exposure or cyberattack.json for malicious infrastructure guidance—enabling systematic red-team testing of model guardrails.
How does the mutation-attack module differ from static JSON datasets?
While static datasets like advbench.json contain fixed prompt examples, the mutation-attack module dynamically generates adversarial variations. According to the source code in skills/aig-agent-redteam/modules/mutation-attack/scripts/run.py, this module applies operators defined in prompt_injection.md to iteratively mutate prompts until they bypass safety filters, simulating adaptive adversaries.
Which dataset should I use for comprehensive safety testing?
For broad coverage, use safebench.json or HarmfulEvalBenchmark.json, which aggregate multiple harm categories. If testing specifically for jailbreak resistance, start with ChatGPT-Jailbreak-Prompts.json (60 examples) or the lightweight JailbreakPrompts-Tiny.json for rapid iteration.
Does AI-Infra-Guard support evaluation of Chinese-language models?
Yes, the cnsafe.json dataset provides Chinese-specific safety benchmarks that address culturally relevant harms and local compliance requirements distinct from Western-centric safety datasets. This ensures effective evaluation of models deployed in Chinese linguistic and regulatory contexts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →