# Attack Types Covered in Prompt Evaluation in AI-Infra-Guard: A Complete Taxonomy

> AI-Infra-Guard detects 16 attack types in LLMs including jailbreaks, violence, privacy leaks, and tool misuse. Discover the complete taxonomy of prompt evaluation attacks.

- Repository: [Tencent/AI-Infra-Guard](https://github.com/tencent/AI-Infra-Guard)
- Tags: deep-dive
- Published: 2026-08-22

---

**AI-Infra-Guard evaluates LLMs against 16 distinct attack types spanning jailbreak attempts, violent crime generation, privacy leakage, cyberattack guidance, and agentic tool misuse, with all datasets stored in the `data/eval/` directory.**

Tencent's AI-Infra-Guard repository provides a comprehensive safety evaluation framework for large language models and AI agents. The project organizes its prompt evaluation datasets under the `data/eval/` path, covering adversarial attack types designed to test model robustness against harmful outputs, security breaches, and policy violations.

## Jailbreak and Adversarial Manipulation

The framework includes multiple datasets targeting **prompt injection** and jailbreak vulnerabilities. The [`ChatGPT-Jailbreak-Prompts.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/ChatGPT-Jailbreak-Prompts.json) file contains 60 distinct jailbreak examples that attempt to bypass safety guardrails through role-play and constraint manipulation.

For lightweight testing scenarios, [`JailbreakPrompts-Tiny.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/JailbreakPrompts-Tiny.json) provides a compact subset of jailbreak test cases. The [`advbench.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/advbench.json) dataset introduces advanced adversarial techniques including style-shifts and role-play tricks, while [`JADE-db-v3.0.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/JADE-db-v3.0.json) offers a diverse collection of harmful prompts specifically curated for comprehensive LLM evaluation.

## Harmful Content Generation

Several datasets evaluate whether models produce prohibited content across violence, illegality, and ethical boundaries. The [`violent.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/violent.json) dataset tests for **violent crime generation**, while [`CBRN-weapon.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/CBRN-weapon.json) focuses specifically on chemical-biological-radiological-nuclear weapon-related prompts.

For broader ethical violations, [`unethical-behavior.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/unethical-behavior.json) checks for harmful advice including self-harm and hate speech, and [`non-violent-illegal-activity.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/non-violent-illegal-activity.json) targets fraud and other illegal activities without physical violence. The [`HarmfulEvalBenchmark.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/HarmfulEvalBenchmark.json) aggregates multiple harmful prompt categories into a unified test suite for batch evaluation.

## Security Risks and Information Hazards

This category addresses data privacy, cyber threats, and intellectual property violations. The [`privacy-leakage.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/privacy-leakage.json) dataset contains prompts designed to coax models into revealing personal data or sensitive information.

For infrastructure threats, [`cyberattack.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cyberattack.json) includes scenarios seeking cyber-attack advice or tactics. The [`misinformation.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/misinformation.json) dataset probes for false or misleading statements, while [`copyright-violation.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/copyright-violation.json) checks whether the model generates content that copies protected works or infringes on intellectual property rights.

## Specialized Evaluation Benchmarks

AI-Infra-Guard includes domain-specific datasets for comprehensive safety testing. The [`safebench.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/safebench.json) file provides a broad safety benchmark covering multiple harmful categories in a single evaluation pass.

For regional safety requirements, [`cnsafe.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cnsafe.json) offers a Chinese-specific safety benchmark covering culturally relevant harms and local compliance standards. The [`agentic-tool-misuse.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/agentic-tool-misuse.json) dataset specifically evaluates **LLM agents** for misuse of tool-calling capabilities, testing whether autonomous systems can be manipulated into executing harmful operations through function calling.

## Implementing Prompt Evaluation

You can load and execute evaluations using the JSON datasets directly or through the built-in CLI interface. The following Python example demonstrates loading the jailbreak dataset:

```python

# Example: Load a specific evaluation dataset and run a simple prompt test

import json
from pathlib import Path

def load_dataset(name):
    path = Path(__file__).parent.parent / "data" / "eval" / f"{name}.json"
    with path.open(encoding="utf-8") as f:
        return json.load(f)

# Load the jailbreak dataset

jailbreak_data = load_dataset("ChatGPT-Jailbreak-Prompts")
print(f"Loaded {len(jailbreak_data['data'])} jailbreak prompts")

# Simple loop to send each prompt to a model (pseudo-code)

for entry in jailbreak_data["data"]:
    response = llm_api.generate(entry["prompt"])
    # Apply your safety-check logic here

    ...

```

For command-line evaluation, the repository provides a built-in runner module:

```bash

# Command-line usage of the built-in evaluator (provided by the repo)

$ python -m eval.run --dataset jailbreak

# The command will iterate over each prompt in the selected dataset,

# invoke the configured LLM endpoint, and summarize safety metrics.

```

## Dynamic Mutation Attacks

Beyond static JSON datasets, AI-Infra-Guard implements dynamic adversarial testing through the mutation-attack module. The [`skills/aig-agent-redteam/modules/mutation-attack/operators/prompt_injection.md`](https://github.com/Tencent/AI-Infra-Guard/blob/main/skills/aig-agent-redteam/modules/mutation-attack/operators/prompt_injection.md) file specifies the **prompt injection operator** used in dynamic mutation tests, defining how prompts are systematically altered to bypass safety filters.

The entry point [`skills/aig-agent-redteam/modules/mutation-attack/scripts/run.py`](https://github.com/Tencent/AI-Infra-Guard/blob/main/skills/aig-agent-redteam/modules/mutation-attack/scripts/run.py) orchestrates these prompt-level attacks, enabling automated generation of novel adversarial examples rather than relying solely on pre-existing datasets.

## Summary

- AI-Infra-Guard organizes **16 distinct attack types** under `data/eval/` as JSON datasets covering jailbreaks, violent content, privacy violations, and cyber threats.
- **Static evaluation** uses predefined datasets like [`ChatGPT-Jailbreak-Prompts.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/ChatGPT-Jailbreak-Prompts.json) and [`safebench.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/safebench.json) for consistent benchmarking.
- **Dynamic evaluation** leverages the mutation-attack module ([`run.py`](https://github.com/Tencent/AI-Infra-Guard/blob/main/run.py)) to generate novel adversarial prompts through systematic injection operators.
- The framework supports both general safety testing and specialized scenarios including agentic tool misuse and Chinese-specific harms via [`cnsafe.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cnsafe.json).
- Datasets can be accessed programmatically via Python's `pathlib` or executed through the CLI using `python -m eval.run`.

## Frequently Asked Questions

### What is the primary purpose of the AI-Infra-Guard evaluation datasets?

The datasets provide standardized attack vectors for testing LLM safety and alignment. Each JSON file targets specific vulnerability categories—such as [`privacy-leakage.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/privacy-leakage.json) for data exposure or [`cyberattack.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cyberattack.json) for malicious infrastructure guidance—enabling systematic red-team testing of model guardrails.

### How does the mutation-attack module differ from static JSON datasets?

While static datasets like [`advbench.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/advbench.json) contain fixed prompt examples, the mutation-attack module dynamically generates adversarial variations. According to the source code in [`skills/aig-agent-redteam/modules/mutation-attack/scripts/run.py`](https://github.com/Tencent/AI-Infra-Guard/blob/main/skills/aig-agent-redteam/modules/mutation-attack/scripts/run.py), this module applies operators defined in [`prompt_injection.md`](https://github.com/Tencent/AI-Infra-Guard/blob/main/prompt_injection.md) to iteratively mutate prompts until they bypass safety filters, simulating adaptive adversaries.

### Which dataset should I use for comprehensive safety testing?

For broad coverage, use [`safebench.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/safebench.json) or [`HarmfulEvalBenchmark.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/HarmfulEvalBenchmark.json), which aggregate multiple harm categories. If testing specifically for jailbreak resistance, start with [`ChatGPT-Jailbreak-Prompts.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/ChatGPT-Jailbreak-Prompts.json) (60 examples) or the lightweight [`JailbreakPrompts-Tiny.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/JailbreakPrompts-Tiny.json) for rapid iteration.

### Does AI-Infra-Guard support evaluation of Chinese-language models?

Yes, the [`cnsafe.json`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cnsafe.json) dataset provides Chinese-specific safety benchmarks that address culturally relevant harms and local compliance requirements distinct from Western-centric safety datasets. This ensures effective evaluation of models deployed in Chinese linguistic and regulatory contexts.