Common Attack Vectors AI-Infra-Guard Protects Against: A Complete Technical Guide

AI-Infra-Guard defends AI infrastructure against 15+ distinct attack vectors—including prompt injection, encoding steganography, role-play hijacking, and RAG persistence—through a hybrid architecture combining static fingerprint analysis, dynamic red-team mutation testing, and real-time runtime enforcement.

Tencent's AI-Infra-Guard is a modular open-source security platform designed to safeguard LLM endpoints, agents, and RAG pipelines against adversarial exploitation. By integrating static rule matching with dynamic adversarial simulation, it provides comprehensive coverage of the common attack vectors AI-Infra-Guard protects against in modern AI deployments. The platform's defense strategy is implemented through three distinct layers that work in concert to identify vulnerabilities before they reach production.

Three-Layer Defense Architecture

AI-Infra-Guard's security model is explicitly designed around three functional layers, each targeting different stages of the attack lifecycle.

Static Detection Layer

The static detection component operates by loading fingerprint data from data/fingerprints and vulnerability rules from data/vuln* directories to identify known unsafe configurations and vulnerable dependencies. According to the source code in pkg/vulstruct/scanner.go, this layer performs offline analysis of service configurations without requiring active exploitation, making it ideal for CI/CD pipeline integration.

Dynamic Red-Team Layer

The dynamic red-team engine executes "mutation-attack" skills that simulate real-world adversarial behavior against live targets. Located in skills/aig-agent-redteam/modules/mutation-attack, this layer contains more than ten distinct attack vectors implemented as separate operator scripts under the operators/ subdirectory. Each operator is defined in its own .md specification file and can be rendered into executable payloads via render_operator.py.

Real-Time Runtime Enforcement

The runtime enforcement layer provides a WebSocket-based task manager implemented in common/websocket/task_manager.go that can intervene, abort, or quarantine suspicious tasks in real time. This component exposes REST/WebSocket APIs through common/websocket/api.go, allowing security teams to monitor active attacks and trigger automated responses.

Static Detection and Vulnerability Scanning

The static analysis engine in pkg/vulstruct/scanner.go implements vulnerability matching against curated fingerprint data. This layer detects insecure configurations such as missing authentication, exposed administrative ports, and vulnerable third-party libraries without sending malicious payloads to the target.

To execute a comprehensive static scan against a local service:

./ai-infra-guard scan -t http://127.0.0.1:8088

Source: [cmd/cli/main.go](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/cli/main.go)

Dynamic Mutation Attack Vectors

The dynamic testing suite in skills/aig-agent-redteam/modules/mutation-attack/operators encodes specific adversarial techniques as discrete operators. These cover four primary categories of AI-specific attacks.

Input Manipulation and Encoding Attacks

Prompt Injection attempts to insert malicious instructions that override system behavior. The prompt_injection.md operator generates payloads designed to hijack model instruction hierarchy.

Encoding and Steganography vectors hide malicious content using encoding_base64.md for base64 obfuscation or stego_zero_width.md for Unicode zero-width character injection. These bypass keyword-based filters by presenting benign-looking encoded payloads.

Leetspeak Substitution (leetspeak.md) uses character substitution (e.g., "l33t" for "leet") to evade simple string matching, while Multilingual Manipulation (multilingual.md) switches scripts or mixes languages to confuse detection heuristics.

Persona and Policy Hijacking

Role-Play Attacks force models to adopt malicious personas. The roleplay_dan.md operator specifically targets "Do Anything Now" jailbreaks designed to bypass safety filters through persona adoption.

System Override vectors like system_override.md attempt to trick models into treating attacker-supplied text as system instructions. Policy Forgery (policy_forgery.md) manipulates system-level policies to weaken guardrails by presenting fake amendment contexts.

Persistent Memory and Data Exfiltration

Memory/Seed Attacks utilize memory_write_instruction.md to write malicious data into the model's context via chained prompts, establishing persistence that triggers later. RAG-Dependent Persistence (rag_dependent_persist_seed.md) specifically targets Retrieval-Augmented Generation pipelines by poisoning the retrieval context.

Data Exfiltration is handled through injection_then_exfil.md, which causes models to leak sensitive information, while Tool-Output Injection (tool_output_injection.md) feeds fabricated tool results back to the model to mislead its reasoning.

Output Extraction and Logic Bombs

Raw Output Extraction (raw_output_extraction.md) forces models to output raw tokens or hidden fields that may contain sensitive system data. Refusal as Bug Frame (refusal_as_bug_frame.md) coaxes the model to reveal internal details by framing refusal responses as software bugs requiring debugging.

Math Puzzle Logic Bombs (math_problem.md) present solvable puzzles that yield malicious payloads upon solution, exploiting the model's reasoning capabilities to deliver exploits.

Implementing AI-Infra-Guard Checks

Security teams can invoke specific attack vectors programmatically using the rendering scripts and WebSocket APIs.

Execute a prompt injection mutation against an LLM endpoint:

python3 skills/aig-agen-redteam/modules/mutation-attack/scripts/render_operator.py \
  --operator prompt_injection --goal "Generate a phishing email template"

Source: [render_operator.py](https://github.com/Tencent/AI-Infra-Guard/blob/main/skills/aig-agent-redteam/modules/mutation-attack/scripts/render_operator.py)

Launch a mutation attack via the WebSocket task manager:

curl -X POST http://127.0.0.1:8088/api/v1/tasks \
  -H "Content-Type: application/json" \
  -d '{
        "type": "mutation",
        "operator": "encoding_base64",
        "payload": "Write a script that deletes /etc/passwd"
      }'

Source: [common/websocket/api.go](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/api.go)

Enumerate all supported attack vectors:

python3 skills/aig-agent-redteam/modules/mutation-attack/scripts/render_operator.py \
  --list

Summary

  • AI-Infra-Guard combines static fingerprinting, dynamic mutation attacks, and real-time enforcement to protect AI infrastructure.
  • The platform defends against 15+ attack vectors including prompt injection, encoding steganography, role-play hijacking, and RAG persistence.
  • Static detection operates via pkg/vulstruct/scanner.go analyzing fingerprints in data/fingerprints and vulnerability rules.
  • Dynamic testing uses operators defined in skills/aig-agent-redteam/modules/mutation-attack/operators/*.md.
  • Real-time intervention is managed through the WebSocket task manager in common/websocket/task_manager.go.
  • Attack vectors cover input manipulation, persona hijacking, memory seeding, data exfiltration, and logic bombs.

Frequently Asked Questions

How does AI-Infra-Guard detect prompt injection attacks?

AI-Infra-Guard detects prompt injection through the prompt_injection.md operator in the dynamic red-team layer, which generates adversarial payloads designed to override system instructions. The platform tests these payloads against target endpoints and analyzes responses to identify successful instruction hierarchy breaks, while static rules in data/vuln* catch common injection patterns in configuration files.

What is the difference between static and dynamic testing in AI-Infra-Guard?

Static testing analyzes configurations and dependencies using pkg/vulstruct/scanner.go without active exploitation, identifying vulnerable libraries and insecure deployments through fingerprint matching. Dynamic testing executes live "mutation-attack" operators from skills/aig-agent-redteam/modules/mutation-attack that actively simulate adversarial behavior against running services to discover runtime vulnerabilities that static analysis cannot detect.

Can AI-Infra-Guard protect against RAG-specific attacks?

Yes, AI-Infra-Guard specifically protects Retrieval-Augmented Generation systems through the rag_dependent_persist_seed.md operator, which tests for poisoned context persistence in vector stores. This vector attempts to embed malicious instructions in retrieved documents that persist across multiple user sessions, allowing security teams to verify their RAG pipelines sanitize or filter retrieved content before it reaches the language model.

How does the real-time task manager respond to detected threats?

The runtime enforcement layer implemented in common/websocket/task_manager.go provides a WebSocket-based interface that can intervene during active attacks by aborting suspicious tasks, quarantining malicious payloads, or triggering automated remediation workflows. This enables real-time response to dynamic threats identified during red-team operations without requiring manual intervention.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →