What Security Threats Does AI-Infra-Guard Address? A Complete Technical Guide
AI-Infra-Guard addresses over 2,000 infrastructure CVEs, MCP protocol abuse, malicious skill logic, jailbreak attacks, and model/API relay threats through specialized scanners that fingerprint services, analyze code statically, and evaluate prompt injection resistance.
Tencent's AI-Infra-Guard (A.I.G) is a comprehensive AI red-team platform designed to identify and mitigate vulnerabilities across the entire AI stack. As AI infrastructure grows more complex, understanding what security threats AI-Infra-Guard addresses becomes critical for securing both public-facing services and internal agent ecosystems. The platform combines infrastructure scanning, static code analysis, and behavioral testing to protect against threats ranging from known CVEs in popular frameworks to sophisticated prompt injection attacks.
AI Infrastructure Vulnerabilities and CVE Detection
The platform's infra scanner probes target URLs and IPs to identify exposed AI services and match them against a curated vulnerability database. This addresses the foundational layer where AI frameworks like vLLM, Ollama, ComfyUI, and Triton Inference Server often run with default configurations or outdated versions.
Known CVE Detection in AI Frameworks
AI-Infra-Guard maintains a YAML-based vulnerability database in data/vuln/ containing over 2,000 known CVEs specific to AI frameworks. When the scanner detects a running service, it queries component versions and matches them against these rules. According to the source code in pkg/vulstruct/scanner.go, the fingerprinting engine gathers version strings from HTTP headers, API endpoints, and default pages, then correlates findings with the vulnerability definitions stored in the data directory.
Version Fingerprinting and API Misconfiguration
Beyond CVE matching, the scanner in cmd/cli/main.go identifies API exposure risks and configuration errors that could lead to unauthorized access. The fingerprint definitions reside in data/fingerprints/, allowing the tool to recognize specific AI service deployments even when they run on non-standard ports or behind reverse proxies.
MCP Server and Agent Skill Threats
The MCP scanner targets the Model Context Protocol (MCP) layer, where AI agents interact with external tools and services. This addresses a critical attack surface where malicious or poorly implemented skills can compromise entire agent ecosystems.
Tool Poisoning and Credential Exfiltration
Static analysis of MCP server code identifies tool-poisoning attempts and credential exfiltration patterns. The scanner applies threat rules defined in data/mcp/ (such as tool-poisoning.yaml) to detect when tools attempt to harvest environment variables, API keys, or session tokens. This covers one of 14 distinct categories of unsafe tool and chain usage that the platform monitors.
Command Injection and Malicious Dependencies
The analysis engine parses server implementations to find command injection vulnerabilities and malicious dependency imports. By examining package manifests and source code in MCP repositories, AI-Infra-Guard flags unsafe system calls, eval statements, and suspicious network requests that could enable remote code execution on the host system.
Skill-Level Logic Attacks (SkillTrustBench)
The skill-scan engine addresses threats within the skill code itself through the SkillTrustBench evaluation framework. This targets the logic layer where AI agents execute user-provided or third-party skills.
The T01-T09 Threat Taxonomy
Skills are evaluated against a comprehensive taxonomy covering nine threat categories:
- T01-T02: Instruction hijacking and memory manipulation attacks
- T03-T04: Code execution and payload download attempts
- T05-T06: Privilege escalation and persistence mechanisms
- T07-T08: Toolchain and dependency supply-chain attacks
- T09: Insecure coding practices and unsafe defaults
Call Graph Analysis
Located in the skill-scan/ directory, the static analysis pipeline parses skill source code to build a call graph, identifying execution flows that could lead to unsafe operations. Results are benchmarked against the SkillTrustBench leaderboard, providing quantitative risk scores for each skill evaluated.
Jailbreak and Prompt Injection Attacks
The prompt-security module addresses adversarial inputs designed to bypass safety filters or extract training data. This targets the interaction layer between users and language models.
Single-Turn and Multi-Turn Prompt Injection
Housed in AIG-PromptSecurity/, this component runs curated jailbreak datasets against target LLMs to measure resistance against prompt-injection attacks, role-play abuse scenarios, and data-leak exploits. The evaluation covers both single-turn attacks (immediate injection) and multi-turn conversations (gradual context manipulation) to test conversation state management vulnerabilities.
Model Integrity and API Relay Verification
The API checker validates that deployed models match their claimed identities and detects unauthorized relay or proxy services intercepting API calls.
Model Fingerprint Mismatches
The tool verifies cryptographic signatures and metadata consistency to detect model fingerprint mismatches, such as unauthorized Claude model signatures or modified weight files. This prevents scenarios where attackers substitute models with compromised versions that behave maliciously while presenting legitimate API endpoints.
Black-Box Relay Auditing
AI-Infra-Guard implements auditing methodologies like PAMELA and Ventor QTest to identify relay black-box services. These checks validate that API responses originate from genuine model providers rather than intermediate proxies that could log, modify, or mine request data.
Running Comprehensive Security Scans
AI-Infra-Guard exposes all functionality through a unified CLI entry point in cmd/cli/main.go. The following commands demonstrate how to execute each scan type:
# Scan an AI-Infra service (e.g., a running vLLM instance)
./ai-infra-guard scan -t http://127.0.0.1:8000
# Returns component versions and matched CVEs from data/vuln/*.yaml
# Run a Skill scan on a local skill repo
pip install aig-skill-scan
export LLM_API_KEY="your-api-key"
aig-skill-scan --repo ./my-skill \
-m deepseek-v4-flash \
--language en \
-o skill-result.json
# JSON output contains T01-T09 risk categories
# Perform MCP security scan on an MCP server repo
./ai-infra-guard mcp-scan -p /path/to/mcp-server
# Detects tool-poisoning, credential exfiltration, command injection, etc.
# Run a jailbreak evaluation against a target LLM
./ai-infra-guard jailbreak-eval \
--model-url http://localhost:8088/api/v1/models \
--dataset data/eval/jailbreak/
# Shows success-rate per attack method
All commands assume the A.I.G service is running via Docker Compose (docker-compose -f docker-compose.images.yml up -d), which exposes the REST API defined in common/websocket/api.go used by both the CLI and web interface.
Summary
- Infrastructure CVEs: AI-Infra-Guard detects over 2,000 known vulnerabilities in AI frameworks like vLLM and Ollama through version fingerprinting and YAML rule matching in
data/vuln/. - MCP Protocol Abuse: The static analyzer in
data/mcp/identifies 14 categories of threats including tool-poisoning and credential exfiltration in agent skills. - SkillTrustBench Coverage: The T01-T09 taxonomy evaluates skills for code execution, privilege escalation, and supply-chain attacks via call-graph analysis.
- Prompt Security: The jailbreak evaluation module tests resistance against single-turn and multi-turn prompt injection using curated datasets.
- API Integrity: Model fingerprint verification and relay auditing ensure deployed LLMs match their claimed identities and aren't proxied by unauthorized intermediates.
Frequently Asked Questions
Does AI-Infra-Guard detect vulnerabilities in vLLM and Ollama deployments?
Yes. The infrastructure scanner specifically targets popular AI serving frameworks including vLLM, Ollama, ComfyUI, and Triton Inference Server. By probing running instances and comparing version fingerprints against the data/vuln/ database, the tool identifies CVEs affecting these specific platforms.
How does the MCP scanner identify tool-poisoning attacks?
The MCP scanner performs static analysis on server code using threat rules defined in data/mcp/tool-poisoning.yaml and related configuration files. It examines function definitions, import statements, and network calls to detect patterns where tools might exfiltrate credentials, execute arbitrary commands, or introduce malicious dependencies.
What is the T01-T09 taxonomy in SkillTrustBench?
The T01-T09 taxonomy is a nine-category classification system for skill-level security threats implemented in the skill-scan/ module. Categories T01-T02 cover instruction and memory hijacking, T03-T04 address code execution and payload downloads, T05-T06 handle privilege escalation and persistence, T07-T08 identify toolchain attacks, and T09 flags insecure coding practices.
Can AI-Infra-Guard test for prompt injection against custom LLMs?
Yes. The jailbreak-eval command accepts a custom --model-url parameter pointing to any OpenAI-compatible API endpoint. This allows security teams to test proprietary or fine-tuned models against standard jailbreak datasets stored in data/eval/jailbreak/ to measure attack success rates and safety robustness.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →