AI Infrastructure Vulnerabilities Detected by AI-Infra-Guard: CVE Coverage and Technical Implementation

AI-Infra-Guard scans for critical AI infrastructure vulnerabilities—including race conditions in vLLM, memory-release flaws in NVIDIA Triton, and authentication bypasses in Upsonic—by matching component versions against a curated YAML database of CVEs.

AI-Infra-Guard is an open-source security scanner maintained by Tencent that specializes in identifying AI infrastructure vulnerabilities across popular machine learning serving stacks. The tool ships with a built-in vulnerability database located in data/vuln_en/ and uses a Go-based scanning engine to detect outdated or misconfigured AI components before they become exploitable in production environments.

vLLM Serving Engine Vulnerabilities

AI-Infra-Guard detects multiple high-impact security flaws in vLLM, a popular large-language-model serving engine. The scanner evaluates version strings against specific CVE rules to identify vulnerable deployments.

  • CVE-2026-73557: Race condition in safe_load_prompt_embeds that can lead to malformed tensor processing. The scanner flags versions >= "0.20.2rc0" && < "0.26.0" and recommends upgrading to vLLM 0.26.0 or later.

  • CVE-2026-73558: Unsafe handling of sparse tensors when enable_prompt_embeds is enabled. Affects the same version range as CVE-2026-73557 (0.20.2rc0 to 0.26.0).

  • CVE-2026-73559: Unsafe deserialization of prompts leading to denial-of-service conditions. Impacts versions between 0.20.2rc0 and 0.26.0.

  • CVE-2026-7141: Insecure default configuration permitting arbitrary code execution. The rule matches versions <= "0.24.0".

Each vulnerability definition resides in the data/vuln_en/vllm/ directory as individual YAML files containing version constraints and remediation advice.

NVIDIA Triton Inference Server Flaws

The scanner monitors NVIDIA Triton Inference Server deployments for memory management and input validation vulnerabilities:

  • CVE-2026-47482: Memory-release denial-of-service vulnerability (CWE-401) affecting versions <= "26.04". Defined in data/vuln_en/triton-inference-server/CVE-2026-47482.yaml.

  • CVE-2025-33238: Unsafe handling of malformed model files leading to potential crashes or code execution in versions <= "25.09".

Additional AI Platform Vulnerabilities

AI-Infra-Guard extends coverage beyond core inference engines to include emerging AI development platforms:

  • Upsonic (CVE-2026-30625): Detects insecure default authentication configurations allowing unauthenticated API access in versions < "2.1.0".

  • WeKnora (CVE-2026-30860): Identifies directory-traversal bugs enabling arbitrary file exposure in versions <= "5.3.2".

  • WeKnora (CVE-2026-30861): Flags insecure deserialization vulnerabilities leading to remote code execution in versions <= "5.3.2".

How the Scanning Engine Works

The vulnerability detection process follows a three-stage pipeline implemented in pkg/vulstruct/scanner.go:

  1. Version Extraction: The scanning agent collects version strings from target services via HTTP headers, /info endpoints, or package manifest files.

  2. Rule Evaluation: The core engine loads YAML vulnerability definitions from data/vuln_en/ and parses the rule field—expressions like version >= "0.20.2rc0" && version < "0.26.0"—evaluating them against extracted component versions.

  3. Report Generation: Matched vulnerabilities return as vulstruct.Info objects, rendered into JSON, SARIF, or custom <vuln> XML formats for downstream AI-agent consumption.

The architecture separates concerns between the data layer (human-editable YAML files), the parser layer (pkg/vulstruct/scanner.go), and the API layer (common/websocket/knowledge_api.go), enabling rapid updates to vulnerability signatures without recompiling the scanner.

CLI Usage and Integration Examples

List all vulnerability templates to verify coverage before scanning:

ai-infra-guard scan --list-vul

Scan a local vLLM instance and receive structured vulnerability reports:

ai-infra-guard scan -t http://127.0.0.1:8000

Generate SARIF output for CI/CD pipeline integration:

ai-infra-guard scan -t http://my-triton:8000 --format sarif > result.sarif

Use the Python wrapper for programmatic access within agent frameworks:

python -m agent-scan.main --repo /path/to/project --agent_provider /path/to/provider.yaml

Summary

  • AI-Infra-Guard maintains a curated database of AI infrastructure vulnerabilities in data/vuln_en/, covering vLLM, NVIDIA Triton, Upsonic, and WeKnora.

  • The scanner detects specific CVEs including race conditions (CVE-2026-73557), memory-release DoS (CVE-2026-47482), and authentication bypasses (CVE-2026-30625) through semantic version matching.

  • Implementation resides in pkg/vulstruct/scanner.go, which evaluates YAML-defined rules against discovered component versions.

  • Output formats include JSON, SARIF, and custom XML, facilitating integration with both manual workflows and automated CI/CD pipelines.

Frequently Asked Questions

What types of AI infrastructure vulnerabilities does AI-Infra-Guard primarily detect?

AI-Infra-Guard focuses on version-specific CVEs affecting AI serving and orchestration platforms. According to the Tencent/AI-Infra-Guard source code, it detects race conditions in tensor processing, unsafe deserialization leading to remote code execution, directory traversal vulnerabilities, and insecure default configurations that permit unauthenticated access.

How does AI-Infra-Guard determine if a component is vulnerable?

The scanner extracts version strings from target endpoints or package files, then evaluates them against constraints defined in YAML files under data/vuln_en/. The pkg/vulstruct/scanner.go engine parses rules such as version >= "0.20.2rc0" && version < "0.26.0" to flag vulnerable vLLM deployments or similar version ranges for other components.

Can AI-Infra-Guard integrate with existing CI/CD pipelines?

Yes. The CLI supports SARIF output via the --format sarif flag, compatible with GitHub Advanced Security and other static analysis integrations. The Python wrapper (agent-scan) also enables programmatic invocation for automated security testing within build processes.

Where are the vulnerability definitions stored and how are they updated?

Vulnerability rules reside in the data/vuln_en/ directory as individual YAML files (e.g., data/vuln_en/vllm/CVE-2026-73557.yaml). This modular data layer allows security teams to add new CVEs by committing new YAML definitions without modifying the Go source code in pkg/vulstruct/scanner.go.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →