# AI Infrastructure Vulnerabilities Detected by AI-Infra-Guard: CVE Coverage and Technical Implementation

> AI-Infra-Guard finds AI infrastructure vulnerabilities like race conditions in vLLM and memory flaws in Triton. Discover CVE coverage and technical details to secure your AI systems.

- Repository: [Tencent/AI-Infra-Guard](https://github.com/tencent/AI-Infra-Guard)
- Tags: deep-dive
- Published: 2026-08-22

---

**AI-Infra-Guard scans for critical AI infrastructure vulnerabilities—including race conditions in vLLM, memory-release flaws in NVIDIA Triton, and authentication bypasses in Upsonic—by matching component versions against a curated YAML database of CVEs.**

AI-Infra-Guard is an open-source security scanner maintained by Tencent that specializes in identifying AI infrastructure vulnerabilities across popular machine learning serving stacks. The tool ships with a built-in vulnerability database located in `data/vuln_en/` and uses a Go-based scanning engine to detect outdated or misconfigured AI components before they become exploitable in production environments.

## vLLM Serving Engine Vulnerabilities

AI-Infra-Guard detects multiple high-impact security flaws in **vLLM**, a popular large-language-model serving engine. The scanner evaluates version strings against specific CVE rules to identify vulnerable deployments.

- **CVE-2026-73557**: Race condition in `safe_load_prompt_embeds` that can lead to malformed tensor processing. The scanner flags versions `>= "0.20.2rc0" && < "0.26.0"` and recommends upgrading to vLLM 0.26.0 or later.

- **CVE-2026-73558**: Unsafe handling of sparse tensors when `enable_prompt_embeds` is enabled. Affects the same version range as CVE-2026-73557 (`0.20.2rc0` to `0.26.0`).

- **CVE-2026-73559**: Unsafe deserialization of prompts leading to denial-of-service conditions. Impacts versions between `0.20.2rc0` and `0.26.0`.

- **CVE-2026-7141**: Insecure default configuration permitting arbitrary code execution. The rule matches versions `<= "0.24.0"`.

Each vulnerability definition resides in the `data/vuln_en/vllm/` directory as individual YAML files containing version constraints and remediation advice.

## NVIDIA Triton Inference Server Flaws

The scanner monitors **NVIDIA Triton Inference Server** deployments for memory management and input validation vulnerabilities:

- **CVE-2026-47482**: Memory-release denial-of-service vulnerability (CWE-401) affecting versions `<= "26.04"`. Defined in [`data/vuln_en/triton-inference-server/CVE-2026-47482.yaml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/data/vuln_en/triton-inference-server/CVE-2026-47482.yaml).

- **CVE-2025-33238**: Unsafe handling of malformed model files leading to potential crashes or code execution in versions `<= "25.09"`.

## Additional AI Platform Vulnerabilities

AI-Infra-Guard extends coverage beyond core inference engines to include emerging AI development platforms:

- **Upsonic (CVE-2026-30625)**: Detects insecure default authentication configurations allowing unauthenticated API access in versions `< "2.1.0"`.

- **WeKnora (CVE-2026-30860)**: Identifies directory-traversal bugs enabling arbitrary file exposure in versions `<= "5.3.2"`.

- **WeKnora (CVE-2026-30861)**: Flags insecure deserialization vulnerabilities leading to remote code execution in versions `<= "5.3.2"`.

## How the Scanning Engine Works

The vulnerability detection process follows a three-stage pipeline implemented in [`pkg/vulstruct/scanner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/pkg/vulstruct/scanner.go):

1. **Version Extraction**: The scanning agent collects version strings from target services via HTTP headers, `/info` endpoints, or package manifest files.

2. **Rule Evaluation**: The core engine loads YAML vulnerability definitions from `data/vuln_en/` and parses the `rule` field—expressions like `version >= "0.20.2rc0" && version < "0.26.0"`—evaluating them against extracted component versions.

3. **Report Generation**: Matched vulnerabilities return as `vulstruct.Info` objects, rendered into JSON, SARIF, or custom `<vuln>` XML formats for downstream AI-agent consumption.

The architecture separates concerns between the **data layer** (human-editable YAML files), the **parser layer** ([`pkg/vulstruct/scanner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/pkg/vulstruct/scanner.go)), and the **API layer** ([`common/websocket/knowledge_api.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/knowledge_api.go)), enabling rapid updates to vulnerability signatures without recompiling the scanner.

## CLI Usage and Integration Examples

List all vulnerability templates to verify coverage before scanning:

```bash
ai-infra-guard scan --list-vul

```

Scan a local vLLM instance and receive structured vulnerability reports:

```bash
ai-infra-guard scan -t http://127.0.0.1:8000

```

Generate SARIF output for CI/CD pipeline integration:

```bash
ai-infra-guard scan -t http://my-triton:8000 --format sarif > result.sarif

```

Use the Python wrapper for programmatic access within agent frameworks:

```bash
python -m agent-scan.main --repo /path/to/project --agent_provider /path/to/provider.yaml

```

## Summary

- AI-Infra-Guard maintains a curated database of AI infrastructure vulnerabilities in `data/vuln_en/`, covering vLLM, NVIDIA Triton, Upsonic, and WeKnora.

- The scanner detects specific CVEs including race conditions (CVE-2026-73557), memory-release DoS (CVE-2026-47482), and authentication bypasses (CVE-2026-30625) through semantic version matching.

- Implementation resides in [`pkg/vulstruct/scanner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/pkg/vulstruct/scanner.go), which evaluates YAML-defined rules against discovered component versions.

- Output formats include JSON, SARIF, and custom XML, facilitating integration with both manual workflows and automated CI/CD pipelines.

## Frequently Asked Questions

### What types of AI infrastructure vulnerabilities does AI-Infra-Guard primarily detect?

AI-Infra-Guard focuses on version-specific CVEs affecting AI serving and orchestration platforms. According to the Tencent/AI-Infra-Guard source code, it detects race conditions in tensor processing, unsafe deserialization leading to remote code execution, directory traversal vulnerabilities, and insecure default configurations that permit unauthenticated access.

### How does AI-Infra-Guard determine if a component is vulnerable?

The scanner extracts version strings from target endpoints or package files, then evaluates them against constraints defined in YAML files under `data/vuln_en/`. The [`pkg/vulstruct/scanner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/pkg/vulstruct/scanner.go) engine parses rules such as `version >= "0.20.2rc0" && version < "0.26.0"` to flag vulnerable vLLM deployments or similar version ranges for other components.

### Can AI-Infra-Guard integrate with existing CI/CD pipelines?

Yes. The CLI supports SARIF output via the `--format sarif` flag, compatible with GitHub Advanced Security and other static analysis integrations. The Python wrapper (`agent-scan`) also enables programmatic invocation for automated security testing within build processes.

### Where are the vulnerability definitions stored and how are they updated?

Vulnerability rules reside in the `data/vuln_en/` directory as individual YAML files (e.g., [`data/vuln_en/vllm/CVE-2026-73557.yaml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/data/vuln_en/vllm/CVE-2026-73557.yaml)). This modular data layer allows security teams to add new CVEs by committing new YAML definitions without modifying the Go source code in [`pkg/vulstruct/scanner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/pkg/vulstruct/scanner.go).