# How AI-Infra-Guard Manages Rules and Data: YAML Storage, Runtime Loading, and Live Updates

> Discover how AI-Infra-Guard manages rules and data using YAML storage, runtime loading, and live updates via WebSocket APIs. Learn about its efficient data management.

- Repository: [Tencent/AI-Infra-Guard](https://github.com/tencent/AI-Infra-Guard)
- Tags: internals
- Published: 2026-08-22

---

**AI-Infra-Guard stores all detection logic and vulnerability definitions as plain-text YAML files under the `data/` directory, loading them at runtime into the scanning engine while exposing CRUD operations via WebSocket APIs.**

Tencent's **AI-Infra-Guard** uses a data-driven architecture where **rules and data management** happens through declarative YAML files rather than hardcoded Go logic. This design allows security teams to add new service fingerprints or vulnerability advisories without recompiling the binary. All rule files live in the repository's top-level `data/` directory and are parsed into Go structs during scanner initialization.

## Where Rules and Data Are Stored

### Fingerprint Definitions

Service identification rules reside in `data/fingerprints/` (approximately 200 files). Each YAML file defines detection patterns, service names, descriptions, and optional severity ratings. For example, the LLama-CPP fingerprint is defined in [`data/fingerprints/llama-cpp.yaml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/data/fingerprints/llama-cpp.yaml). According to the source code in [`common/runner/runner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/runner/runner.go), the scanner reads this directory via the `fpDir` variable and parses each file using `common/fingerprints/parser`.

### Vulnerability Advisories

CVE-style entries are stored in `data/vuln/` (primary) and `data/vuln_en/` (English translations). These files contain version ranges, impact summaries, and remediation suggestions. The MLflow 2.8.0 advisory at [`data/vuln/mlflow/CVE-2023-6977.yaml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/data/vuln/mlflow/CVE-2023-6977.yaml) demonstrates the structure. During initialization, [`common/runner/runner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/runner/runner.go) calls `initVulnerabilityDB()`, which invokes `vulstruct.AdvisoryEngine.LoadFromDirectory(dir)` to populate the advisory engine.

### MCP-Specific Rules and Evaluation Data

The platform maintains plugin-level checks in `data/mcp/` for the MCP scanner, loaded by [`internal/mcp/plugins.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/internal/mcp/plugins.go). Evaluation datasets for prompt-security benchmarks are stored as JSON in `data/eval/` and consumed by `AIG-PromptSecurity` modules.

## Runtime Loading and Initialization

The CLI accepts configuration flags defined in [`internal/options/options.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/internal/options/options.go), including `--fps` for the fingerprint directory and `--vul` for the vulnerability directory. When the runner starts, [`common/runner/runner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/runner/runner.go) executes `ShowFpAndVulList()`, which orchestrates the loading sequence.

First, the fingerprint parser scans the directory and converts YAML definitions into Go structs. Then the vulnerability engine builds a version-based advisory index. This separation allows the scanning engine to match discovered services against fingerprints before checking version strings against the advisory database.

```go
// Runner initialization (excerpt from common/runner/runner.go)
fpDir := "data/fingerprints"
vulDir := "data/vuln"

fps, err := parser.LoadDirectory(fpDir)          // parses every *.yaml in fpDir
if err != nil { log.Fatalf("load fp: %v", err) }

advEngine, err := vulstruct.NewAdvisoryEngine()
if err != nil { log.Fatalf("adv engine: %v", err) }
if err = advEngine.LoadFromDirectory(vulDir); err != nil {
    log.Fatalf("load vulns: %v", err)
}

```

## CRUD Operations via WebSocket API

AI-Infra-Guard exposes full lifecycle management through the WebSocket API implemented in [`common/websocket/knowledge_api.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/knowledge_api.go). The system provides dedicated handlers for both data types:

- **Fingerprints**: `HandleListFingerprints`, `HandleCreateFingerprint`, `HandleEditFingerprint`, and `HandleDeleteFingerprint`
- **Vulnerabilities**: `HandleListVulnerabilities`, `HandleCreateVulnerability`, and corresponding update/delete methods

These endpoints enable live reload capabilities. When you create or modify a rule through the API, the change persists to disk and immediately reflects in subsequent scans without requiring a process restart.

```bash

# Create a new fingerprint via the WebSocket API

curl -X POST http://localhost:8088/api/v1/knowledge/fingerprints \
    -H "Content-Type: application/json" \
    -d '{
          "name": "myservice",
          "description": "Detects My Service",
          "match": {"url": "myservice.example.com"},
          "version": {"header": "X-My-Version"},
          "severity": "high"
        }'

```

## Version-Aware Matching and Scanning

During task execution, the engine produces a `Result` struct defined in [`common/runner/result.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/runner/result.go). This result includes both `fingerprints` and `vulnerabilities` arrays. The matching logic uses version constraints (e.g., `>=`, `<`) defined in the YAML advisories to determine if a detected service version falls within vulnerable ranges.

You can trigger scans via CLI and retrieve results through the API:

```bash

# Start a scan

./ai-infra-guard scan -t http://127.0.0.1:8088

# Retrieve results including fingerprints and vulnerabilities

curl http://localhost:8088/api/v1/tasks/<task-id>/result

# JSON response includes:

#   "fingerprints": [{...}],

#   "vulnerabilities": [{...}]

```

## Adding Rules Without Recompiling

Because all detection logic lives as **plain-text YAML**, you can extend coverage by dropping new files into the appropriate `data/` subdirectories. The built-in `yamlcheck` command validates schema compliance before deployment, ensuring structural integrity without code changes.

## Summary

- **AI-Infra-Guard** manages rules and data through YAML files in `data/fingerprints/`, `data/vuln/`, and `data/mcp/`.
- The runner initializes by loading these files via [`common/runner/runner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/runner/runner.go) and parsing them with `common/fingerprints/parser` and `pkg/vulstruct`.
- **WebSocket APIs** in [`common/websocket/knowledge_api.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/knowledge_api.go) provide real-time CRUD operations for fingerprints and vulnerabilities.
- **Version-aware matching** correlates detected service versions against CVE-style advisories using constraint operators.
- The architecture supports **live updates** and **hot-reloading** without binary recompilation.

## Frequently Asked Questions

### How do I add a new service fingerprint to AI-Infra-Guard?

Create a new YAML file in `data/fingerprints/` following the existing schema with fields for `name`, `description`, `match` patterns, and optional `version` extraction rules. Run the `yamlcheck` command to validate syntax, then restart the scanner or use the WebSocket API to create the fingerprint dynamically. The parser in [`common/fingerprints/parser/parser.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/fingerprints/parser/parser.go) will automatically load the new definition on the next scan cycle.

### Can I update vulnerability rules while the scanner is running?

Yes. The WebSocket API exposed in [`common/websocket/knowledge_api.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/knowledge_api.go) supports live updates through `HandleCreateVulnerability` and `HandleEditVulnerability` endpoints. When you modify an advisory, the system writes changes to the YAML file in `data/vuln/` and updates the in-memory advisory engine immediately, allowing subsequent scans to use the new rules without downtime.

### What is the difference between `data/vuln/` and `data/vuln_en/`?

The `data/vuln/` directory contains the primary vulnerability advisories, while `data/vuln_en/` stores English translations of the same entries. Both directories follow identical YAML schemas with CVE-style metadata, version constraints, and remediation guidance. The advisory engine can load from either location based on locale configuration.

### How does the scanner validate YAML rule syntax?

AI-Infra-Guard includes a built-in `yamlcheck` command that validates fingerprint and vulnerability files against the expected schema. This validation ensures that required fields like `name`, `match` conditions, and version constraints are present and correctly formatted before the runtime parser attempts to load them in [`common/runner/runner.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/runner/runner.go).