How AI-Infra-Guard Manages Rules and Data: YAML Storage, Runtime Loading, and Live Updates
AI-Infra-Guard stores all detection logic and vulnerability definitions as plain-text YAML files under the data/ directory, loading them at runtime into the scanning engine while exposing CRUD operations via WebSocket APIs.
Tencent's AI-Infra-Guard uses a data-driven architecture where rules and data management happens through declarative YAML files rather than hardcoded Go logic. This design allows security teams to add new service fingerprints or vulnerability advisories without recompiling the binary. All rule files live in the repository's top-level data/ directory and are parsed into Go structs during scanner initialization.
Where Rules and Data Are Stored
Fingerprint Definitions
Service identification rules reside in data/fingerprints/ (approximately 200 files). Each YAML file defines detection patterns, service names, descriptions, and optional severity ratings. For example, the LLama-CPP fingerprint is defined in data/fingerprints/llama-cpp.yaml. According to the source code in common/runner/runner.go, the scanner reads this directory via the fpDir variable and parses each file using common/fingerprints/parser.
Vulnerability Advisories
CVE-style entries are stored in data/vuln/ (primary) and data/vuln_en/ (English translations). These files contain version ranges, impact summaries, and remediation suggestions. The MLflow 2.8.0 advisory at data/vuln/mlflow/CVE-2023-6977.yaml demonstrates the structure. During initialization, common/runner/runner.go calls initVulnerabilityDB(), which invokes vulstruct.AdvisoryEngine.LoadFromDirectory(dir) to populate the advisory engine.
MCP-Specific Rules and Evaluation Data
The platform maintains plugin-level checks in data/mcp/ for the MCP scanner, loaded by internal/mcp/plugins.go. Evaluation datasets for prompt-security benchmarks are stored as JSON in data/eval/ and consumed by AIG-PromptSecurity modules.
Runtime Loading and Initialization
The CLI accepts configuration flags defined in internal/options/options.go, including --fps for the fingerprint directory and --vul for the vulnerability directory. When the runner starts, common/runner/runner.go executes ShowFpAndVulList(), which orchestrates the loading sequence.
First, the fingerprint parser scans the directory and converts YAML definitions into Go structs. Then the vulnerability engine builds a version-based advisory index. This separation allows the scanning engine to match discovered services against fingerprints before checking version strings against the advisory database.
// Runner initialization (excerpt from common/runner/runner.go)
fpDir := "data/fingerprints"
vulDir := "data/vuln"
fps, err := parser.LoadDirectory(fpDir) // parses every *.yaml in fpDir
if err != nil { log.Fatalf("load fp: %v", err) }
advEngine, err := vulstruct.NewAdvisoryEngine()
if err != nil { log.Fatalf("adv engine: %v", err) }
if err = advEngine.LoadFromDirectory(vulDir); err != nil {
log.Fatalf("load vulns: %v", err)
}
CRUD Operations via WebSocket API
AI-Infra-Guard exposes full lifecycle management through the WebSocket API implemented in common/websocket/knowledge_api.go. The system provides dedicated handlers for both data types:
- Fingerprints:
HandleListFingerprints,HandleCreateFingerprint,HandleEditFingerprint, andHandleDeleteFingerprint - Vulnerabilities:
HandleListVulnerabilities,HandleCreateVulnerability, and corresponding update/delete methods
These endpoints enable live reload capabilities. When you create or modify a rule through the API, the change persists to disk and immediately reflects in subsequent scans without requiring a process restart.
# Create a new fingerprint via the WebSocket API
curl -X POST http://localhost:8088/api/v1/knowledge/fingerprints \
-H "Content-Type: application/json" \
-d '{
"name": "myservice",
"description": "Detects My Service",
"match": {"url": "myservice.example.com"},
"version": {"header": "X-My-Version"},
"severity": "high"
}'
Version-Aware Matching and Scanning
During task execution, the engine produces a Result struct defined in common/runner/result.go. This result includes both fingerprints and vulnerabilities arrays. The matching logic uses version constraints (e.g., >=, <) defined in the YAML advisories to determine if a detected service version falls within vulnerable ranges.
You can trigger scans via CLI and retrieve results through the API:
# Start a scan
./ai-infra-guard scan -t http://127.0.0.1:8088
# Retrieve results including fingerprints and vulnerabilities
curl http://localhost:8088/api/v1/tasks/<task-id>/result
# JSON response includes:
# "fingerprints": [{...}],
# "vulnerabilities": [{...}]
Adding Rules Without Recompiling
Because all detection logic lives as plain-text YAML, you can extend coverage by dropping new files into the appropriate data/ subdirectories. The built-in yamlcheck command validates schema compliance before deployment, ensuring structural integrity without code changes.
Summary
- AI-Infra-Guard manages rules and data through YAML files in
data/fingerprints/,data/vuln/, anddata/mcp/. - The runner initializes by loading these files via
common/runner/runner.goand parsing them withcommon/fingerprints/parserandpkg/vulstruct. - WebSocket APIs in
common/websocket/knowledge_api.goprovide real-time CRUD operations for fingerprints and vulnerabilities. - Version-aware matching correlates detected service versions against CVE-style advisories using constraint operators.
- The architecture supports live updates and hot-reloading without binary recompilation.
Frequently Asked Questions
How do I add a new service fingerprint to AI-Infra-Guard?
Create a new YAML file in data/fingerprints/ following the existing schema with fields for name, description, match patterns, and optional version extraction rules. Run the yamlcheck command to validate syntax, then restart the scanner or use the WebSocket API to create the fingerprint dynamically. The parser in common/fingerprints/parser/parser.go will automatically load the new definition on the next scan cycle.
Can I update vulnerability rules while the scanner is running?
Yes. The WebSocket API exposed in common/websocket/knowledge_api.go supports live updates through HandleCreateVulnerability and HandleEditVulnerability endpoints. When you modify an advisory, the system writes changes to the YAML file in data/vuln/ and updates the in-memory advisory engine immediately, allowing subsequent scans to use the new rules without downtime.
What is the difference between data/vuln/ and data/vuln_en/?
The data/vuln/ directory contains the primary vulnerability advisories, while data/vuln_en/ stores English translations of the same entries. Both directories follow identical YAML schemas with CVE-style metadata, version constraints, and remediation guidance. The advisory engine can load from either location based on locale configuration.
How does the scanner validate YAML rule syntax?
AI-Infra-Guard includes a built-in yamlcheck command that validates fingerprint and vulnerability files against the expected schema. This validation ensures that required fields like name, match conditions, and version constraints are present and correctly formatted before the runtime parser attempts to load them in common/runner/runner.go.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →