How to Customize AI-Infra-Guard Rules: Complete YAML Configuration Guide
AI-Infra-Guard stores detection logic as YAML rule files under the data/ directory, allowing you to customize AI-Infra-Guard rules by adding or modifying fingerprint definitions in data/fingerprints/ and vulnerability entries in data/vuln/ or data/vuln_en/, with changes auto-detected on service restart.
Tencent/AI-Infra-Guard is an open-source security scanner designed to identify AI infrastructure components and their vulnerabilities. The project uses a flexible rule-based architecture that enables security teams to extend detection capabilities without modifying the core Go codebase. You can customize AI-Infra-Guard rules by creating or editing YAML files that the scanner loads at runtime from configurable directories.
Understanding AI-Infra-Guard Rule Types
The scanner distinguishes between two primary rule categories, each serving distinct detection purposes. Both types share a common parsing schema implemented in common/fingerprints/parser/parser.go and load automatically when the service starts.
Fingerprint Rules
Fingerprint rules identify specific AI services, frameworks, or components by matching HTTP responses, headers, or body content. These rules reside in data/fingerprints/*.yaml, such as data/fingerprints/vllm.yaml. Each file defines matchers that recognize unique characteristics of target services.
Vulnerability Rules
Vulnerability rules catalog known CVEs and security issues affecting specific component versions. These are stored in data/vuln/*.yaml (Chinese descriptions) and data/vuln_en/*.yaml (English descriptions), exemplified by data/vuln_en/weknora/CVE-2026-30861.yaml. Entries map CVE identifiers to affected version ranges and severity ratings.
Step-by-Step Guide to Customizing AI-Infra-Guard Rules
Follow this workflow to add custom detection logic to your AI-Infra-Guard deployment. The process requires no Go programming knowledge—only YAML editing and file system access.
1. Create Rule Files in the Appropriate Directory
Place new fingerprint definitions under data/fingerprints/ and vulnerability entries under data/vuln/ or data/vuln_en/. Use existing files as templates to ensure schema compliance. The scanner recursively loads all .yaml files from these locations at startup.
2. Define Required Rule Fields
For fingerprint rules, specify the service name, detection type (typically http), and matchers containing regex patterns, header checks, or body content conditions.
For vulnerability rules, include the cve identifier, description, severity level, affectedVersions list, and reference URLs.
3. Validate Syntax Using yamlcheck
Run the built-in validation tool before deploying changes. The yamlcheck binary, defined in cmd/yamlcheck/main.go, verifies YAML syntax and schema compliance.
./yamlcheck data/fingerprints data/vuln data/vuln_en
Any structural errors or missing required fields will be reported with file paths and line numbers.
4. Reload the Scanner and Verify
Restart the AI-Infra-Guard server or re-run the CLI command to load updated rules. The runner.go module in common/runner/ handles the file loading process. Verify your changes by querying the knowledge base API or running a test scan against a known target:
# Check loaded fingerprints via API
curl http://localhost:8080/api/v1/knowledge/fingerprints
# Run scan with custom fingerprint directory
./ai-infra-guard scan -t http://example.com -fps ./my_custom_rules
Rule File Structure Examples
Use these templates when creating custom rules to ensure compatibility with the parser package.
Minimal Fingerprint Template
name: example-service
type: http
matchers:
- method: GET
url: "/api/v1/status"
header:
X-Example: "true"
body: "example version"
This example demonstrates HTTP-based detection using header and body matchers. Store this content as data/fingerprints/example.yaml to detect the service immediately upon restart.
Minimal Vulnerability Template
cve: CVE-2025-0001
description: "Example service remote code execution"
severity: critical
affectedVersions:
- "1.0.0"
- "1.1.x"
references:
- "https://example.com/issue/details"
Place vulnerability definitions in hierarchical directories like data/vuln_en/vendor-name/ for organizational clarity. The scanner indexes these by CVE ID for correlation with detected fingerprints.
Technical Implementation of Rule Loading
Understanding the underlying code helps troubleshoot custom rule issues. The entry point in cmd/cli/main.go initializes the scanning process, while common/runner/runner.go orchestrates rule loading.
The parser package exposes LoadFromDir() to recursively read YAML files:
import "github.com/Tencent/AI-Infra-Guard/common/fingerprints/parser"
func LoadFingerprints(dir string) ([]parser.FingerPrint, error) {
fps, err := parser.LoadFromDir(dir) // dir = "data/fingerprints"
if err != nil {
return nil, err
}
return fps, nil
}
By default, the CLI uses the --fps (or -fps) flag to specify the fingerprint directory, defaulting to data/fingerprints. You can override this to use entirely custom rule sets without modifying the original distribution files.
Summary
- AI-Infra-Guard rules are YAML files stored in
data/fingerprints/,data/vuln/, anddata/vuln_en/directories according to the Tencent/AI-Infra-Guard source code. - Fingerprint rules identify AI components via HTTP matchers, while vulnerability rules map CVEs to affected versions.
- Use the
./yamlcheckvalidation tool to verify syntax before deployment. - The scanner loads rules automatically from directories specified by the
--fpsflag on startup. - Changes take effect immediately after service restart without requiring code recompilation.
Frequently Asked Questions
Where are AI-Infra-Guard rules stored by default?
Default rules reside in the data/ directory of the repository. Fingerprint definitions are located in data/fingerprints/*.yaml, while vulnerability entries are split between data/vuln/*.yaml (Chinese) and data/vuln_en/*.yaml (English). You can override these locations using the -fps CLI flag or by modifying the configuration in common/runner/runner.go.
How do I validate custom rule syntax before deployment?
AI-Infra-Guard includes a dedicated yamlcheck binary compiled from cmd/yamlcheck/main.go. Run ./yamlcheck data/fingerprints data/vuln data/vuln_en to validate all rule files against the parser schema. The tool reports specific line numbers and structural errors, preventing runtime loading failures.
Do I need to restart the service after updating rules?
Yes, you must restart the AI-Infra-Guard server or re-run the CLI command. The parser loads rules from disk only during initialization via parser.LoadFromDir() as implemented in common/fingerprints/parser/parser.go. There is no hot-reload mechanism; changes are picked up on the next startup cycle.
Can I use custom rule directories without modifying the original data folder?
Yes, specify an alternative directory using the -fps flag when running scans. For example: ./ai-infra-guard scan -t http://example.com -fps ./my_custom_rules. This allows you to maintain proprietary detection logic outside the main repository while still leveraging the core scanning engine from cmd/cli/main.go.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →