How to Contribute Custom Rules to AI-Infra-Guard: A Complete Guide
To contribute custom rules to AI-Infra-Guard, create a YAML file in the appropriate data/ subdirectory, validate it locally using the yamlcheck tool in cmd/yamlcheck/main.go, and submit a pull request for automated CI review.
Tencent/AI-Infra-Guard relies on a YAML-based rule system to detect service fingerprints and vulnerability patterns. All rules reside in the data/ directory and must pass strict validation before merging into the main branch. This guide walks through the exact workflow for contributing new detection rules while adhering to the project's quality standards.
Understanding the Rule Architecture
AI-Infra-Guard organizes detection logic into two distinct rule categories. Service fingerprint rules identify AI infrastructure components like model serving frameworks, while vulnerability rules match specific CVE patterns or security misconfigurations.
Rule files follow a single-purpose YAML document structure. Fingerprint definitions live in data/fingerprints/*.yaml, vulnerability signatures reside in data/vuln/*.yaml (Chinese) and data/vuln_en/*.yaml (English), and the cmd/yamlcheck/main.go utility validates syntax against the internal schema documented in its Rule Schema comment block.
Step-by-Step Contribution Workflow
1. Fork and Clone the Repository
Create a personal fork of Tencent/AI-Infra-Guard on GitHub, then clone your fork locally. This establishes your development environment for creating and testing rules.
2. Create Your Rule File
Determine whether you are adding a service fingerprint or a vulnerability signature. Create a new YAML file in the corresponding directory:
- Fingerprints:
data/fingerprints/your-service.yaml - Vulnerabilities:
data/vuln/your-cve.yamlordata/vuln_en/your-cve.yaml
Use existing files like data/fingerprints/vllm.yaml as templates for proper structure and naming conventions.
3. Define the Rule Schema
Populate the required info section and at least one http section. The info block requires name, author, severity, and desc fields, while the http array defines request methods, paths, and response matchers. Optionally include a version section to extract version strings via regular expressions.
4. Validate Locally with yamlcheck
Run the built-in validation tool to catch syntax errors before submission:
go run cmd/yamlcheck/main.go data/fingerprints data/vuln data/vuln_en
This command parses all YAML files against the strict schema defined in cmd/yamlcheck/main.go and reports validation failures with line-specific diagnostics.
5. Run Unit Tests
Execute the full test suite to ensure your rule integrates correctly with the parser:
go test ./...
The test suite automatically loads all YAML files under data/ and verifies that the rule engine can parse them without runtime errors.
6. Submit Your Pull Request
Commit your new rule file and push to your fork. Open a pull request targeting the upstream main branch. In the description, explain the rule's purpose, reference the relevant CVE or product, and confirm that local validation passed. The CI pipeline defined in .github/workflows/yaml-lint.yml will automatically run yamlcheck and the Go test matrix.
Rule Structure and Schema
Below is a minimal fingerprint rule for a hypothetical service. Save this as data/fingerprints/my-service.yaml:
info:
name: my-service
author: Your Name
severity: info
desc: Detects the My-Service API endpoint.
metadata:
product: my-service
vendor: my-company
http:
- method: GET
path: '/status'
matchers:
- body="\"status\":\"ok\""
version:
- method: GET
path: '/version'
extractor:
part: body
group: 1
regex: '"version":"([\d\.]+)"'
Key components explained:
info: Required metadata describing the rule purpose, author, and severity level.http: Defines request patterns (method,path) and response matchers for detection logic.version: Optional extraction rules using regex groups to capture software version strings from responses.
Validation and Quality Gates
The yamlcheck binary serves as the primary quality gate for all rule contributions. Located in cmd/yamlcheck/main.go, this tool enforces schema compliance by verifying that every YAML file contains valid info and http structures before they enter the main branch.
When you open a pull request, GitHub Actions automatically executes the workflow defined in .github/workflows/yaml-lint.yml, which runs yamlcheck against the entire data/ tree. This prevents repository-wide breakage and maintains consistent rule quality across the project.
Summary
- AI-Infra-Guard uses YAML-based rules stored in
data/fingerprints/anddata/vuln/directories. - Always validate new rules locally using
go run cmd/yamlcheck/main.gobefore submitting. - Rule files require
infoandhttpsections, with optionalversionextractors for semantic versioning. - The CI pipeline in
.github/workflows/yaml-lint.ymlautomatically validates all pull requests. - Reference existing rules like
vllm.yamlfor structure and schema examples.
Frequently Asked Questions
What is the difference between fingerprint and vulnerability rules?
Fingerprint rules in data/fingerprints/ identify specific AI services and infrastructure components (like vLLM or Triton), while vulnerability rules in data/vuln/ and data/vuln_en/ match specific security flaws or CVE patterns. Both use the same YAML structure but serve different detection purposes in the scanning engine.
How do I validate my rule before submitting a PR?
Run the local validation command go run cmd/yamlcheck/main.go data/fingerprints data/vuln data/vuln_en from the repository root. This executes the schema validator defined in cmd/yamlcheck/main.go and reports syntax errors, missing required fields, or malformed regex patterns before you commit.
What fields are required in a rule YAML file?
Every rule must contain an info section with name, author, severity, and desc fields, plus at least one http array entry defining method, path, and matchers. The version section is optional but recommended for services where version extraction is possible.
Where can I find examples of existing rules?
Browse the data/fingerprints/ directory for service detection examples like vllm.yaml, or explore data/vuln/ for vulnerability signatures. These files demonstrate proper field usage, matcher syntax, and metadata conventions accepted by the AI-Infra-Guard parser.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →