AI-Infra-Guard Best Practices: A Complete Guide to Secure Deployment and Scanning

AI-Infra-Guard best practices include deploying behind a firewall with Docker, managing tasks via the HTTP API or CLI, validating YAML rules with the built-in checker, and tuning the Go worker pool to prevent resource exhaustion.

AI-Infra-Guard (AIG) is a hybrid-stack AI red-team platform developed by Tencent that combines a Go-based core engine with Python scanning modules. Understanding its architecture—from the WebSocket task management in common/websocket/server.go to the plugin loader in common/fingerprints/preload/preload.go—is essential for operating the tool safely in production environments.

Secure Deployment Architecture

The platform ships without built-in authentication mechanisms, making network isolation critical.

Container-First Deployment

Use the official Docker Compose configuration to ensure consistent dependencies and resource isolation.

curl https://raw.githubusercontent.com/Tencent/AI-Infra-Guard/refs/heads/main/docker.sh | bash

This script installs Docker if missing and launches the service via docker compose -f docker-compose.images.yml up -d.

Network Security

  • Run AI-Infra-Guard behind a firewall or VPN; expose port 8088 only on trusted LAN segments
  • allocate ≥ 4 GB RAM and ≥ 10 GB disk per the Quick-Start specifications in the repository root

Task Management via API and CLI

The core engine in cmd/cli/main.go exposes both a command-line interface and RESTful endpoints for orchestrating scans.

Creating Tasks Programmatically

Initiate infrastructure scans through the HTTP API rather than manual UI interaction:

curl -X POST http://localhost:8088/api/v1/tasks \
  -H "Content-Type: application/json" \
  -d '{"type":"infra","target":"http://127.0.0.1:8000"}'

Monitor real-time progress through the Server-Sent Events (SSE) endpoint at /api/v1/events or poll the task status endpoint:

curl http://localhost:8088/api/v1/tasks/<task-id>

CLI Execution

For ad-hoc scanning, use the binary directly with explicit target flags:

./ai-infra-guard scan -t http://127.0.0.1:8088

The -t parameter specifies the target URL of the AI service (e.g., a vLLM endpoint).

Concurrency Control

Prevent out-of-memory (OOM) errors in production by adjusting the worker pool. Edit the max_workers configuration in cmd/cli/main.go or pass the --workers flag to limit parallel scan operations:

./ai-infra-guard scan --workers 4 -t http://target:8080

Extending Detection Rules

AI-Infra-Guard uses YAML-based rule files loaded dynamically at runtime. The plugin system in common/fingerprints/preload/preload.go automatically discovers new definitions placed in the data/ directory.

Adding Custom Fingerprints

Create new service identification rules in data/fingerprints/:

name: custom-service
type: http
fingerprint:
  - header: "Server"
    pattern: "CustomAI/([0-9\.]+)"

Always validate syntax before deployment using the dedicated checker tool:

go build -o yamlcheck ./cmd/yamlcheck
./yamlcheck data/fingerprints

Vulnerability and MCP Rules

  • Place CVE definitions in data/vuln/, ensuring the id field follows the existing GSHA schema
  • Store MCP security plugins in data/mcp/; these are auto-loaded on service startup without requiring recompilation

Python Scanner Integration

The repository distributes three standalone Python packages (aig-skill-scan, aig-mcp-scan, aig-agent-scan) that operate independently or call the Go service via HTTP.

Installation and Configuration

Install via pip and configure the LLM API key for semantic analysis features:

pip install aig-skill-scan
export LLM_API_KEY="YOUR_KEY"

CI/CD Pipeline Usage

Run scanners in containerized environments to ensure reproducible results and capture structured output:

aig-skill-scan --repo ./my-skill \
               -m deepseek-v4-flash \
               --language en \
               -o skill-report.json

Store skill-report.json as a pipeline artifact for downstream security gates.

Performance Optimization

Model selection and caching strategies significantly impact detection accuracy and scan velocity.

LLM Selection

Higher-performing models (Claude Opus 4.6, GLM 5.1) yield superior F1 scores for vulnerability detection compared to lightweight alternatives. Specify models explicitly via the -m flag in Python scanners or the model field in API payloads.

HTTP Caching

Enable cache headers (CACHE_CONTROL) to avoid redundant fingerprint lookups when rescanning identical targets. The HTTP client implementation in pkg/httpx/httpx.go respects these directives to reduce network overhead.

Monitoring and Logging

The platform uses the centralized gologger package defined in common/gologger/gologger.go for structured logging.

Log Level Configuration

Enable debug traces during troubleshooting:

./ai-infra-guard webserver --log-level debug

Configure external log rotation (e.g., Docker logging drivers) to prevent disk exhaustion, as the service generates verbose output during active scans.

Summary

  • Deploy via Docker using the provided docker.sh script and isolate the service behind a firewall
  • Manage tasks through the HTTP API (POST /api/v1/tasks) or CLI, tracking progress via SSE endpoints
  • Validate rules with yamlcheck before placing new fingerprints in data/fingerprints/ or vulnerabilities in data/vuln/
  • Control concurrency by adjusting max_workers in cmd/cli/main.go or using the --workers CLI flag
  • Integrate Python scanners in CI/CD pipelines with explicit model selection and JSON output capture
  • Monitor logs using the gologger framework with appropriate level filtering and external rotation

Frequently Asked Questions

Does AI-Infra-Guard require authentication for the web interface?

No, the platform does not implement built-in authentication mechanisms. You must deploy it behind a reverse proxy with OAuth, VPN access, or firewall rules that restrict port 8088 to authorized networks only.

How do I add a new vulnerability check without restarting the service?

Place your YAML rule file in data/vuln/ following the existing schema. The plugin loader in common/fingerprints/preload/preload.go dynamically loads these files at runtime, though you may need to trigger a new scan task for the rules to take effect on existing targets.

What is the difference between the Go core and Python scanners?

The Go core (cmd/cli/main.go) provides the web server, task orchestration, and infrastructure fingerprinting capabilities. The Python scanners (skill-scan, mcp-scan, agent-scan) perform specialized semantic analysis of AI skills and agents, often requiring LLM API keys for operation. They can run standalone or feed results back to the Go service.

How can I prevent AI-Infra-Guard from exhausting system memory during large scans?

Limit the worker pool size using the --workers flag or modify max_workers in the source before compilation. Additionally, avoid running multiple concurrent large-scale scans; queue tasks sequentially and monitor memory usage through the SSE status endpoints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →