AI-Infra-Guard Performance Benchmarks: Accuracy, Speed, and Security Metrics

AI-Infra-Guard achieves a 0.9848 top-score on the SkillTrustBench accuracy benchmark while supporting sub-second latency for MCP scans and multi-model jailbreak resistance testing.

AI-Infra-Guard (AIG) is a hybrid-stack AI security platform developed by Tencent that combines a Go-based core service with Python-based scanning modules to deliver comprehensive performance benchmarks. The platform measures both accuracy (how effectively it detects security vulnerabilities) and efficiency (how quickly it processes scans) across multiple dimensions including skill-based validation, jailbreak resistance, and harmful content detection. Understanding these performance benchmarks for AI-Infra-Guard is essential for deploying the tool in production environments where both precision and speed are critical.

Key Performance Dimensions

The platform evaluates performance across five primary technical dimensions, each targeting specific attack surfaces and operational workloads.

Skill-Scan Accuracy (SkillTrustBench)

The SkillTrustBench benchmark evaluates the precision of skill-based security checks within the scanning engine. According to the CHANGELOG.md at line 466, the latest release reports a top-score of 0.9848 for the skill-scan engine. This metric reflects the platform's ability to correctly identify and validate AI system capabilities without false positives.

Jailbreak Resistance Testing

Multi-turn jailbreak attacks—including Many-Shot, PAIR, GOAT, and ActorAttack methodologies—are exercised against target models using standardized datasets. The benchmark datasets are defined in api.md (lines 386-400) and include:

  • JailBench-Tiny
  • JailbreakPrompts-Tiny
  • ChatGPT-Jailbreak-Prompts

These datasets allow developers to measure how well different LLMs resist adversarial prompt injection across varying complexity levels.

Harmful Content Detection (HarmfulEvalBenchmark)

The HarmfulEvalBenchmark aggregates diverse harmful prompt categories to measure detection recall and precision. As documented in api.md (lines 398-401), this benchmark tests the platform's ability to identify toxic, biased, or dangerous outputs across multiple content categories, providing standardized metrics for content safety validation.

MCP Scan Throughput

The MCP (Model-Component-Protection) scanner processes code repositories in parallel to detect vulnerabilities in AI infrastructure components. The concurrency limits are explicitly documented in api.md (lines 1347-1350) to prevent performance degradation during large-scale repository scans, ensuring optimal throughput without overwhelming system resources.

End-to-End Latency Visualization

Performance latency is visualized through interactive radar charts in the web interface. The PromptResultDisplay.tsx component (lines 535-568) renders real-time performance data for jailbreak and attack-method evaluations, enabling comparative analysis across different LLMs including qwen3-max and claude-opus-4.1.

Architecture Supporting Benchmarks

The platform's benchmark capabilities rely on a three-tier architecture that separates orchestration, execution, and visualization concerns.

Go Backend (cmd/cli/main.go) orchestrates task queues and WebSocket communication while enforcing concurrency limits to maintain benchmark consistency.

Python Modules (agent-scan, mcp-scan, AIG-PromptSecurity) implement the actual scanning logic and expose benchmark datasets under the data/eval/ directory.

Web UI renders benchmark results and performance charts, allowing operators to compare model behavior across different security dimensions as documented in common/websocket/static/aigdocs/docs/faq_en.md (lines 166-170).

Running Performance Benchmarks

Execute benchmarks using the CLI interface or Python library to validate system performance against standardized datasets.

CLI Benchmark Execution

Run the full skill-scan benchmark against a local endpoint:

ai-infra-guard scan -t http://127.0.0.1:8088 \
    --scan-type skill \
    --benchmark skilltrustbench \
    -o result.json

Execute jailbreak resistance testing with a specific model:

ai-infra-guard scan -t http://127.0.0.1:8088 \
    --scan-type jailbreak \
    --dataset JailBench-Tiny \
    --model qwen3-max \
    -o jailbreak_report.json

Run harmful content detection validation:

ai-infra-guard scan -t http://127.0.0.1:8088 \
    --scan-type harmful \
    --dataset HarmfulEvalBenchmark \
    --model claude-opus-4.1 \
    -o harmful_report.json

Python Library Integration

Programmatically access benchmarks using the agent-scan module:

from agent_scan import AgentScanner

scanner = AgentScanner(
    server="http://127.0.0.1:8088",
    model="qwen3-max",
)

# Run the SkillTrustBench benchmark

skill_report = scanner.run_benchmark(
    benchmark="skilltrustbench",
    scan_type="skill"
)

print("SkillTrustBench score:", skill_report["score"])

Key Source Files for Benchmark Data

The following files define how AI-Infra-Guard measures, stores, and displays performance metrics:

Summary

  • SkillTrustBench accuracy reaches 0.9848 for skill-based security validation, as recorded in the project changelog.
  • Jailbreak resistance is tested against standardized datasets including JailBench-Tiny and ChatGPT-Jailbreak-Prompts.
  • Harmful content detection utilizes the HarmfulEvalBenchmark to measure recall and precision across toxic content categories.
  • MCP scan throughput is governed by concurrency limits defined in api.md to prevent system overload.
  • Performance visualization occurs through the React frontend component PromptResultDisplay.tsx, which renders interactive radar charts for comparative model analysis.

Frequently Asked Questions

What is the highest accuracy score reported for AI-Infra-Guard benchmarks?

The highest reported accuracy is 0.9848 on the SkillTrustBench benchmark, documented in the CHANGELOG.md file. This score represents the precision of the skill-scan engine in correctly identifying security-relevant AI capabilities without generating false positives.

Which datasets are used for jailbreak resistance testing?

AI-Infra-Guard utilizes multiple jailbreak datasets including JailBench-Tiny, JailbreakPrompts-Tiny, and ChatGPT-Jailbreak-Prompts. These datasets are referenced in api.md (lines 386-400) and contain adversarial prompts designed to test multi-turn attack resistance against target LLMs.

How does the platform prevent performance degradation during large scans?

The platform implements concurrency limits in the MCP (Model-Component-Protection) scanner, documented in api.md (lines 1347-1350). These limits control parallel processing of code repositories to maintain consistent throughput without overwhelming system resources or degrading scan accuracy.

Can I compare performance across different AI models using the web interface?

Yes. The web UI includes a radar chart visualization in PromptResultDisplay.tsx (lines 535-568) that enables comparative analysis across different LLMs such as qwen3-max and claude-opus-4.1. The FAQ documentation (lines 166-170) confirms this capability for comparing model behavior across jailbreak and attack-method evaluations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →