Risk Categories for Agent Skill Security Audits in Tencent AI-Infra-Guard

The Agent Skill scanner in Tencent/AI-Infra-Guard evaluates skills against six specific risk categories: Malicious behavior, Permission abuse, Privacy access, High-risk operations, Hard-coded secrets, and LLM jailbreak attempts.

Agent skills extend the capabilities of AI systems, but they also introduce security boundaries that must be validated. The Tencent/AI-Infra-Guard project provides a comprehensive static analysis framework that audits these skills according to a strict taxonomy defined in skills/edgeone-skill-scanner/SKILL.md. Understanding these risk categories for Agent Skill security audits is essential for developers and security teams who need to evaluate third-party skills before deployment.

The Six Risk Categories for Agent Skill Audits

The scanner examines every skill for evidence falling into six distinct risk categories. Each category targets specific attack vectors or security misconfigurations that could compromise the host system or user data.

Malicious Behavior

This category detects code attempting to perform active attacks. The scanner searches for reverse shells, backdoors, cryptominers, and trojan-like payload delivery mechanisms.

When a skill contains functions that open network connections to external command-and-control servers or deploy mining scripts, the scanner flags this as malicious behavior. Typical findings include literal detections of "reverse shell", "cryptomining", or "backdoor" patterns in the codebase.

Permission Abuse

Permission abuse occurs when a skill requests or utilizes more privileges than its declared purpose requires. The scanner checks for unauthorized file system access, permission changes, and admin-level actions without explicit consent.

Concrete findings include attempts at "broad deletion", "disk wipe/format" operations, or "dangerous permission changes" such as modifying system files like /etc/hosts without user approval.

Privacy Access

This category identifies access to privacy-sensitive data including photos, documents, chat logs, authentication tokens, passwords, and cryptographic key files. The skills/edgeone-skill-scanner/SKILL.md specification explicitly flags any skill that attempts to read private user data outside its intended scope.

Typical detections include "access to private files" and "credential exfiltration" patterns where skills transmit sensitive data to external endpoints.

High-Risk Operations

High-risk operations encompass actions that can disrupt the host environment or runtime stability. The scanner monitors for host-disruptive commands and system-level changes that could crash the system or corrupt data.

Findings in this category include "host-disruptive actions" and "tool tampering" where a skill modifies its own execution environment or other installed tools.

Hard-Coded Secrets

The scanner detects real credentials, API tokens, cryptographic keys, or passwords embedded directly in code or configuration files. Unlike other categories that examine behavior, this category focuses on static secrets exposure.

When the scanner identifies strings matching token patterns or finds credentials in YAML/JSON config files, it reports "hard-coded credentials" findings with specific line references.

LLM Jailbreak and Prompt Override

This advanced category detects encoded payloads and obfuscation tricks designed to bypass model safety guards. The scanner decodes base64 strings, Unicode smuggling attempts, and ROT13/hex-encoded directives that attempt to override system prompts.

Findings include "base64-encoded overrides" and "Unicode smuggling" where malicious instructions are hidden in seemingly benign text to trigger unauthorized behavior in the underlying language model.

Severity Classification and Audit Verdicts

When the scanner identifies issues within these risk categories, it assigns one of two severity grades according to the specification in skills/edgeone-skill-scanner/SKILL.md lines 188-194.

Risk Severity Levels

  • ⚠️ Needs Attention: Indicates medium-or-higher risk evidence that does not constitute immediate high risk. These findings require manual review but may not block deployment.
  • 🔴 High Risk: Reserved for clear malicious intent, confirmed credential exfiltration, destructive host actions, or attempts to bypass sandbox/approval mechanisms.

Capability vs. Abuse Distinction

The audit process distinguishes between capability (what a skill can do) and abuse (what it actually does harmfully). A skill might possess the capability to delete files, but the scanner only elevates this to a high-risk finding when the code actively abuses that capability without proper safeguards or user consent.

Practical Example of a Security Audit Report

The following JSON structure illustrates how the scanner formats findings when multiple risk categories are triggered:

{
  "skill": "example_skill",
  "risk_findings": [
    {
      "category": "Permission abuse",
      "description": "Skill writes to /etc/hosts without user consent",
      "severity": "high",
      "recommendation": "Remove host-file modification or request explicit user approval"
    },
    {
      "category": "Hard-coded secrets",
      "description": "API token `abc123XYZ` is hard-coded in config.yaml",
      "severity": "medium",
      "recommendation": "Store the token in an environment variable or secret manager"
    }
  ],
  "overall_verdict": "🔴 发现风险"
}

For user-facing reports, the scanner generates markdown summaries:


## 🔴 example_skill 发现安全风险

**不建议直接安装或继续使用。**

这个 skill 存在以下问题:它会在你不知情的情况下修改系统文件,并且硬编码了敏感的 API token。

**建议**:
1. 先停用这个 skill  
2. 联系 skill 的开发者确认是否为正常行为  
3. 在确认安全前不要重新启用

These reports directly reference the risk categories and provide actionable remediation steps based on the severity classification defined in the source code.

Summary

  • The Agent Skill scanner in skills/edgeone-skill-scanner/SKILL.md defines six risk categories: Malicious behavior, Permission abuse, Privacy access, High-risk operations, Hard-coded secrets, and LLM jailbreak/Prompt override.
  • Findings are classified as either "⚠️ Needs Attention" (medium risk) or "🔴 High Risk" (clear malicious intent or destructive capability).
  • The audit framework distinguishes between a skill's capability and actual abuse of that capability.
  • Hard-coded secrets and privacy access violations target static code analysis, while malicious behavior and high-risk operations focus on dynamic execution patterns.
  • Each finding includes specific remediation recommendations tailored to the risk category and severity level.

Frequently Asked Questions

How does the scanner differentiate between normal functionality and permission abuse?

The scanner analyzes whether the skill uses privileges beyond its declared purpose. According to the Tencent/AI-Infra-Guard source code, permission abuse requires evidence that the skill performs admin-level actions—such as modifying system files or changing permissions—without explicit user consent or outside the scope documented in the skill's manifest. Normal functionality stays within declared boundaries, while abuse involves "broad deletion" or "dangerous permission changes."

What constitutes a High Risk vs. Needs Attention severity rating?

High Risk findings indicate clear malicious intent, confirmed credential exfiltration, destructive host actions, or sandbox bypass attempts. Needs Attention covers medium-or-higher risks that lack immediate destructive capability, such as theoretical vulnerabilities or excessive permissions that aren't actively exploited in the current code path. The distinction is defined in lines 188-194 of the skill scanner specification.

Can the scanner detect obfuscated malicious code in skills?

Yes, the LLM jailbreak / Prompt override category specifically targets obfuscation techniques. The scanner identifies base64-encoded overrides, Unicode smuggling, and ROT13/hex-encoded directives that attempt to hide malicious payloads. This category addresses evasion techniques designed to bypass both static analysis and model safety filters.

Where are the risk category definitions maintained in the repository?

The canonical definitions reside in skills/edgeone-skill-scanner/SKILL.md at lines 165-166, which establishes the six-category taxonomy. Additional example implementations demonstrating these categories in practice can be found in agent-scan/agent_scan/prompt/skills/web-exfiltration-detection/SKILL.md (privacy access) and agent-scan/agent_scan/prompt/skills/unexpected-code-execution-detection/SKILL.md (malicious behavior).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →