How the Claude Plugins Security Scanner Detects Credential Exfiltration
The Claude Plugins security scanner detects credential exfiltration by using an LLM-based policy prompt that analyzes code flows to identify when credentials from system stores (Keychain, environment variables, AWS files, etc.) are routed to external services, blocking any plugin that attempts this cross-service data flow.
The anthropics/claude-plugins-community repository implements a comprehensive security scanning system that prevents malicious plugins from stealing user credentials. At the heart of this system lies a specialized scanner that runs during the scan-plugins CI workflow, examining every file in a plugin submission—including hidden directories and test scripts—to identify patterns where sensitive credentials might be exfiltrated to attacker-controlled endpoints.
The Three-Stage Detection Pipeline
The scanner operates through a coordinated three-stage process defined across multiple files in the repository.
Stage 1: Policy Prompt Definition
The detection logic begins in .github/actions/scan-plugins/policy/prompt.md, which contains the textual rubric that instructs the LLM scanner what constitutes a violation.
This policy explicitly instructs the model to flag credential/secret exfiltration by detecting code that performs two simultaneous actions:
- Reads a credential store such as macOS Keychain (
security find-generic-password), Linuxsecret-tool, Windowscmdkey,~/.aws/credentials,~/.claude/.credentials, or environment variables (process.env.*) - Routes the value to a service other than its origin, creating a cross-service data flow
The prompt also defines negative examples—clarifying that normal same-service usage does not constitute a violation—reducing false positives during scanning.
Stage 2: LLM-Based Policy Enforcement
The scan-plugins action invokes the LLM with the policy prompt and feeds it the complete plugin payload. Located in .github/actions/scan-plugins/scripts/scan.sh, this script:
- Collects all files under the plugin directory, including non-surface files
- Calls the LLM with the exfiltration detection rules
- Parses the model's analysis to identify any credential-source calls and subsequent code paths
- Inspects whether retrieved secrets are sent to a different endpoint than their origin
When the model identifies a credential read followed by transmission to an external service, it reports a "FAILS policy" violation.
Stage 3: CI Feedback and Enforcement
Violations surface in the GitHub Actions workflow defined in .github/workflows/validate-plugins.yml. When the scanner detects exfiltration, the workflow generates GitHub-compatible error annotations using the format:
::error scan-plugins: my-plugin FAILS policy — credential exfiltration detected in scripts/upload.sh
This output halts the CI pipeline and prevents the plugin from being merged or listed in the marketplace until developers remove the exfiltration code.
What Constitutes a Credential Exfiltration Violation
The scanner specifically targets the cross-service credential flow pattern. Valid credential access that remains within the originating service—such as reading AWS credentials to make AWS API calls—passes inspection. However, code that extracts credentials from the macOS Keychain and subsequently transmits them to a remote logging server triggers an immediate failure.
The policy covers multiple credential storage mechanisms:
- System keychains: macOS Keychain, Windows Credential Manager (
cmdkey), Linux Secret Service (secret-tool) - Configuration files:
~/.aws/credentials,~/.claude/.credentials, and other dotfiles - Environment variables: Any access to
process.envor shell environment variables containing secrets
Running the Scanner Locally
Developers can execute the security scanner locally to validate plugins before submitting pull requests. The scanner requires no additional parameters beyond the plugin directory structure.
Execute the scanning script directly:
bash .github/actions/scan-plugins/scripts/scan.sh
The script performs the following operations:
- Recursively collects all files in the plugin directory
- Sends file contents to the LLM with the exfiltration policy prompt
- Outputs GitHub-compatible error annotations for any violations found
For CI integration, the workflow automatically invokes the scanner:
# .github/workflows/validate-plugins.yml (excerpt)
- name: Scan plugins
uses: ./.github/actions/scan-plugins
with:
# Automatically reads the plugin directory and runs the policy prompt
# No additional parameters required
Integration with Runtime Security Layers
According to .claude-plugin/marketplace.json (line 809), the exfiltration detection scanner forms part of the seven security layers that guard every Claude Code session. The marketplace description lists "exfiltration detection" alongside path validation, command validation, injection scanning, and credential redaction.
This architecture ensures that detection happens at two levels: statically during the CI/CD pipeline via the scan-plugins workflow, and dynamically during runtime execution through the Claude Code security guardrails.
Summary
- The scanner uses an LLM-based analysis defined in
.github/actions/scan-plugins/policy/prompt.mdto interpret code semantics and identify credential exfiltration patterns. - It specifically flags code that reads from credential stores (Keychain, environment variables, AWS files) and transmits data to external services.
- Violations generate
::errorannotations in.github/workflows/validate-plugins.yml, blocking plugin publication. - The system examines all files including hidden directories and test scripts, leaving no code path unanalyzed.
- Exfiltration detection operates as one of seven security layers described in the marketplace configuration.
Frequently Asked Questions
How does the scanner distinguish between legitimate credential use and exfiltration?
The LLM policy prompt explicitly defines that credentials must remain within their originating service context. Reading ~/.aws/credentials to make an AWS API call constitutes legitimate use, while extracting those same credentials and sending them to a third-party logging endpoint triggers a violation. The scanner analyzes data-flow patterns to determine whether the credential value crosses service boundaries.
Can the scanner detect obfuscated exfiltration attempts in minified or encoded files?
Yes. The scanner in .github/actions/scan-plugins/scripts/scan.sh processes every file in the plugin directory regardless of extension or visibility, including hidden directories and encoded content. Because the LLM analyzes the semantic behavior of the code rather than relying on simple string matching, it can identify exfiltration patterns even in obfuscated or minified JavaScript, base64-encoded strings, or indirectly referenced credential accessors.
What happens if a plugin fails the credential exfiltration check?
The CI workflow defined in .github/workflows/validate-plugins.yml treats any "FAILS policy" output as a blocking error. The workflow prints a GitHub annotation in the format scan-plugins: <plugin> FAILS policy — <details> and terminates with a non-zero exit code. The plugin cannot be merged into the repository or listed in the Claude marketplace until developers remove the exfiltration code and the scanner returns a passing status.
Does the scanner only check environment variables, or does it analyze system keychain access too?
The scanner analyzes all common credential storage mechanisms. The policy prompt explicitly lists macOS Keychain (security find-generic-password), Linux secret-tool, Windows cmdkey, AWS credential files, Claude-specific credential files, and environment variables. This comprehensive coverage ensures that plugins cannot bypass detection by choosing alternative credential stores.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →