What Is VulnClaw? An AI-Powered Penetration Testing Framework

VulnClaw is an open-source, AI-driven penetration testing framework that translates natural language instructions into automated security workflows, orchestrating the entire testing lifecycle from reconnaissance to exploitation and report generation through a large language model (LLM) agent.

Developed by the team at Unclecheng-li/VulnClaw, this tool integrates a modular architecture with a Model-Context-Protocol (MCP) toolchain to execute real-world security actions. Unlike traditional scanners, VulnClaw employs a goal-driven OODA loop (Observe-Orient-Decide-Act) that plans, validates, and documents each step while preventing redundant operations through a centralized state graph.

Core Architecture and Components

VulnClaw’s architecture consists of tightly integrated layers that handle everything from user input to final report delivery. The system is designed to prevent LLM hallucinations through an evidence-level anti-hallucination gate that only accepts flags proven in real tool output.

User Interface Layer

VulnClaw provides multiple entry points to accommodate different workflows:

Each interface connects to the underlying agent core, allowing seamless switching between interactive and automated modes.

LLM Agent and MCP Toolchain

The vulnclaw/agent/core.py module serves as the central orchestrator, interpreting natural language intent and selecting appropriate skills. The agent communicates with external tools via the Model-Context-Protocol (MCP), implemented in vulnclaw/mcp/registry.py.

Available MCP services include:

  • fetch: HTTP request handling and web scraping
  • memory: Persistent storage for session data
  • chrome-devtools: Browser automation and dynamic analysis
  • burp: Integration with Burp Suite for proxy-based testing

Blackboard State Management

The vulnclaw/agent/blackboard.py implements a Fact/Intent graph that serves as the system's central memory. Facts represent verified observations (e.g., "Port 80 is open"), while Intents represent pending exploration directions (e.g., "Check for SQL injection on login form"). This graph prevents the agent from repeating actions and ensures comprehensive coverage.

Solve Engine and OODA Loop

The vulnclaw/agent/solver.py implements the default goal-driven engine following the OODA loop pattern:

  1. Reason: Analyze current state and constraints
  2. Explore: Select and execute the next best action
  3. Conclude: Validate findings against the goal

The engine terminates when the target goal is reached, the frontier is exhausted, or a safety budget limit is triggered.

Reasoning and Reflexion System

Advanced reasoning capabilities are handled by vulnclaw/agent/reasoning_state.py and vulnclaw/agent/reflexion.py. The system stores attack-chain candidates and constraints, while the reflexion module auto-classifies failures and upgrades payload sophistication from L0 to L4 across execution cycles.

Key Features and Capabilities

Evidence-Based Validation

VulnClaw guards against false positives through a strict validation mechanism. The LLM cannot claim vulnerability discovery without concrete evidence from tool outputs, ensuring that reports contain only verified findings.

Skill and Knowledge Base System

The framework ships with over 21 built-in skills covering core and specialized penetration testing techniques. The skill loader (vulnclaw/skills/loader.py) dynamically invokes capabilities, while the knowledge base (vulnclaw/kb/store.py) provides searchable references for the LLM to consult during execution.

Plugin Runtime

The vulnclaw/plugins/ directory contains low-coupling plugins for specific checks (e.g., header analysis, JWT inspection). These plugins run safely in isolation and feed results back into the session state without affecting the core agent loop.

Report Generation

The vulnclaw/report/generator.py and vulnclaw/report/poc_builder.py modules transform accumulated session data into structured Markdown reports. Each report includes a runnable Python Proof-of-Concept (PoC) script that can reproduce the discovered vulnerabilities.

Getting Started with VulnClaw

Installation

Install the latest release from PyPI:

pip install vulnclaw

Command-Line Usage

Execute a quick scan using the default goal-driven engine:

vulnclaw run http://target.example.com

Run with an explicit objective for CTF-style challenges:

vulnclaw solve http://ctf.example.com --goal "obtain the flag"

Perform reconnaissance only:

vulnclaw recon http://target.example.com

Execute a specific plugin on saved HTTP headers:

vulnclaw plugins run builtin.web.headers --input headers.json --session session.json

Launch the Web UI (defaults to 127.0.0.1:7788):

vulnclaw web

Open the terminal UI workbench:

vulnclaw tui

Python API Integration

For advanced automation, use the Python API directly:

from vulnclaw.agent.core import AgentCore
from vulnclaw.config.settings import Settings

# Load configuration from ~/.vulnclaw/config.yaml

settings = Settings.load()
agent = AgentCore(settings)

# Execute full scan workflow

result = agent.run_target("http://target.example.com")
print(result.report_path)  # Path to generated Markdown report

Summary

  • VulnClaw is an AI-driven penetration testing framework that uses natural language to orchestrate security assessments via an LLM agent.
  • The architecture centers on a Fact/Intent blackboard (vulnclaw/agent/blackboard.py) that prevents redundant actions and an OODA solve engine (vulnclaw/agent/solver.py) that drives goal-oriented testing.
  • MCP toolchain integration (vulnclaw/mcp/registry.py) provides real tool execution capabilities including HTTP requests, browser automation, and proxy integration.
  • Anti-hallucination safeguards ensure only verified findings are reported, with automatic generation of Markdown reports and Python PoC scripts.
  • The system supports 21+ built-in skills, extensible plugins, and multiple interfaces (CLI, TUI, Web).

Frequently Asked Questions

What makes VulnClaw different from traditional vulnerability scanners?

Unlike signature-based scanners, VulnClaw uses an LLM agent to plan and adapt testing strategies in real-time. The OODA loop in vulnclaw/agent/solver.py allows the tool to make context-aware decisions, while the blackboard architecture ensures it learns from each action rather than following a static script. This enables handling of complex, multi-step vulnerabilities that require logical reasoning.

How does VulnClaw prevent AI hallucinations in security findings?

The framework implements an evidence-level anti-hallucination gate that requires all vulnerability claims to be backed by concrete tool output from the MCP toolchain. The blackboard system in vulnclaw/agent/blackboard.py separates Facts (verified observations) from Intents (hypotheses), ensuring reports contain only validated security issues with supporting evidence.

Can VulnClaw integrate with existing security tools like Burp Suite?

Yes, through the MCP registry (vulnclaw/mcp/registry.py), VulnClaw includes a burp service that enables integration with Burp Suite. Additionally, the chrome-devtools service provides browser automation, while the fetch service handles HTTP interactions. These integrations allow VulnClaw to leverage existing enterprise tools while adding AI-driven orchestration.

What types of penetration testing tasks can VulnClaw automate?

VulnClaw can automate the entire testing lifecycle including information gathering (vulnclaw recon), vulnerability exploitation (via the solve engine), and report generation. With over 21 built-in skills and a plugin system (vulnclaw/plugins/), it handles web application testing, API security assessments, header analysis, JWT inspection, and custom attack chains defined in natural language goals.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →