# What is VulnClaw? An AI-Powered Penetration Testing Framework Explained

> Discover VulnClaw, the AI-powered open-source pen testing framework. Uncover vulnerabilities from recon to exploit with LLM orchestration and evidence validation. Learn more!

- Repository: [Unclecheng/VulnClaw](https://github.com/Unclecheng-li/VulnClaw)
- Tags: getting-started
- Published: 2026-07-03

---

**VulnClaw is an AI-driven, open-source penetration-testing framework that uses large language models to orchestrate the entire security testing lifecycle—from reconnaissance to exploitation—while preventing hallucinations through an evidence-level validation system.**

VulnClaw is an open-source project hosted at `Unclecheng-li/VulnClaw` that redefines automated security testing by combining natural language processing with traditional penetration testing tools. This framework allows security professionals to describe testing objectives in plain English while an LLM agent coordinates specialized tools, validates findings against real tool output, and generates structured reports with proof-of-concept code.

## Architecture Overview

VulnClaw's architecture follows a modular design that separates user interfaces, agent intelligence, tool execution, and state management. The system processes natural language commands through multiple entry points and translates them into concrete security testing actions via tightly integrated layers.

### Entry Points and Interfaces

Users can interact with VulnClaw through several interfaces implemented in distinct modules. The command-line interface provides scriptable automation, while [`vulnclaw/cli/tui_textual.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/cli/tui_textual.py) implements a rich terminal UI workbench for interactive sessions. For browser-based access, [`vulnclaw/web/app.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/web/app.py) launches a web interface defaulting to `127.0.0.1:7788`.

### The LLM Agent Core

The central intelligence resides in [`vulnclaw/agent/core.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/core.py), which implements the **AgentCore** class responsible for orchestrating the testing lifecycle. This module interprets natural language intents and coordinates with [`vulnclaw/agent/prompts.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/prompts.py) to build system prompts that inject available skills and MCP tools into the LLM context.

### MCP Toolchain Integration

VulnClaw implements the **Model Context Protocol (MCP)** to interface with external security tools. The [`vulnclaw/mcp/registry.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/mcp/registry.py) module registers services such as `fetch`, `memory`, `chrome-devtools`, and `burp`, allowing the LLM to execute real HTTP requests, browser automation, and integration with existing security scanners.

### The Blackboard State Graph

To prevent the agent from repeating actions or hallucinating progress, [`vulnclaw/agent/blackboard.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/blackboard.py) maintains a central state graph distinguishing between **Facts** (verified observations) and **Intents** (pending exploration directions). This graph structure ensures the agent tracks what has been proven versus what remains to be tested.

### The Solve Engine and OODA Loop

The default goal-driven engine implemented in [`vulnclaw/agent/solver.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/solver.py) follows an **OODA loop** (Observe, Orient, Decide, Act) methodology. The engine operates through three phases: **Reason**, **Explore**, and **Conclude**. It terminates when the target goal is reached, the exploration frontier is exhausted, or a safety budget is depleted.

### Reasoning and Reflexion

Advanced cognitive capabilities reside in [`vulnclaw/agent/reasoning_state.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/reasoning_state.py) and [`vulnclaw/agent/reflexion.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/reflexion.py). The reasoning state stores facts, constraints, and attack-chain candidates, while the reflexion module auto-classifies failures and upgrades payloads across severity levels (L0-L4) through iterative cycles.

### Skills and Knowledge Base

VulnClaw includes over 21 built-in skills accessible through [`vulnclaw/skills/loader.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/skills/loader.py). These specialized capabilities range from header analysis to JWT inspection. The knowledge base in [`vulnclaw/kb/store.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/kb/store.py) provides searchable references that the LLM can invoke on demand without loading everything into context.

### Plugin Runtime

The `vulnclaw/plugins/` directory contains low-coupling plugins that execute vulnerability detection logic safely and feed results back into the session state. These plugins run with minimal coupling to the core system, allowing for extensible detection capabilities.

### Report Generation

After completing the testing cycle, [`vulnclaw/report/generator.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/report/generator.py) processes the accumulated session state to produce structured Markdown reports. The companion module [`vulnclaw/report/poc_builder.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/report/poc_builder.py) generates runnable Python proof-of-concept scripts based on successful exploitation paths.

## Anti-Hallucination Mechanisms

A critical differentiator in VulnClaw is its **evidence-level anti-hallucination gate**. Unlike generic LLM agents that might invent vulnerabilities, VulnClaw only accepts flags and findings that are proven in real tool output. This validation occurs at the boundary between the LLM agent and the MCP toolchain, ensuring that reported vulnerabilities correspond to actual observed behavior rather than model-generated speculation.

## Installation and Usage

Install VulnClaw from PyPI and execute testing workflows through the CLI or Python API.

Install the framework:

```bash
pip install vulnclaw

```

Run a quick scan against a target:

```bash
vulnclaw run http://target.example.com

```

Use goal-driven mode to capture specific flags:

```bash
vulnclaw solve http://ctf.example.com --goal "obtain the flag"

```

Execute specific reconnaissance phases:

```bash
vulnclaw recon http://target.example.com

```

Run built-in plugins on existing data:

```bash
vulnclaw plugins run builtin.web.headers --input headers.json --session session.json

```

Launch the web interface:

```bash
vulnclaw web

```

Or open the terminal UI:

```bash
vulnclaw tui

```

For programmatic access, instantiate the AgentCore directly:

```python
from vulnclaw.agent.core import AgentCore
from vulnclaw.config.settings import Settings

settings = Settings.load()
agent = AgentCore(settings)
result = agent.run_target("http://target.example.com")
print(result.report_path)

```

## Summary

- **VulnClaw** is an AI-driven penetration-testing framework that automates the security testing lifecycle through natural language interaction.
- The architecture separates concerns across user interfaces ([`vulnclaw/cli/tui_textual.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/cli/tui_textual.py), [`vulnclaw/web/app.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/web/app.py)), agent intelligence ([`vulnclaw/agent/core.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/core.py)), and tool integration ([`vulnclaw/mcp/registry.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/mcp/registry.py)).
- An **evidence-level anti-hallucination gate** ensures only validated findings from real tool output are reported, preventing false positives from LLM speculation.
- The **Blackboard** state graph in [`vulnclaw/agent/blackboard.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/blackboard.py) tracks Facts and Intents to prevent redundant actions and maintain testing state.
- The **OODA loop** implementation in [`vulnclaw/agent/solver.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/solver.py) provides a goal-driven engine that reasons, explores, and concludes based on defined objectives.
- Over **21 built-in skills** and a plugin architecture allow extensible vulnerability detection without modifying core framework code.

## Frequently Asked Questions

### What does the name VulnClaw represent?

VulnClaw combines "Vuln" (short for vulnerabilities) with "Claw," suggesting the framework's ability to dig deep into applications and extract security weaknesses. The name reflects its purpose as a tool that aggressively hunts for vulnerabilities while maintaining precision through its evidence-based validation system.

### How does VulnClaw prevent LLM hallucinations during testing?

VulnClaw implements an **evidence-level anti-hallucination gate** that strictly validates findings against real tool output from the MCP toolchain. The system only accepts flags that are proven through actual HTTP responses, browser automation results, or scanner output. Additionally, the **Blackboard** state graph in [`vulnclaw/agent/blackboard.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/blackboard.py) maintains verified Facts separate from speculative Intents, ensuring the agent operates on observed reality rather than generated assumptions.

### What is the MCP protocol and why does VulnClaw use it?

MCP stands for **Model Context Protocol**, a standardized interface that allows VulnClaw to connect with external security tools and services. Implemented in [`vulnclaw/mcp/registry.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/mcp/registry.py), this protocol enables the LLM agent to invoke real tools like `fetch` for HTTP requests, `chrome-devtools` for browser automation, and `burp` for integration with existing security scanners. Using MCP ensures that the agent's actions translate to actual system interactions rather than simulated responses.

### Can VulnClaw integrate with existing security tools like Burp Suite?

Yes, VulnClaw supports integration with existing security infrastructure through its MCP toolchain. The [`vulnclaw/mcp/registry.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/mcp/registry.py) module specifically includes a `burp` service for connecting with Burp Suite, allowing the AI agent to leverage existing proxy configurations, scan results, and manual testing artifacts within its automated workflow. This integration enables security teams to augment their current tooling with AI-driven orchestration without replacing established workflows.