# How VulnClaw Performs Vulnerability Analysis: Inside the LLM-Driven Penetration Testing Workflow

> Discover how VulnClaw performs vulnerability analysis with its autonomous LLM-driven loop. Explore its automated security tool execution and structured report generation.

- Repository: [Unclecheng/VulnClaw](https://github.com/Unclecheng-li/VulnClaw)
- Tags: internals
- Published: 2026-07-03

---

**VulnClaw performs vulnerability analysis through an autonomous LLM-driven loop that parses natural language requests, executes built-in security tools like nmap and Python sandboxes, and iteratively refines findings until generating a structured report.**

VulnClaw is an open-source LLM-driven autonomous penetration-testing agent that transforms plain text instructions into systematic security assessments. Understanding how VulnClaw performs vulnerability analysis reveals a sophisticated architecture combining natural language processing, constrained tool execution, and state-machine-driven workflow management. The system operates through a tight feedback loop that cycles between inference, action, and reflection until the analysis is complete.

## The Seven-Stage Vulnerability Analysis Workflow

VulnClaw implements a deterministic seven-stage pipeline that converts user intent into executed security tests. Each stage is implemented in specific source files within the `vulnclaw/agent/` directory.

### Stage 1: Parsing Natural Language Input

The analysis begins in [`vulnclaw/agent/input_analysis.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/input_analysis.py), where the system extracts structured parameters from free-form text. The module provides four key extraction functions:

- **`detect_target`** – Identifies URLs, IP addresses, or hostnames from user input
- **`detect_phase`** – Determines the penetration testing phase (reconnaissance, vulnerability discovery, exploitation, etc.)
- **`extract_task_constraints`** – Parses allow/deny lists for ports, hosts, and paths
- **`extract_user_vuln_hint`** – Captures explicit vulnerability mentions (e.g., "test for SQL injection on the login parameter")

When a user mentions a specific vulnerability, `extract_user_vuln_hint` generates a directive containing proof-of-concept payload examples, forcing the LLM to craft immediate tests rather than performing broad reconnaissance.

### Stage 2: Dynamic System Prompt Construction

The `AgentCore._build_system_prompt` method in [`vulnclaw/agent/core.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/core.py) (lines 85-115) assembles the LLM context. This dynamic prompt injection combines the detected target, current pentest phase, skill context, MCP tool schemas, and optional knowledge-base data. The prompt engineering ensures the LLM understands its current operational constraints and available toolset before generating a response.

### Stage 3: LLM Inference and Function Calling

The [`vulnclaw/agent/llm_client.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/llm_client.py) module handles communication with OpenAI-compatible endpoints. The LLM returns either plain text analysis, structured findings, or a **function-call request** specifying which security tool to invoke. This decision point determines whether the agent proceeds to tool execution or updates its internal state based on the LLM's reasoning.

### Stage 4: Tool Execution and Constraint Enforcement

When the LLM requests tool execution, the dispatcher in [`vulnclaw/agent/builtin_tools.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/builtin_tools.py) validates and runs the operation. The module implements **constraint enforcement** through `enforce_port_constraints` and `enforce_host_path_constraints`, ensuring scans respect user-defined boundaries before invoking native binaries.

Key tools for vulnerability analysis include:

- **`nmap_scan`** – Executes native nmap scans with configurable timing templates and port ranges (lines 84-119)
- **`python_execute`** – Runs sandboxed Python code for custom HTTP requests, payload generation, or response parsing (lines 133-173)
- **`space_search`, `subdomain_enum`, `js_recon`, `dir_enum`, `unauth_test`** – Passive asset discovery, subdomain brute-forcing, JavaScript endpoint extraction, directory enumeration, and unauthenticated endpoint testing (lines 165-235)

The Python sandbox blocks dangerous patterns via `BLOCKED_PATTERNS` while allowing flexible vulnerability verification scripts.

### Stage 5: Context Update and Phase Detection

After tool execution, the system updates the agent's working memory with outputs and parses findings. The `detect_phase_from_output` function in [`vulnclaw/agent/anti_loop.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/anti_loop.py) (lines 1-30) analyzes LLM responses to determine if the current pentest phase is complete or if the workflow should advance (e.g., from `RECON` to `VULN_DISCOVERY`). This state transition logic prevents premature exploitation before reconnaissance finishes.

### Stage 6: Autonomous Loop Control

The high-level orchestration resides in [`vulnclaw/agent/loop_controller.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/loop_controller.py), invoked by `AgentCore.auto_pentest`. This controller manages the **auto-pentest cycle**, implementing:

- Maximum round count enforcement
- Persistent cycle detection and abortion
- Reflexion memory updates to avoid repeating failed paths

The loop continues until the LLM signals completion, the round limit is reached, or a persistent cycle is detected.

### Stage 7: Report Generation

Finally, [`vulnclaw/web/services/report_service.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/web/services/report_service.py) compiles accumulated findings into markdown or HTML reports. The service aggregates tool outputs, LLM reasoning chains, and generated attack summaries into a structured deliverable.

## Core Architectural Components

VulnClaw's vulnerability analysis capability relies on specialized components working in concert:

**State Management** – The `PentestPhase` enum in [`vulnclaw/agent/context.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/context.py) drives the state machine through distinct phases: `RECON` → `VULN_DISCOVERY` → `EXPLOITATION` → `POST_EXPLOITATION` → `REPORTING`.

**Tool Schema Generation** – The `build_openai_tools` function in [`vulnclaw/agent/builtin_tools.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/builtin_tools.py) dynamically generates OpenAI function-calling schemas for both built-in tools and MCP-provided extensions.

**Reconnaissance Utilities** – [`vulnclaw/agent/recon_tools.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/recon_tools.py) houses passive discovery helpers including subdomain enumeration and JavaScript endpoint harvesting.

**Safety Enforcement** – Constraint validation occurs at the tool layer, ensuring that even LLM-generated commands cannot exceed user-specified scopes for target hosts or network ports.

## Built-In Security Tools for Vulnerability Discovery

The tool suite in [`vulnclaw/agent/builtin_tools.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/builtin_tools.py) provides the concrete capabilities used during analysis:

**Network Scanning** – The `nmap_scan` tool wraps the native nmap binary, supporting timing templates (`-T4`, `-T5`) and specific port lists while validating targets against user constraints.

**Custom Payload Execution** – The `python_execute` tool creates a sandboxed environment for running Python scripts that craft custom HTTP requests, process responses, or generate cryptographic payloads. The sandbox inspects code against `BLOCKED_PATTERNS` before execution.

**Asset Discovery** – Reconnaissance tools include `subdomain_enum` for DNS brute-forcing, `js_recon` for parsing JavaScript files for API endpoints, and `dir_enum` for path brute-forcing using wordlists.

## Practical Implementation Examples

### Running a Full Automated Pentest

Initialize the agent and execute a complete vulnerability analysis workflow:

```python
from vulnclaw.agent.core import AgentCore
from vulnclaw.config.schema import VulnClawConfig

config = VulnClawConfig.load_default()
agent = AgentCore(config)

# Execute autonomous vulnerability analysis

results = await agent.auto_pentest(
    user_input="帮我对 https://demo.example.com 进行渗透测试，重点查找 SQL 注入和 XSS。",
    target="demo.example.com",
    max_rounds=12,
)
for r in results:
    print(r.output)   # LLM replies + tool outputs per round

```

This invokes `AgentCore.auto_pentest` → `run_auto_pentest` in the loop controller, managing up to 12 rounds of analysis.

### Extracting Explicit Vulnerability Hints

Force the system to focus on specific vulnerability types:

```python
from vulnclaw.agent.input_analysis import extract_user_vuln_hint

hint = extract_user_vuln_hint(
    "https://target.com/login?uid=1 参数存在 XSS 漏洞，请直接测试"
)
print(hint)   # Returns directive with PoC payload examples

```

The function returns a multi-line directive incorporating payload examples from `get_payload_examples`, bypassing generic reconnaissance.

### LLM Function Calling for Nmap

When the LLM decides port scanning is necessary, it returns:

```json
{
  "name": "nmap_scan",
  "arguments": {
    "target": "192.168.10.5",
    "scan_type": "top_ports",
    "timing": 4
  }
}

```

The agent's `execute_mcp_tool` dispatcher validates host constraints through `enforce_host_path_constraints` before invoking `execute_nmap`.

### Sandboxed Python Execution for Custom Checks

Probe specific endpoints with custom logic:

```json
{
  "name": "python_execute",
  "arguments": {
    "code": "import requests; r = requests.get('https://demo.example.com/search?q=<>'); print(r.text[:200])",
    "purpose": "test XSS payload"
  }
}

```

The sandbox enforces safety patterns and respects task constraints defined in the initial user input.

### Generating Attack Summaries

Produce narrative reports of the attack chain:

```python
summary = await agent._generate_attack_summary()
print(summary)   # LLM-generated narrative of exploitation steps

```

The `_generate_attack_summary` method in [`vulnclaw/agent/core.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/core.py) (lines 50-56) builds a prompt containing all round outputs and requests a polished description for the final report.

## Summary

- **VulnClaw** performs vulnerability analysis through an autonomous seven-stage pipeline defined in [`vulnclaw/agent/core.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/core.py) and [`vulnclaw/agent/loop_controller.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/loop_controller.py).
- The **input analysis module** ([`input_analysis.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/input_analysis.py)) transforms natural language into structured targets, phases, and constraints using regex and keyword matching.
- **Dynamic prompt construction** injects context, tool schemas, and knowledge-base data before each LLM call.
- **Built-in tools** ([`builtin_tools.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/builtin_tools.py)) provide sandboxed execution of nmap, Python scripts, and reconnaissance utilities with mandatory constraint enforcement.
- **Phase detection** ([`anti_loop.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/anti_loop.py)) and **reflexion memory** prevent infinite loops and guide the state machine through proper pentest phases.
- The system generates **structured reports** via [`report_service.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/report_service.py) after completing the autonomous analysis loop.

## Frequently Asked Questions

### How does VulnClaw handle user-defined constraints during scanning?

VulnClaw enforces constraints at the tool execution layer in [`vulnclaw/agent/builtin_tools.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/builtin_tools.py). The functions `enforce_port_constraints` and `enforce_host_path_constraints` validate all tool arguments against allow/deny lists extracted during the initial input analysis phase. This prevents the LLM from accidentally scanning out-of-scope hosts or restricted network ports even when generating tool calls autonomously.

### What LLM models does VulnClaw support for vulnerability analysis?

According to the source code in [`vulnclaw/agent/llm_client.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/llm_client.py), VulnClaw supports OpenAI-compatible endpoints. The system uses standard OpenAI function-calling schemas generated by `build_openai_tools` in [`builtin_tools.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/builtin_tools.py), allowing integration with GPT-4, GPT-3.5-turbo, or any API-compatible alternative that supports tool use and streaming responses.

### How does VulnClaw prevent infinite loops during autonomous testing?

The framework implements multiple safeguards in [`vulnclaw/agent/loop_controller.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/loop_controller.py) and [`vulnclaw/agent/anti_loop.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/anti_loop.py). The `detect_phase_from_output` function tracks phase transitions, while the loop controller monitors for persistent cycles and enforces a `max_rounds` limit. Additionally, [`vulnclaw/agent/reflexion.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/reflexion.py) maintains memory of failed paths to prevent the agent from retrying identical failed actions repeatedly.

### Can VulnClaw execute custom Python code for specialized vulnerability checks?

Yes, through the `python_execute` tool defined in [`vulnclaw/agent/builtin_tools.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/builtin_tools.py) (lines 133-173). Users (or the LLM) can submit Python scripts for custom HTTP requests, payload generation, or data processing. The execution occurs in a sandboxed environment that validates code against `BLOCKED_PATTERNS` to prevent dangerous operations while respecting the task constraints defined in the initial user input.