How VulnClaw Performs Vulnerability Analysis: Inside the LLM-Driven Penetration Testing Workflow
VulnClaw performs vulnerability analysis through an autonomous LLM-driven loop that parses natural language requests, executes built-in security tools like nmap and Python sandboxes, and iteratively refines findings until generating a structured report.
VulnClaw is an open-source LLM-driven autonomous penetration-testing agent that transforms plain text instructions into systematic security assessments. Understanding how VulnClaw performs vulnerability analysis reveals a sophisticated architecture combining natural language processing, constrained tool execution, and state-machine-driven workflow management. The system operates through a tight feedback loop that cycles between inference, action, and reflection until the analysis is complete.
The Seven-Stage Vulnerability Analysis Workflow
VulnClaw implements a deterministic seven-stage pipeline that converts user intent into executed security tests. Each stage is implemented in specific source files within the vulnclaw/agent/ directory.
Stage 1: Parsing Natural Language Input
The analysis begins in vulnclaw/agent/input_analysis.py, where the system extracts structured parameters from free-form text. The module provides four key extraction functions:
detect_target– Identifies URLs, IP addresses, or hostnames from user inputdetect_phase– Determines the penetration testing phase (reconnaissance, vulnerability discovery, exploitation, etc.)extract_task_constraints– Parses allow/deny lists for ports, hosts, and pathsextract_user_vuln_hint– Captures explicit vulnerability mentions (e.g., "test for SQL injection on the login parameter")
When a user mentions a specific vulnerability, extract_user_vuln_hint generates a directive containing proof-of-concept payload examples, forcing the LLM to craft immediate tests rather than performing broad reconnaissance.
Stage 2: Dynamic System Prompt Construction
The AgentCore._build_system_prompt method in vulnclaw/agent/core.py (lines 85-115) assembles the LLM context. This dynamic prompt injection combines the detected target, current pentest phase, skill context, MCP tool schemas, and optional knowledge-base data. The prompt engineering ensures the LLM understands its current operational constraints and available toolset before generating a response.
Stage 3: LLM Inference and Function Calling
The vulnclaw/agent/llm_client.py module handles communication with OpenAI-compatible endpoints. The LLM returns either plain text analysis, structured findings, or a function-call request specifying which security tool to invoke. This decision point determines whether the agent proceeds to tool execution or updates its internal state based on the LLM's reasoning.
Stage 4: Tool Execution and Constraint Enforcement
When the LLM requests tool execution, the dispatcher in vulnclaw/agent/builtin_tools.py validates and runs the operation. The module implements constraint enforcement through enforce_port_constraints and enforce_host_path_constraints, ensuring scans respect user-defined boundaries before invoking native binaries.
Key tools for vulnerability analysis include:
nmap_scan– Executes native nmap scans with configurable timing templates and port ranges (lines 84-119)python_execute– Runs sandboxed Python code for custom HTTP requests, payload generation, or response parsing (lines 133-173)space_search,subdomain_enum,js_recon,dir_enum,unauth_test– Passive asset discovery, subdomain brute-forcing, JavaScript endpoint extraction, directory enumeration, and unauthenticated endpoint testing (lines 165-235)
The Python sandbox blocks dangerous patterns via BLOCKED_PATTERNS while allowing flexible vulnerability verification scripts.
Stage 5: Context Update and Phase Detection
After tool execution, the system updates the agent's working memory with outputs and parses findings. The detect_phase_from_output function in vulnclaw/agent/anti_loop.py (lines 1-30) analyzes LLM responses to determine if the current pentest phase is complete or if the workflow should advance (e.g., from RECON to VULN_DISCOVERY). This state transition logic prevents premature exploitation before reconnaissance finishes.
Stage 6: Autonomous Loop Control
The high-level orchestration resides in vulnclaw/agent/loop_controller.py, invoked by AgentCore.auto_pentest. This controller manages the auto-pentest cycle, implementing:
- Maximum round count enforcement
- Persistent cycle detection and abortion
- Reflexion memory updates to avoid repeating failed paths
The loop continues until the LLM signals completion, the round limit is reached, or a persistent cycle is detected.
Stage 7: Report Generation
Finally, vulnclaw/web/services/report_service.py compiles accumulated findings into markdown or HTML reports. The service aggregates tool outputs, LLM reasoning chains, and generated attack summaries into a structured deliverable.
Core Architectural Components
VulnClaw's vulnerability analysis capability relies on specialized components working in concert:
State Management – The PentestPhase enum in vulnclaw/agent/context.py drives the state machine through distinct phases: RECON → VULN_DISCOVERY → EXPLOITATION → POST_EXPLOITATION → REPORTING.
Tool Schema Generation – The build_openai_tools function in vulnclaw/agent/builtin_tools.py dynamically generates OpenAI function-calling schemas for both built-in tools and MCP-provided extensions.
Reconnaissance Utilities – vulnclaw/agent/recon_tools.py houses passive discovery helpers including subdomain enumeration and JavaScript endpoint harvesting.
Safety Enforcement – Constraint validation occurs at the tool layer, ensuring that even LLM-generated commands cannot exceed user-specified scopes for target hosts or network ports.
Built-In Security Tools for Vulnerability Discovery
The tool suite in vulnclaw/agent/builtin_tools.py provides the concrete capabilities used during analysis:
Network Scanning – The nmap_scan tool wraps the native nmap binary, supporting timing templates (-T4, -T5) and specific port lists while validating targets against user constraints.
Custom Payload Execution – The python_execute tool creates a sandboxed environment for running Python scripts that craft custom HTTP requests, process responses, or generate cryptographic payloads. The sandbox inspects code against BLOCKED_PATTERNS before execution.
Asset Discovery – Reconnaissance tools include subdomain_enum for DNS brute-forcing, js_recon for parsing JavaScript files for API endpoints, and dir_enum for path brute-forcing using wordlists.
Practical Implementation Examples
Running a Full Automated Pentest
Initialize the agent and execute a complete vulnerability analysis workflow:
from vulnclaw.agent.core import AgentCore
from vulnclaw.config.schema import VulnClawConfig
config = VulnClawConfig.load_default()
agent = AgentCore(config)
# Execute autonomous vulnerability analysis
results = await agent.auto_pentest(
user_input="帮我对 https://demo.example.com 进行渗透测试,重点查找 SQL 注入和 XSS。",
target="demo.example.com",
max_rounds=12,
)
for r in results:
print(r.output) # LLM replies + tool outputs per round
This invokes AgentCore.auto_pentest → run_auto_pentest in the loop controller, managing up to 12 rounds of analysis.
Extracting Explicit Vulnerability Hints
Force the system to focus on specific vulnerability types:
from vulnclaw.agent.input_analysis import extract_user_vuln_hint
hint = extract_user_vuln_hint(
"https://target.com/login?uid=1 参数存在 XSS 漏洞,请直接测试"
)
print(hint) # Returns directive with PoC payload examples
The function returns a multi-line directive incorporating payload examples from get_payload_examples, bypassing generic reconnaissance.
LLM Function Calling for Nmap
When the LLM decides port scanning is necessary, it returns:
{
"name": "nmap_scan",
"arguments": {
"target": "192.168.10.5",
"scan_type": "top_ports",
"timing": 4
}
}
The agent's execute_mcp_tool dispatcher validates host constraints through enforce_host_path_constraints before invoking execute_nmap.
Sandboxed Python Execution for Custom Checks
Probe specific endpoints with custom logic:
{
"name": "python_execute",
"arguments": {
"code": "import requests; r = requests.get('https://demo.example.com/search?q=<>'); print(r.text[:200])",
"purpose": "test XSS payload"
}
}
The sandbox enforces safety patterns and respects task constraints defined in the initial user input.
Generating Attack Summaries
Produce narrative reports of the attack chain:
summary = await agent._generate_attack_summary()
print(summary) # LLM-generated narrative of exploitation steps
The _generate_attack_summary method in vulnclaw/agent/core.py (lines 50-56) builds a prompt containing all round outputs and requests a polished description for the final report.
Summary
- VulnClaw performs vulnerability analysis through an autonomous seven-stage pipeline defined in
vulnclaw/agent/core.pyandvulnclaw/agent/loop_controller.py. - The input analysis module (
input_analysis.py) transforms natural language into structured targets, phases, and constraints using regex and keyword matching. - Dynamic prompt construction injects context, tool schemas, and knowledge-base data before each LLM call.
- Built-in tools (
builtin_tools.py) provide sandboxed execution of nmap, Python scripts, and reconnaissance utilities with mandatory constraint enforcement. - Phase detection (
anti_loop.py) and reflexion memory prevent infinite loops and guide the state machine through proper pentest phases. - The system generates structured reports via
report_service.pyafter completing the autonomous analysis loop.
Frequently Asked Questions
How does VulnClaw handle user-defined constraints during scanning?
VulnClaw enforces constraints at the tool execution layer in vulnclaw/agent/builtin_tools.py. The functions enforce_port_constraints and enforce_host_path_constraints validate all tool arguments against allow/deny lists extracted during the initial input analysis phase. This prevents the LLM from accidentally scanning out-of-scope hosts or restricted network ports even when generating tool calls autonomously.
What LLM models does VulnClaw support for vulnerability analysis?
According to the source code in vulnclaw/agent/llm_client.py, VulnClaw supports OpenAI-compatible endpoints. The system uses standard OpenAI function-calling schemas generated by build_openai_tools in builtin_tools.py, allowing integration with GPT-4, GPT-3.5-turbo, or any API-compatible alternative that supports tool use and streaming responses.
How does VulnClaw prevent infinite loops during autonomous testing?
The framework implements multiple safeguards in vulnclaw/agent/loop_controller.py and vulnclaw/agent/anti_loop.py. The detect_phase_from_output function tracks phase transitions, while the loop controller monitors for persistent cycles and enforces a max_rounds limit. Additionally, vulnclaw/agent/reflexion.py maintains memory of failed paths to prevent the agent from retrying identical failed actions repeatedly.
Can VulnClaw execute custom Python code for specialized vulnerability checks?
Yes, through the python_execute tool defined in vulnclaw/agent/builtin_tools.py (lines 133-173). Users (or the LLM) can submit Python scripts for custom HTTP requests, payload generation, or data processing. The execution occurs in a sandboxed environment that validates code against BLOCKED_PATTERNS to prevent dangerous operations while respecting the task constraints defined in the initial user input.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →