What Are the Limitations of VulnClaw? A Technical Deep Dive
VulnClaw is constrained by LLM dependency, mandatory API requirements, hard-coded engine caps on reasoning steps, and reliance on external MCP services, requiring explicit authorization and careful configuration for reliable penetration testing.
VulnClaw is an AI-driven penetration-testing CLI developed by Unclecheng-li that orchestrates LLM reasoning, MCP tool calls, and a modular skill set. While powerful, the framework imposes specific technical, operational, and security boundaries that define what it can and cannot execute. Understanding these limitations is essential for authorized security researchers who need to configure the engine for reliable, compliant assessments.
Legal and Operational Constraints
Scope Authorization Requirements
VulnClaw is designed exclusively for authorized security testing. The tool must not be executed against systems without explicit permission. This limitation is enforced by policy rather than technical controls, as the CLI will accept any target URL but operates under a strict "Scope-Authorized Only" security statement defined in the repository documentation.
LLM and API Dependencies
Model Reliability and Hallucination Risks
Results are fundamentally driven by large language models. While VulnClaw implements an "evidence-level anti-hallucination gate" to validate claims against actual tool output, the quality of findings remains dependent on model capability and prompt engineering. If the LLM suggests a vulnerability flag that does not appear in any collected evidence, the claim is automatically discarded. This protects against false positives but may cause the solver to fail reaching "goal achieved" on weaker models or complex targets.
Mandatory API Key Requirements
The engine cannot function without a compatible LLM API key. You must supply credentials for OpenAI, MiniMax, DeepSeek, or other supported providers via the llm.api_key configuration. Without valid authentication, VulnClaw cannot generate reasoning chains or execute tool calls, rendering the framework inoperable.
Architectural and Engine Limitations
Reasoning Step Caps and Termination Logic
The default solve engine (implemented in vulnclaw/agent/solver.py) terminates only when a goal is reached, the blackboard front-line is exhausted, or a safety budget is hit. In practice, exploration is constrained by configurable caps:
session.solve_max_stepsdefaults to 40 reasoning stepssession.max_roundsdefaults to 15 for the legacy engine
You can adjust these limits via CLI commands:
vulnclaw config set session.max_rounds 30
vulnclaw config set session.solve_max_steps 20
These values are validated against the Pydantic schema in vulnclaw/config/schema.py and stored in vulnclaw/config/settings.py.
Parallel Intent and Tool Round Constraints
The solver imposes strict parallelism limits to prevent runaway exploration:
- Maximum 3 intents per Reason (
session.solve_max_intents = 3) - Maximum 6 tool rounds per Intent (
session.solve_max_tool_rounds = 6)
Exceeding these limits silently aborts further exploration of that branch, forcing the engine to backtrack or terminate. This design prevents resource exhaustion but may truncate deep attack paths.
Resource Budgets for Persistent Sessions
Persistent mode operates under hard caps defined in the configuration table:
session.persistent_rounds_per_cycle = 100session.persistent_max_cycles = 10(set to 0 for unlimited)
Setting persistent_max_cycles=0 removes the cycle limit but risks disk space exhaustion from log accumulation. The default session.max_rounds (15) additionally caps the old fixed-round engine.
External Service and Tool Constraints
MCP Server Dependencies
Advanced features—including browser automation, HTTP replay, and Chrome DevTools integration—require external MCP servers to be running. If these services are unavailable, dependent skills are silently skipped. You can disable MCP-dependent tools to avoid failures in restricted environments:
vulnclaw config set mcp.servers.chrome-devtools.enabled false
vulnclaw config set mcp.servers.burp.enabled false
Configuration handling resides in vulnclaw/config/settings.py where the mcp section is merged into runtime parameters.
Built-in Tool Sandbox Restrictions
Only five built-in tools execute locally without MCP:
python_executenmap_scancrypto_decodebrute_force_loginload_skill_reference
All other actions route through MCP, which may impose network restrictions or privilege limitations. If the MCP layer is unreachable (e.g., in air-gapped environments), only these five tools remain functional.
Knowledge Base and Detection Coverage
The embedded knowledge base contains a finite set of references (approximately 180 documents). If a target requires vulnerability knowledge not present in the pre-loaded corpus, the agent falls back to "KB fallback" logic, potentially reducing detection accuracy for novel CVEs or emerging attack techniques. The KB module is implemented in vulnclaw/kb/store.py.
Environmental and Runtime Constraints
Dependency Requirements
VulnClaw requires Python 3.10+, Node.js for MCP tool execution, and optionally nmap for port scanning functionality. Missing dependencies abort related skills without halting the entire session. Environment checks are performed at startup, and unsupported configurations will skip specific tools.
Non-Deterministic Behavior
Because LLM responses vary between calls, repeated runs against the same target may produce different exploration paths and findings. The non-determinism is mitigated by the evidence gate but not eliminated, making automated regression testing or consistent reproduction of results challenging.
Anti-Loop Detection Sensitivity
The engine detects dead-loops after a configurable number of "stale rounds" (session.stale_rounds_threshold = 5). In highly constrained environments with strict rate limits, this detector may prematurely abort useful probing sequences, mistaking slow response times for infinite loops.
How to Configure and Mitigate These Limits
Inspect current engine constraints using the configuration CLI:
vulnclaw config list
Enable the goal-driven solve engine with explicit step limits for controlled exploration:
vulnclaw config set session.engine solve
vulnclaw config set session.solve_max_steps 20
vulnclaw config set session.solve_max_intents 3
vulnclaw solve http://target.example --goal "obtain flag"
Inspect the blackboard graph to diagnose why exploration was terminated:
vulnclaw repl
> blackboard dump
The vulnclaw/agent/blackboard.py module stores Fact and Intent nodes, enforcing uniqueness to prevent cycles while respecting the stale_rounds_threshold.
Summary
- Authorization Required: VulnClaw operates under strict legal scope limitations for authorized testing only.
- LLM Dependency: Requires valid API keys and suffers from potential hallucinations mitigated by an evidence gate.
- Engine Caps: Hard limits on reasoning steps (40), intents per reason (3), and tool rounds per intent (6) prevent runaway execution.
- MCP Reliance: Advanced features require external MCP servers; only 5 tools run locally without dependencies.
- Resource Limits: Persistent sessions default to 100 rounds per cycle and 10 cycles maximum.
- KB Boundaries: Finite knowledge base (~180 docs) may miss novel vulnerabilities.
- Environment: Requires Python 3.10+, Node.js, and optional system tools like nmap.
Frequently Asked Questions
Can VulnClaw run without an internet connection?
Partially. The framework requires an internet connection to reach LLM APIs (OpenAI, DeepSeek, etc.) and MCP servers. However, if you disable MCP-dependent tools and work with only the five built-in tools (python_execute, nmap_scan, etc.), and if the LLM endpoint is accessible via local network or self-hosted model, limited functionality may persist. The tool cannot generate reasoning or tool calls without API access.
Why does VulnClaw stop exploring after 40 steps?
The default solve_max_steps setting in vulnclaw/config/settings.py caps exploration at 40 reasoning steps to prevent resource exhaustion and runaway costs. You can increase this limit using vulnclaw config set session.solve_max_steps <value>, though higher values increase token consumption and execution time.
How does VulnClaw prevent false positives from LLM hallucinations?
The framework implements an evidence-level anti-hallucination gate that validates every LLM claim against actual tool output. If the model asserts a vulnerability or flag that does not appear in collected evidence, the claim is discarded. This safety mechanism protects against false positives but may cause the solver to miss valid findings if the evidence is ambiguous or the model is imprecise.
What happens if MCP servers are unavailable during a scan?
When MCP servers (Chrome DevTools, Burp, etc.) are unreachable, VulnClaw silently skips skills that depend on those services and continues with available tools. You can proactively disable these dependencies via configuration to avoid runtime errors in restricted network environments, though this limits the framework to its five built-in tools.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →