# How VulnClaw's LLM Agent Orchestrates Pentest Skills: Dynamic Skill Dispatch and Execution Control

> Discover how VulnClaw's LLM agent orchestrates pentest skills with its three-stage pipeline for intent analysis, dynamic skill injection, and autonomous execution control for efficient security testing.

- Repository: [Unclecheng/VulnClaw](https://github.com/Unclecheng-li/VulnClaw)
- Tags: architecture
- Published: 2026-06-30

---

**VulnClaw's LLM agent orchestrates pentest skills through a three-stage pipeline that analyzes user intent, dynamically injects skill-specific markdown content into system prompts, and manages autonomous execution loops with reflexion-based stagnation detection.**

VulnClaw is an open-source penetration testing framework that leverages large language models to automate security assessments. According to the Unclecheng-li/VulnClaw repository source code, the system centers on an **LLM agent** ([`vulnclaw/agent/core.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/core.py)) that coordinates specialized pentest capabilities through dynamic skill loading and runtime prompt composition. The orchestration flow bridges natural language requests with structured tool execution across three distinct stages: input analysis, prompt construction, and loop-controlled execution.

## Input Analysis and Skill Dispatch

The orchestration process begins when user input enters the system through `AgentCore.chat` or `AgentCore.auto_pentest`. The agent immediately delegates intent recognition to a specialized dispatch layer that maps free-form requests to structured skill definitions.

### Detecting Targets and Phases

In [`vulnclaw/agent/input_analysis.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/input_analysis.py), the system extracts operational parameters before selecting a skill. The `detect_target` and `detect_phase` functions parse the user input to identify the target host and the current penetration testing phase (reconnaissance, exploitation, post-exploitation, etc.). These constraints feed into the `SkillDispatcher` located in [`vulnclaw/skills/dispatcher.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/skills/dispatcher.py), which scores every intent pattern in `SKILL_INTENT_MAP` against the lower-cased input to find the highest-matching skill name.

### Loading Skill Definitions

Once matched, `SkillDispatcher.dispatch(user_input)` triggers `load_skill_by_name` from [`vulnclaw/skills/loader.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/skills/loader.py). For directory-style skills, the loader returns a dictionary containing the `content` (raw markdown procedural knowledge) and a list of `references` (supporting documentation filenames). This content becomes the foundational knowledge context that guides the LLM's subsequent behavior.

## Dynamic System Prompt Composition

After skill selection, the **LLM agent** constructs a specialized system prompt that combines the skill's procedural knowledge with runtime configuration. This happens in `AgentCore._build_system_prompt`, which aggregates multiple context sources before sending the final prompt to the LLM via `call_llm`.

### Building the Skill Context

The core imports contextual data through `get_active_skill_context` in [`vulnclaw/agent/skill_context.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/skill_context.py). This function returns the selected skill's markdown body plus a printable reference list, ensuring the model receives explicit procedural instructions for the chosen pentest domain. The `build_dynamic_system_prompt` function in [`vulnclaw/agent/system_prompt.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/system_prompt.py) then assembles these pieces into the final prompt string that instructs the LLM on available tools, constraints, and execution patterns.

### Integrating MCP Tools and Recon Dimensions

The system prompt builder also incorporates **MCP tool schemas** when an MCP manager is attached, allowing the agent to invoke external Model Context Protocol services. Additionally, it injects **recon-dimension flags** (such as OSINT "personnel" mode) and optional **KB context** from semantic knowledge bases, creating a comprehensive instruction set that guides the LLM toward specific pentest procedures and tool invocations.

## The Execution Loop and Reflexion

With the system prompt established, the agent enters an execution loop that translates LLM outputs into concrete security testing actions. The `run_auto_pentest` function in [`vulnclaw/agent/loop_controller.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/loop_controller.py) manages this cycle, supporting three operational modes: single-turn chat, auto-pentest, and persistent pentesting.

### Auto-Pentest Loop Control

The auto-pentest loop executes up to `max_rounds` (default 15) iterations. After each LLM response, the controller checks for completion signals via `is_completion_signal`, monitors for stagnation using `detect_attack_path`, and validates progress through `is_meaningful_step`. If the loop detects the agent has stalled on the same technique, the **reflexion** subsystem re-weights subsequent prompts using a limited history of failed paths stored in `_REFLEXION_ATTEMPT_MEMORY`, forcing the agent to explore alternative attack vectors.

### Runtime State Management

During execution, `AgentCore._maybe_auto_save_session` persists session state if auto-save is enabled. The `FindingParser.parse(response_text)` method extracts vulnerability findings from LLM outputs and stores them in the session state. When the LLM returns tool calls using the OpenAI tool-calling schema, the core routes them through dedicated execution helpers: `AgentCore._execute_nmap`, `AgentCore._execute_python`, or `AgentCore._execute_mcp_tool`, bridging the gap between LLM reasoning and actual network scanning or code execution.

## Orchestrator Integration

Outside the core agent, `run_agent_task` in [`vulnclaw/orchestrator.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/orchestrator.py) provides a unified entry point used by both the CLI ([`cli/main.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/cli/main.py)) and the web server ([`web/app.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/web/app.py)). This wrapper loads or creates a `SessionRestoreResult`, executes the supplied runner (typically `agent.auto_pentest`), and generates a task summary via `build_task_session_summary`. This architecture separates the low-level agent mechanics from the delivery mechanism, allowing the same LLM agent orchestration logic to power both command-line and web interfaces.

## Code Examples

**Example 1 – Run an autonomous pentest against a target:**

```python
from vulnclaw.agent.core import AgentCore
from vulnclaw.config.settings import load_config

cfg = load_config()                     # loads VulnClawConfig from defaults / env

agent = AgentCore(cfg)                  # initialise the LLM agent

# "recon" intent will cause the dispatcher to load the `recon` skill

results = await agent.auto_pentest(
    user_input="帮我做一次对 10.10.10.5 的渗透测试，包括端口扫描和子域名枚举",
    target="10.10.10.5",
    max_rounds=12,
)

for r in results:
    print(r.output)                     # LLM-generated step description / findings

```

**Example 2 – Manually fetch a Skill's markdown content:**

```python
from vulnclaw.skills.loader import load_skill_by_name

skill = load_skill_by_name("web-security-advanced")
print(skill["content"])                 # raw markdown that will become part of the system prompt

print(skill["references"])              # optional reference docs

```

**Example 3 – Use the orchestrator helper (CLI entry point):**

```python
from vulnclaw.orchestrator import run_agent_task
from vulnclaw.agent.core import AgentCore
from vulnclaw.config.settings import load_config

cfg = load_config()
agent = AgentCore(cfg)

run_result = await run_agent_task(
    agent=agent,
    command="pentest",
    target="192.168.1.42",
    resume=False,
    runner=lambda a: a.auto_pentest(
        user_input="全流程渗透测试",
        target="192.168.1.42",
    ),
)

print(run_result.summary)               # concise session report

```

## Summary

- **VulnClaw's LLM agent** orchestrates pentest skills by combining intent detection, dynamic skill loading, and reflexion-based loop control in [`vulnclaw/agent/core.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/core.py).
- The **SkillDispatcher** ([`vulnclaw/skills/dispatcher.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/skills/dispatcher.py)) maps natural language to procedural skill definitions stored as markdown files.
- **Dynamic system prompts** are constructed in [`vulnclaw/agent/system_prompt.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/system_prompt.py) by merging skill context, MCP tool schemas, and recon-dimension flags.
- The **auto-pentest loop** ([`vulnclaw/agent/loop_controller.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/loop_controller.py)) executes up to 15 rounds by default, using stagnation detection and reflexion memory to avoid repetitive failed attempts.
- The **orchestrator** ([`vulnclaw/orchestrator.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/orchestrator.py)) provides a unified wrapper for CLI and web interfaces, managing session persistence and summary generation.

## Frequently Asked Questions

### What is the role of SkillDispatcher in VulnClaw?

The `SkillDispatcher` in [`vulnclaw/skills/dispatcher.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/skills/dispatcher.py) acts as the intent-to-skill router. It scores user input against patterns defined in `SKILL_INTENT_MAP` to determine which pentest skill (stored as markdown documentation) should provide the knowledge context for the LLM. This ensures the agent applies domain-specific procedures (web security, reconnaissance, etc.) based on the user's natural language request.

### How does VulnClaw prevent the LLM agent from repeating failed attack paths?

VulnClaw implements a **reflexion** subsystem within the execution loop that tracks failed attempts in `_REFLEXION_ATTEMPT_MEMORY`. When `detect_attack_path` identifies stagnation or `is_meaningful_step` returns false, the system re-weights subsequent prompts to discourage previously attempted techniques. This forces the LLM agent to explore alternative approaches rather than looping on ineffective methods.

### Can VulnClaw integrate with external tools like MCP servers?

Yes. During system prompt construction in `AgentCore._build_system_prompt`, the agent checks for an attached MCP manager and injects MCP tool schemas into the context. When the LLM returns tool calls, `AgentCore._execute_mcp_tool` routes these to the appropriate Model Context Protocol servers, allowing seamless integration with external security tools and data sources beyond the built-in nmap and Python executors.

### What is the difference between single-turn chat and auto-pentest modes?

**Single-turn chat** processes a single user input and returns an immediate response without persisting multi-step state. **Auto-pentest** mode activates the full `run_auto_pentest` loop in [`vulnclaw/agent/loop_controller.py`](https://github.com/Unclecheng-li/VulnClaw/blob/main/vulnclaw/agent/loop_controller.py), which iterates up to `max_rounds` (default 15), maintains reflexion memory, auto-saves session state, and continues execution until completion signals or stagnation is detected. Auto-pentest is designed for autonomous, multi-phase penetration testing workflows.