VulnClaw's LLM Agent Core Architecture: A Modular State-Driven Orchestration Layer

VulnClaw's LLM agent core is a modular, state-driven orchestration system implemented in vulnclaw/agent/core.py that coordinates the LLM, tool ecosystem, and session knowledge base through three tightly coupled components: AgentCore, Context/SessionState, and RuntimeState.

The VulnClaw LLM agent core serves as the central nervous system of the Unclecheng-li/VulnClaw repository, handling everything from single-turn conversations to autonomous multi-cycle penetration testing. This architecture separates persistent session data from ephemeral runtime state while maintaining a pluggable interface for tools and knowledge bases.

Core Architectural Components

The architecture revolves around three primary abstractions that manage different aspects of the agent's lifecycle.

AgentCore: The Central Orchestrator

AgentCore acts as the top-level entry point for all user interactions. Located in vulnclaw/agent/core.py, this class initializes the entire subsystem and coordinates the request lifecycle.

During initialization (AgentCore.__init__), the constructor wires together the context manager, creates a fresh runtime container, and prepares optional semantic search backends:

self.context = ContextManager()
self.runtime = RuntimeState()
self._reset_runtime_state()
self._kb_retriever = KnowledgeRetriever()   # optional KB back-end

This initialization sequence appears at lines 60-71 of core.py. The class also handles LLM client resolution at lines 31-63, managing API keys, OAuth flows, or local ChatGPT proxy auto-start.

ContextManager and SessionState: Persistent Session Data

The ContextManager class in vulnclaw/agent/context.py maintains the conversational history and owns a SessionState object (line 99). SessionState serves as the persistent brain of the operation, storing:

  • Target specifications and current phase
  • Discovered findings and task constraints
  • Recon-dimension flags and blackboard data
  • Reflexion snapshots for cross-cycle learning

The extensive definition spans lines 78-165 and 180-300 in context.py, implementing a durable knowledge layer that survives across multiple autonomous loops.

RuntimeState: Ephemeral Execution Context

RuntimeState in vulnclaw/agent/runtime_state.py tracks mutable per-run data used by autonomous loops. This dataclass (lines 50-80) maintains:

  • Current attack path and reflexion engine instances
  • CTF mode flags and user-provided hints
  • Counters for stalls, errors, and execution metrics

Unlike SessionState, this component resets between high-level execution phases, ensuring clean state for each distinct operation.

Request Processing Pipeline

The AgentCore.chat() method implements the primary request handling workflow at lines 66-112 of core.py. This pipeline executes six distinct phases:

  1. Target Detection: _detect_target and _detect_phase analyze input using input_analysis utilities
  2. Context Storage: self.context.add_user_message(user_input) persists the query
  3. Dynamic Prompt Construction: _build_system_prompt merges MCP tool schemas, skill contexts, knowledge-base snippets, and phase-aware templates (lines 85-95)
  4. LLM Invocation: await call_llm(self, system_prompt, stream_sink) in vulnclaw/agent/llm_client.py
  5. Finding Extraction: self._finding_parser.parse(response_text) structures discovered vulnerabilities
  6. Auto-Persistence: _maybe_auto_save_session() writes JSON snapshots when enabled

Autonomous Execution Loops

Three high-level controllers in vulnclaw/agent/loop_controller.py (lines 24-44) drive multi-turn execution, each invoked through AgentCore methods:

Auto Pentest Loop

AgentCore.auto_pentest triggers run_auto_pentest for fixed-round, deterministic testing against a single target. This loop runs a predetermined number of rounds (default 15) before terminating, making it suitable for bounded vulnerability scans.

Persistent Pentest Loop

AgentCore.persistent_pentest launches run_persistent_pentest, which repeats auto-pentest cycles until encountering stop conditions or user interrupts. This mode supports automatic report generation across multiple testing cycles.

Solver Loop

AgentCore.solve invokes run_solve, implementing a goal-oriented "blackboard-OODA" loop without preset round limits. This controller uses the blackboard pattern defined in vulnclaw/agent/blackboard.py to coordinate complex, multi-stage exploitation chains.

Safety and Anti-Loop Mechanisms

The architecture incorporates sophisticated safeguards against execution stagnation. The vulnclaw/agent/anti_loop.py module (lines 7-12) implements:

  • Attack path detection via detect_attack_path
  • Meaningfulness validation through is_meaningful_step
  • Failure tracking with track_failed_target

All tool executions pass through tool_call_manager.safe_parse_tool_args for security-checked JSON parsing, while builtin tools (nmap, Python executor, IP checks) reside in vulnclaw/agent/builtin_tools.py.

Optional Knowledge Base Integration

When the vulnclaw[kb] extra is installed, AgentCore lazily instantiates KnowledgeRetriever and injects semantic search results via build_kb_context. This integration modifies the system prompt assembly at lines 73-101 of core.py, enriching LLM context with CVE data and historical findings without altering core orchestration logic.

Practical Implementation Examples

Running Single-Turn Reconnaissance

import asyncio
from vulnclaw.agent.core import AgentCore
from vulnclaw.config.settings import load_config

async def main():
    cfg = load_config()                 # reads config/VulnClawConfig

    core = AgentCore(cfg)

    result = await core.chat(
        user_input="对 10.10.10.5 进行信息收集",
        target="10.10.10.5"
    )
    print(result.output)               # LLM response

    print(core.session_state.findings) # any discovered findings

asyncio.run(main())

This example utilizes the chat method (lines 66-80 of core.py) for immediate feedback without autonomous loop execution.

Executing Bounded Autonomous Testing

import asyncio
from vulnclaw.agent.core import AgentCore
from vulnclaw.config.settings import load_config

async def auto():
    cfg = load_config()
    core = AgentCore(cfg)

    # Run up to 10 rounds (default is 15)

    history = await core.auto_pentest(
        user_input="对 192.168.0.10 进行渗透测试",
        target="192.168.0.10",
        max_rounds=10,
    )
    for step in history:
        print(step.output)

asyncio.run(auto())

The auto_pentest method forwards to run_auto_pentest in loop_controller.py (lines 28-36 of core.py), executing a finite testing sequence.

Persistent Multi-Cycle Engagement

import asyncio
from vulnclaw.agent.core import AgentCore
from vulnclaw.config.settings import load_config

async def persistent():
    cfg = load_config()
    core = AgentCore(cfg)

    cycles = await core.persistent_pentest(
        user_input="持续对 example.com 渗透",
        target="example.com",
        rounds_per_cycle=20,
        max_cycles=5,
        auto_report=True,
    )
    print(f"Generated {len(cycles)} cycle reports")
    for c in cycles:
        print(c.report_path)

asyncio.run(persistent())

This triggers run_persistent_pentest (lines 80-90 of core.py), enabling long-running engagements with automatic report generation between cycles.

Summary

  • AgentCore in vulnclaw/agent/core.py serves as the central orchestrator, managing LLM clients, prompt construction, and workflow delegation
  • SessionState maintains persistent pentest context including targets, findings, and reconnaissance status across multiple execution cycles
  • RuntimeState tracks ephemeral per-run data such as attack paths, reflexion engines, and error counters
  • Three distinct loop controllers (auto_pentest, persistent_pentest, solve) provide flexibility for bounded, continuous, or goal-oriented testing scenarios
  • Anti-loop mechanisms in vulnclaw/agent/anti_loop.py prevent stagnation through attack-path detection and meaningfulness validation
  • Optional knowledge-base integration enriches prompts with semantic search results without modifying core orchestration logic

Frequently Asked Questions

What is the primary role of AgentCore in VulnClaw?

AgentCore functions as the top-level orchestration class that receives user requests, builds dynamic system prompts, calls the LLM through vulnclaw/agent/llm_client.py, parses tool calls, and updates session state. It serves as the single entry point for both interactive chat and autonomous penetration testing modes.

How does VulnClaw prevent infinite loops during autonomous testing?

The system implements anti-loop logic in vulnclaw/agent/anti_loop.py that detects repeated failures through detect_attack_path and is_meaningful_step functions. The RuntimeState tracks stall counters and failed targets, forcing path switches when the agent encounters stagnation or redundant operations.

What distinguishes SessionState from RuntimeState?

SessionState maintains persistent data structures including targets, findings, and reflexion snapshots that survive across multiple testing cycles, while RuntimeState holds ephemeral per-run information such as current attack paths, CTF flags, and execution counters that reset between distinct operation phases.

How does the knowledge base integration enhance the LLM agent core?

When enabled, the optional KnowledgeRetriever injects semantic search results into system prompts via build_kb_context in vulnclaw/agent/kb_context.py. This allows the agent to reference CVE data and historical findings during tool selection and exploitation without requiring modifications to the core orchestration logic in AgentCore.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →