Typical Use Cases for Pentagi: AI-Driven Penetration Testing Workflows

Pentagi is a self-hosted, AI-driven penetration-testing platform that automates red-team engagements, integrates into CI/CD pipelines for continuous security testing, and supports security research through isolated multi-agent workflows.

Pentagi, developed by vxcontrol as an open-source project, combines a multi-agent AI system with secure Docker sandboxing and persistent vector-memory storage. The platform's modular architecture enables security teams to deploy automated testing workflows across diverse environments while maintaining full control over data and infrastructure. This article explores the typical use cases for Pentagi based on its implementation in the vxcontrol/pentagi repository.

Automated Red-Team Engagements

Pentagi orchestrates end-to-end penetration testing workflows through its Flow → Task → Subtask hierarchy managed in backend/pkg/flow/flow_worker.go. The system automatically decomposes high-level security goals into actionable subtasks executed by specialized agents including the Pentester, Coder, and Installer agents.

The platform ships with a curated set of 20+ security tools—including nmap, metasploit, and sqlmap—that run inside isolated Docker containers. As implemented in backend/pkg/docker/client.go, each tool execution occurs within a dedicated container with optional NET_ADMIN capabilities for network-heavy scans, ensuring that malicious payloads or unstable exploits cannot compromise the host system.

All artefacts, tool outputs, and intermediate findings are stored in a searchable vector store using pgvector, enabling the system to reference prior reconnaissance data during extended engagements.

Continuous Security Testing in CI/CD Pipelines

Pentagi exposes REST and GraphQL APIs with Bearer-token authentication, as defined in backend/pkg/server/router.go and backend/pkg/graph/schema.graphqls. This allows CI/CD pipelines to trigger security flows programmatically and retrieve structured JSON results.

Security teams can fail builds automatically when Pentagi identifies critical vulnerabilities during staging deployments. The following example demonstrates how to integrate Pentagi into a CI pipeline using standard HTTP requests:


# Deploy Pentagi services

docker compose up -d

# Submit a flow via the REST API

curl -X POST https://localhost:8443/api/v1/flows \
  -H "Authorization: Bearer $PENTAGI_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"title":"CI Nightly Scan","target":"https://staging.example.com"}'

# Poll for completion

while true; do
  status=$(curl -s -H "Authorization: Bearer $PENTAGI_API_TOKEN" \
    https://localhost:8443/api/v1/flows/$FLOW_ID/status | jq -r .status)
  [[ $status == "completed" ]] && break
  sleep 10
done

# Retrieve the security report

curl -O -H "Authorization: Bearer $PENTAGI_API_TOKEN" \
  https://localhost:8443/api/v1/flows/$FLOW_ID/report

Vulnerability Assessment and Reporting

The Reporter agent generates comprehensive, templated vulnerability reports with exploitation guidance by consuming all tool outputs and memory artefacts stored in the vector database. This process leverages the data models defined in backend/pkg/database/models.go to query pgvector tables for relevant findings.

Users can retrieve specific exploit guides or previous assessment data through GraphQL queries against the knowledge store:

query SearchGuide {
  searchGuide(query: "SQL injection on MySQL 8") {
    id
    content
    metadata
  }
}

Security Research and Exploit Development

Pentagi supports advanced security research through its Coder and Installer agents, which can write, compile, and test custom exploit code inside controlled Docker environments. As documented in the flow execution architecture, researchers can iterate on proof-of-concept payloads without risking host system compromise.

The isolation mechanism in backend/pkg/docker/client.go ensures that even deliberately malicious code runs only within the ephemeral container boundaries, with network capabilities configurable per-task.

Knowledge-Graph-Powered Threat Intelligence

Through optional Graphiti integration, Pentagi stores semantic relationships in Neo4j, enabling fast "search-in-memory" of prior findings. This capability allows security teams to reuse intelligence from past investigations when approaching new engagements, reducing duplicated effort and improving assessment consistency.

The vector memory system, built on pgvector and defined in backend/pkg/database/models.go, maintains long-term storage of guides, code snippets, and tool outputs that feed into the knowledge graph.

Multi-LLM Provider Flexibility

Pentagi supports 10+ LLM providers including OpenAI, Anthropic, Gemini, Bedrock, Ollama, DeepSeek, GLM, Kimi, and Qwen, as well as aggregators like OpenRouter and DeepInfra. The provider architecture in backend/pkg/providers/providers.go enables rapid registration of new backends through a three-step implementation process.

Teams can configure local models via examples/configs/ollama-llama318b.provider.yml to run Pentagi entirely on-premises with zero external API costs, or select high-capability cloud models for complex reasoning tasks.

Observability and Auditing

The platform includes a comprehensive observability stack featuring OpenTelemetry, Grafana, Loki, Jaeger, and Langfuse for LLM-level tracing. Configuration in observability/otel/config.yml sets up telemetry pipelines that capture every tool invocation, LLM reasoning step, and container action.

Security managers can review complete audit trails for compliance requirements, with all data flowing through the instrumentation defined in the backend observability documentation.

Educational and Training Environments

Pentagi supports one-click Docker Compose deployment and isolated worker nodes for hands-on security training. The interactive installer in cmd/installer/main.go provides a TUI that guides users through provider and worker-node setup, while examples/guides/worker_node.md documents the two-node architecture for high-security environments.

Instructors can spin up sandboxed instances for labs without exposing real production assets, allowing students to practice exploit techniques safely within the Docker-isolated environment.

Summary

  • Pentagi is a self-hosted, AI-driven penetration testing platform combining multi-agent workflows with Docker sandboxing.
  • Automated red-team engagements utilize the Flow → Task → Subtask hierarchy in backend/pkg/flow/flow_worker.go with isolated tool execution.
  • CI/CD integration is supported via REST and GraphQL APIs defined in backend/pkg/server/router.go and backend/pkg/graph/schema.graphqls.
  • Security research benefits from the Coder and Installer agents executing within Docker containers managed by backend/pkg/docker/client.go.
  • Knowledge retention uses pgvector storage in backend/pkg/database/models.go with optional Neo4j integration for threat intelligence.
  • Flexible deployment supports 10+ LLM providers through backend/pkg/providers/providers.go and local models via Ollama.

Frequently Asked Questions

What makes Pentagi suitable for automated red-team operations?

Pentagi automates red-team engagements through a hierarchical multi-agent system that decomposes high-level goals into executable subtasks. The Flow Worker in backend/pkg/flow/flow_worker.go orchestrates specialized agents—such as the Pentester, Coder, and Installer—while the Docker client in backend/pkg/docker/client.go ensures all offensive tools run in isolated containers with configurable network capabilities.

How does Pentagi integrate with existing CI/CD pipelines?

The platform exposes REST and GraphQL APIs with Bearer-token authentication, allowing pipeline scripts to trigger security flows programmatically. As defined in backend/pkg/server/router.go, these endpoints accept JSON payloads specifying targets and profiles, return structured vulnerability data, and support polling for completion status. Teams can configure build-failure rules based on severity thresholds identified during automated scans.

Can Pentagi operate entirely offline with local AI models?

Yes, Pentagi supports 10+ LLM providers including local deployment via Ollama and vLLM, as configured in files like examples/configs/ollama-llama318b.provider.yml. The provider architecture in backend/pkg/providers/providers.go enables registration of custom endpoints, allowing organizations to run the complete platform—including AI reasoning—without transmitting data to external cloud services.

What security controls protect the host system during exploit testing?

Pentagi implements Docker-based isolation where every tool and exploit runs inside ephemeral containers managed by backend/pkg/docker/client.go. These containers operate with restricted capabilities, optionally including NET_ADMIN for network scans, but cannot access the host filesystem or resources outside their namespace. The Installer and Coder agents compile and test custom exploits within these sandboxes, ensuring that malicious or unstable code cannot compromise the Pentagi host or target networks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →