How Wigolo's Autonomous Agent Plans, Searches, Fetches, Extracts, and Synthesizes Data

Wigolo's autonomous agent orchestrates a five-stage pipeline—planning, searching, fetching, extracting, and synthesizing—through a thin Python SDK that delegates complex orchestration to a Rust-based server-side daemon.

The Wigolo project (available at KnockOutEZ/wigolo) exposes a single-purpose agent tool that transforms high-level prompts into structured research results. Understanding how this autonomous agent coordinates data gathering requires examining both the lightweight client SDK and the server-side orchestration logic.

The Five-Stage Data Pipeline

The agent workflow follows a logical sequence from initial prompt to final output. Each stage maps to specific SDK methods and REST API endpoints.

Stage 1: Planning

The planning phase parses the user prompt and determines the sequence of actions required. When you invoke client.agent(), defined in sdks/python/src/wigolo/_client.py at lines 87-99, the method forwards the request to the server's /v1/agent endpoint. The server-side planner, implemented in Rust, analyzes the prompt and constructs a step-by-step execution plan.

During the search phase, the agent executes web queries to discover relevant URLs. The CrewAI wrapper function run_search in packages/wigolo-crewai/wigolo_crewai/_core.py (lines 46-73) prepares parameters and calls client.search(). Under the hood, the Client.search method uses _call_tool to POST to /v1/search, as implemented in sdks/python/src/wigolo/_client.py at lines 49-56.

Stage 3: Fetch

The fetch stage retrieves raw HTML or rendered content from discovered URLs. The wrapper run_fetch in wigolo_crewai/_core.py (lines 77-99) delegates to client.fetch(), which submits requests to the /v1/fetch endpoint via the same _call_tool mechanism.

Stage 4: Extract

Extraction applies structured schemas or CSS selectors to convert web pages into JSON objects. The run_extract wrapper in wigolo_crewai/_core.py (lines 60-78) calls client.extract(), mapping to the /v1/extract endpoint. This stage transforms unstructured HTML into structured data according to your specified schema.

Stage 5: Synthesis

Finally, synthesis aggregates extracted data into the final output. The server-side agent implementation handles this aggregation, optionally running a language model "research" step over the collected data. When a JSON schema is supplied, the server shapes the output to match it; otherwise, it returns a free-form markdown summary. The client receives this result as the return value of client.agent() at lines 87-100 in _client.py.

Client-Side Architecture

The Python SDK maintains a stateless, transport-only design, delegating all complex logic to the server.

The Tool Manifest

The SDK uses a manifest system defined in sdks/python/src/wigolo/_manifest.py (lines 17-19) to map tool names to their REST paths and default timeouts. This ensures the client stays synchronized with server capabilities without hardcoding endpoint URLs throughout the codebase.

The HTTP Client

All high-level tool methods—search, fetch, extract, research, crawl, and agent—delegate to _call_tool in sdks/python/src/wigolo/_client.py (lines 49-56). This method looks up the tool's manifest entry and performs a single POST request with the supplied arguments, keeping the client implementation thin and maintainable.

Server-Side Orchestration

While the client handles transport, the server-side daemon manages stateful orchestration:

  1. Plan Generation — Parses the prompt and decides on sequences of search → fetch → extract or research calls
  2. Execution — Runs each step in order, respecting user-provided limits such as max_pages and max_time_ms
  3. Result Aggregation — Collects extracted data and runs optional LLM processing to produce the final output

This separation of concerns allows the Python SDK to remain lightweight while the Rust-based server handles the heavy lifting of autonomous decision-making.

Implementation Example

The following example demonstrates the complete end-to-end flow using the Python SDK, mirroring the official examples/sdk-python-agent/gather.py script:

import json
from wigolo import local_client

SCHEMA = {
    "type": "object",
    "properties": {
        "latest_version": {"type": "string", "description": "latest SQLite release version"},
        "release_date":  {"type": "string", "description": "date that release shipped"},
    },
}

with local_client() as client:
    # Verify the daemon is running

    print("Daemon status:", client.health().get("status"))

    # Run the autonomous agent

    result = client.agent(
        prompt="What is the latest SQLite release version and release date?",
        schema=SCHEMA,
        urls=["https://sqlite.org/"],          # optional seed URLs

        max_pages=3,
        max_time_ms=90_000,
    )

    # The agent returns a dict with rich metadata

    print("Pages fetched:", result.get("pages_fetched"))
    print("Steps taken:", len(result.get("steps", [])))
    print("Total time (ms):", result.get("total_time_ms"))

    # Structured result (matches SCHEMA) or a free‑form markdown report

    if isinstance(result.get("result"), dict):
        print(json.dumps(result["result"], indent=2))
    else:
        print("\n".join(result["result"].splitlines()[:20]))

This script initializes a local client, verifies daemon health, and invokes the agent with a structured schema. The method returns rich metadata including pages_fetched, steps, and total_time_ms, plus either structured JSON matching your schema or a markdown summary.

Summary

  • Five-stage pipeline: Wigolo's autonomous agent executes planning, search, fetch, extract, and synthesis operations sequentially.
  • Thin client design: The Python SDK in sdks/python/src/wigolo/_client.py delegates complex logic to a Rust-based server via REST endpoints.
  • Manifest-driven: Tool configurations live in sdks/python/src/wigolo/_manifest.py, ensuring client-server synchronization.
  • Flexible output: Supply a JSON schema for structured data extraction, or receive free-form markdown summaries.
  • Resource limits: Control execution via max_pages and max_time_ms parameters to prevent runaway operations.

Frequently Asked Questions

How does the agent decide which URLs to fetch?

The server-side planner analyzes the initial prompt during the planning stage to determine relevant search queries. It executes these queries via the /v1/search endpoint, then selects URLs from the results based on relevance scoring. You can optionally provide seed URLs via the urls parameter in client.agent() to guide the initial discovery phase.

What is the difference between the client and server components?

The client (sdks/python/src/wigolo/_client.py) is a stateless HTTP transport layer that marshals arguments and forwards them to REST endpoints. The server (implemented in Rust) contains the autonomous logic, maintaining state across the multi-step pipeline, executing searches, managing browser fetching, and running synthesis operations. This separation keeps the SDK lightweight while centralizing complex orchestration in the daemon.

Can I customize the extraction schema?

Yes. Pass a JSON schema dictionary to the schema parameter in client.agent() to enforce structured output. During the extract stage, the server applies this schema to convert HTML into matching JSON objects. If you omit the schema, the agent returns a free-form markdown summary instead of structured data.

What happens if the agent exceeds time or page limits?

The agent respects the max_pages and max_time_ms parameters passed to client.agent(). The server-side orchestrator tracks resource consumption across the pipeline; when limits are reached, it gracefully terminates the current operation and returns whatever data has been collected up to that point, along with metadata indicating the partial completion status.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →