# How Wigolo's Autonomous Agent Plans, Searches, Fetches, Extracts, and Synthesizes Data

> Discover how Wigolo's autonomous agent plans, searches, fetches, extracts, and synthesizes data using a Python SDK and Rust server daemon for efficient orchestration.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: how-to-guide
- Published: 2026-07-19

---

**Wigolo's autonomous agent orchestrates a five-stage pipeline—planning, searching, fetching, extracting, and synthesizing—through a thin Python SDK that delegates complex orchestration to a Rust-based server-side daemon.**

The Wigolo project (available at KnockOutEZ/wigolo) exposes a single-purpose **agent** tool that transforms high-level prompts into structured research results. Understanding how this autonomous agent coordinates data gathering requires examining both the lightweight client SDK and the server-side orchestration logic.

## The Five-Stage Data Pipeline

The agent workflow follows a logical sequence from initial prompt to final output. Each stage maps to specific SDK methods and REST API endpoints.

### Stage 1: Planning

The **planning** phase parses the user prompt and determines the sequence of actions required. When you invoke `client.agent()`, defined in [`sdks/python/src/wigolo/_client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/sdks/python/src/wigolo/_client.py) at lines 87-99, the method forwards the request to the server's `/v1/agent` endpoint. The server-side planner, implemented in Rust, analyzes the prompt and constructs a step-by-step execution plan.

### Stage 2: Search

During the **search** phase, the agent executes web queries to discover relevant URLs. The CrewAI wrapper function `run_search` in [`packages/wigolo-crewai/wigolo_crewai/_core.py`](https://github.com/KnockOutEZ/wigolo/blob/main/packages/wigolo-crewai/wigolo_crewai/_core.py) (lines 46-73) prepares parameters and calls `client.search()`. Under the hood, the `Client.search` method uses `_call_tool` to POST to `/v1/search`, as implemented in [`sdks/python/src/wigolo/_client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/sdks/python/src/wigolo/_client.py) at lines 49-56.

### Stage 3: Fetch

The **fetch** stage retrieves raw HTML or rendered content from discovered URLs. The wrapper `run_fetch` in [`wigolo_crewai/_core.py`](https://github.com/KnockOutEZ/wigolo/blob/main/wigolo_crewai/_core.py) (lines 77-99) delegates to `client.fetch()`, which submits requests to the `/v1/fetch` endpoint via the same `_call_tool` mechanism.

### Stage 4: Extract

**Extraction** applies structured schemas or CSS selectors to convert web pages into JSON objects. The `run_extract` wrapper in [`wigolo_crewai/_core.py`](https://github.com/KnockOutEZ/wigolo/blob/main/wigolo_crewai/_core.py) (lines 60-78) calls `client.extract()`, mapping to the `/v1/extract` endpoint. This stage transforms unstructured HTML into structured data according to your specified schema.

### Stage 5: Synthesis

Finally, **synthesis** aggregates extracted data into the final output. The server-side agent implementation handles this aggregation, optionally running a language model "research" step over the collected data. When a JSON schema is supplied, the server shapes the output to match it; otherwise, it returns a free-form markdown summary. The client receives this result as the return value of `client.agent()` at lines 87-100 in [`_client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/_client.py).

## Client-Side Architecture

The Python SDK maintains a **stateless, transport-only** design, delegating all complex logic to the server.

### The Tool Manifest

The SDK uses a manifest system defined in [`sdks/python/src/wigolo/_manifest.py`](https://github.com/KnockOutEZ/wigolo/blob/main/sdks/python/src/wigolo/_manifest.py) (lines 17-19) to map tool names to their REST paths and default timeouts. This ensures the client stays synchronized with server capabilities without hardcoding endpoint URLs throughout the codebase.

### The HTTP Client

All high-level tool methods—`search`, `fetch`, `extract`, `research`, `crawl`, and `agent`—delegate to `_call_tool` in [`sdks/python/src/wigolo/_client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/sdks/python/src/wigolo/_client.py) (lines 49-56). This method looks up the tool's manifest entry and performs a single POST request with the supplied arguments, keeping the client implementation thin and maintainable.

## Server-Side Orchestration

While the client handles transport, the server-side daemon manages stateful orchestration:

1. **Plan Generation** — Parses the prompt and decides on sequences of `search → fetch → extract` or `research` calls
2. **Execution** — Runs each step in order, respecting user-provided limits such as `max_pages` and `max_time_ms`
3. **Result Aggregation** — Collects extracted data and runs optional LLM processing to produce the final output

This separation of concerns allows the Python SDK to remain lightweight while the Rust-based server handles the heavy lifting of autonomous decision-making.

## Implementation Example

The following example demonstrates the complete end-to-end flow using the Python SDK, mirroring the official [`examples/sdk-python-agent/gather.py`](https://github.com/KnockOutEZ/wigolo/blob/main/examples/sdk-python-agent/gather.py) script:

```python
import json
from wigolo import local_client

SCHEMA = {
    "type": "object",
    "properties": {
        "latest_version": {"type": "string", "description": "latest SQLite release version"},
        "release_date":  {"type": "string", "description": "date that release shipped"},
    },
}

with local_client() as client:
    # Verify the daemon is running

    print("Daemon status:", client.health().get("status"))

    # Run the autonomous agent

    result = client.agent(
        prompt="What is the latest SQLite release version and release date?",
        schema=SCHEMA,
        urls=["https://sqlite.org/"],          # optional seed URLs

        max_pages=3,
        max_time_ms=90_000,
    )

    # The agent returns a dict with rich metadata

    print("Pages fetched:", result.get("pages_fetched"))
    print("Steps taken:", len(result.get("steps", [])))
    print("Total time (ms):", result.get("total_time_ms"))

    # Structured result (matches SCHEMA) or a free‑form markdown report

    if isinstance(result.get("result"), dict):
        print(json.dumps(result["result"], indent=2))
    else:
        print("\n".join(result["result"].splitlines()[:20]))

```

This script initializes a local client, verifies daemon health, and invokes the agent with a structured schema. The method returns rich metadata including `pages_fetched`, `steps`, and `total_time_ms`, plus either structured JSON matching your schema or a markdown summary.

## Summary

- **Five-stage pipeline**: Wigolo's autonomous agent executes planning, search, fetch, extract, and synthesis operations sequentially.
- **Thin client design**: The Python SDK in [`sdks/python/src/wigolo/_client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/sdks/python/src/wigolo/_client.py) delegates complex logic to a Rust-based server via REST endpoints.
- **Manifest-driven**: Tool configurations live in [`sdks/python/src/wigolo/_manifest.py`](https://github.com/KnockOutEZ/wigolo/blob/main/sdks/python/src/wigolo/_manifest.py), ensuring client-server synchronization.
- **Flexible output**: Supply a JSON schema for structured data extraction, or receive free-form markdown summaries.
- **Resource limits**: Control execution via `max_pages` and `max_time_ms` parameters to prevent runaway operations.

## Frequently Asked Questions

### How does the agent decide which URLs to fetch?

The server-side planner analyzes the initial prompt during the **planning** stage to determine relevant search queries. It executes these queries via the `/v1/search` endpoint, then selects URLs from the results based on relevance scoring. You can optionally provide seed URLs via the `urls` parameter in `client.agent()` to guide the initial discovery phase.

### What is the difference between the client and server components?

The **client** ([`sdks/python/src/wigolo/_client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/sdks/python/src/wigolo/_client.py)) is a stateless HTTP transport layer that marshals arguments and forwards them to REST endpoints. The **server** (implemented in Rust) contains the autonomous logic, maintaining state across the multi-step pipeline, executing searches, managing browser fetching, and running synthesis operations. This separation keeps the SDK lightweight while centralizing complex orchestration in the daemon.

### Can I customize the extraction schema?

Yes. Pass a JSON schema dictionary to the `schema` parameter in `client.agent()` to enforce structured output. During the **extract** stage, the server applies this schema to convert HTML into matching JSON objects. If you omit the schema, the agent returns a free-form markdown summary instead of structured data.

### What happens if the agent exceeds time or page limits?

The agent respects the `max_pages` and `max_time_ms` parameters passed to `client.agent()`. The server-side orchestrator tracks resource consumption across the pipeline; when limits are reached, it gracefully terminates the current operation and returns whatever data has been collected up to that point, along with metadata indicating the partial completion status.