# How to Set Up wigolo with LangChain, CrewAI, or LlamaIndex for RAG Applications

> Learn how to set up wigolo with LangChain, CrewAI, or LlamaIndex for RAG applications. Install wigolo, initialize the client, and ingest web content into your vector index easily.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: how-to-guide
- Published: 2026-07-19

---

**To set up wigolo for RAG applications, install the `wigolo-langchain` or `wigolo-llamaindex` package, initialize the `WigoloMcpClient` to spawn the local MCP server, and use the provided retrievers, tools, or readers to ingest web content into your vector index.**

**wigolo** is a local-first web-search MCP (Model Context Protocol) server that exposes `search` and `fetch` capabilities without requiring external API keys. By integrating wigolo with **LangChain**, **CrewAI**, or **LlamaIndex**, you can build Retrieval-Augmented Generation (RAG) pipelines that retrieve live web data while keeping all processing local.

## Architecture Overview

The integration relies on four core components that communicate asynchronously:

- **wigolo server**: A subprocess launched via `npx wigolo` that implements the MCP JSON-RPC protocol and performs actual web crawling and rendering.
- **`WigoloMcpClient`**: An async Python wrapper (implemented in [`packages/wigolo-langchain/wigolo_langchain/client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/packages/wigolo-langchain/wigolo_langchain/client.py) and [`packages/wigolo-llamaindex/wigolo_llamaindex/client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/packages/wigolo-llamaindex/wigolo_llamaindex/client.py)) that manages the subprocess lifecycle, connection reuse, and clean shutdown.
- **LangChain integration** (`wigolo-langchain`): Provides `WigoloSearchRetriever` for retrieval pipelines and `WigoloSearchTool`/`WigoloFetchTool` for agent frameworks. CrewAI consumes these LangChain tools natively.
- **LlamaIndex integration** (`wigolo-llamaindex`): Supplies `WigoloWebReader` and `WigoloSearchReader` in [`packages/wigolo-llamaindex/wigolo_llamaindex/reader.py`](https://github.com/KnockOutEZ/wigolo/blob/main/packages/wigolo-llamaindex/wigolo_llamaindex/reader.py), returning `Document` objects compatible with `VectorStoreIndex`.

Because the client maintains a persistent connection to the subprocess, repeated calls avoid the overhead of spawning new processes.

## Prerequisites and Installation

First, ensure the wigolo server is available on your system. Then install the appropriate Python SDK for your framework.

1. **Install and run the wigolo server**:

```bash

# Global installation (recommended)

npm install -g wigolo

# Or run on-demand without installation

npx wigolo

```

2. **Install the Python integration packages**:

```bash

# For LangChain or CrewAI

pip install wigolo-langchain

# For LlamaIndex

pip install wigolo-llamaindex

```

## LangChain and CrewAI Integration

The `wigolo-langchain` package exposes retrievers for RAG pipelines and tools for agentic workflows. CrewAI agents inherit LangChain tool compatibility directly.

### Configuring the WigoloSearchRetriever

Use `WigoloSearchRetriever` when you need to retrieve web search results as LangChain `Document` objects for downstream processing.

```python
from wigolo_langchain import WigoloMcpClient, WigoloSearchRetriever

async def run_retriever():
    async with WigoloMcpClient() as client:
        retriever = WigoloSearchRetriever(
            client=client,
            max_results=5,
            include_domains=["docs.python.org"],
            category="docs",  # Options: "general", "code", "news", "papers"

        )
        docs = await retriever.ainvoke("Python async/await tutorial")
        for doc in docs:
            print(doc.metadata["title"], "→", doc.metadata["url"])

```

The `WigoloMcpClient` context manager defined in [`packages/wigolo-langchain/wigolo_langchain/client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/packages/wigolo-langchain/wigolo_langchain/client.py) handles JSON-RPC request marshaling and ensures the `npx wigolo` subprocess terminates cleanly on exit.

### Building CrewAI Agents with Wigolo Tools

For agentic frameworks like CrewAI, instantiate `WigoloSearchTool` and `WigoloFetchTool` and pass them to the `Agent` constructor.

```python
from wigolo_langchain import WigoloMcpClient, WigoloSearchTool, WigoloFetchTool
from crewai import Agent, Task, Crew

async def crew_agent():
    async with WigoloMcpClient() as client:
        tools = [
            WigoloSearchTool(client=client),
            WigoloFetchTool(client=client),
        ]

        researcher = Agent(
            role="Web Researcher",
            goal="Find up-to-date technical documentation",
            backstory="You have access to the wigolo web-search MCP.",
            tools=tools,
            verbose=True,
        )

        task = Task(
            description="Search for the latest Python 3.12 release notes and fetch the page.",
            expected_output="A short summary and the URL",
            agent=researcher
        )

        crew = Crew(agents=[researcher], tasks=[task])
        await crew.kickoff()

```

## LlamaIndex Integration

The `wigolo-llamaindex` package provides `BaseReader` implementations that convert web content into LlamaIndex `Document` objects for indexing.

### Fetching Static Content with WigoloWebReader

Use `WigoloWebReader` to crawl specific URLs and convert them into documents. The `render_js` parameter controls JavaScript execution.

```python
from wigolo_llamaindex import WigoloMcpClient, WigoloWebReader
from llama_index.core import VectorStoreIndex

async def build_index():
    async with WigoloMcpClient() as client:
        reader = WigoloWebReader(client=client, render_js="auto")
        docs = await reader.aload_data(urls=[
            "https://docs.python.org/3/library/asyncio.html",
            "https://docs.python.org/3/library/typing.html",
        ])
        index = VectorStoreIndex.from_documents(docs)
        return index

```

The `WigoloWebReader` class is implemented in [`packages/wigolo-llamaindex/wigolo_llamaindex/reader.py`](https://github.com/KnockOutEZ/wigolo/blob/main/packages/wigolo-llamaindex/wigolo_llamaindex/reader.py) and accepts standard LlamaIndex reader conventions.

### Dynamic Search with WigoloSearchReader

For search-driven document generation, use `WigoloSearchReader`, which queries the wigolo server and returns results as documents ready for indexing.

```python
from wigolo_llamaindex import WigoloMcpClient, WigoloSearchReader

async def search_to_docs():
    async with WigoloMcpClient() as client:
        reader = WigoloSearchReader(
            client=client,
            max_results=5,
            include_domains=["docs.python.org"],
            category="docs",
        )
        docs = await reader.aload_data(query="asyncio best practices")
        # docs are LlamaIndex Document objects ready for VectorStoreIndex

```

### End-to-End RAG Pipeline

Combine `WigoloWebReader` with LlamaIndex’s `VectorStoreIndex` and an LLM to complete the RAG loop.

```python
import os
from wigolo_llamaindex import WigoloMcpClient, WigoloWebReader
from llama_index.core import VectorStoreIndex, ServiceContext
from llama_index.llms.openai import OpenAI

async def rag_pipeline():
    # Fetch web pages via wigolo

    async with WigoloMcpClient() as client:
        web_reader = WigoloWebReader(client=client)
        docs = await web_reader.aload_data(urls=[
            "https://react.dev/learn",
            "https://react.dev/reference/react",
        ])

    # Build vector index

    llm = OpenAI(model="gpt-4o-mini", api_key=os.getenv("OPENAI_API_KEY"))
    service_ctx = ServiceContext.from_defaults(llm=llm)
    index = VectorStoreIndex.from_documents(docs, service_context=service_ctx)

    # Query with retrieved context

    query_engine = index.as_query_engine()
    response = query_engine.query("How do I use React useEffect with async functions?")
    print(response)

```

## Summary

- **wigolo** runs entirely locally via `npx wigolo`, eliminating the need for external search API keys while maintaining privacy.
- The **`WigoloMcpClient`** class (found in both [`packages/wigolo-langchain/wigolo_langchain/client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/packages/wigolo-langchain/wigolo_langchain/client.py) and [`packages/wigolo-llamaindex/wigolo_llamaindex/client.py`](https://github.com/KnockOutEZ/wigolo/blob/main/packages/wigolo-llamaindex/wigolo_llamaindex/client.py)) manages the async subprocess connection and protocol handling.
- **LangChain** users should utilize `WigoloSearchRetriever` for retrieval chains or `WigoloSearchTool`/`WigoloFetchTool` for agents consumed by CrewAI.
- **LlamaIndex** users should import `WigoloWebReader` and `WigoloSearchReader` from [`packages/wigolo-llamaindex/wigolo_llamaindex/reader.py`](https://github.com/KnockOutEZ/wigolo/blob/main/packages/wigolo-llamaindex/wigolo_llamaindex/reader.py) to generate `Document` objects for `VectorStoreIndex`.
- All integrations support filtering via `include_domains`, `category`, and `max_results` parameters.

## Frequently Asked Questions

### Does wigolo require API keys or external services?

No API keys are required for wigolo itself. The server runs locally on your machine using `npx wigolo`. You only need API keys for the LLM providers (such as OpenAI) if you are integrating with LangChain or LlamaIndex completion models.

### How does the MCP client handle multiple concurrent requests?

The `WigoloMcpClient` opens the `npx wigolo` subprocess once upon entering the async context manager and reuses the same JSON-RPC connection for all subsequent calls. This minimizes process spawn overhead and supports concurrent async operations until the context manager exits and cleanly shuts down the subprocess.

### Can I restrict searches to specific domains or content types?

Yes. Both the LangChain retrievers/tools and LlamaIndex readers accept an `include_domains` list parameter to whitelist specific sites. They also support a `category` parameter (values: `"general"`, `"code"`, `"news"`, `"papers"`, `"docs"`) to narrow result types.

### Is CrewAI integration separate from LangChain?

No. CrewAI builds directly on top of LangChain’s tool and agent architecture. Therefore, you use the `wigolo-langchain` package and pass `WigoloSearchTool` or `WigoloFetchTool` instances to your CrewAI `Agent` objects exactly as you would with standard LangChain agents.