How to Set Up wigolo with LangChain, CrewAI, or LlamaIndex for RAG Applications

To set up wigolo for RAG applications, install the wigolo-langchain or wigolo-llamaindex package, initialize the WigoloMcpClient to spawn the local MCP server, and use the provided retrievers, tools, or readers to ingest web content into your vector index.

wigolo is a local-first web-search MCP (Model Context Protocol) server that exposes search and fetch capabilities without requiring external API keys. By integrating wigolo with LangChain, CrewAI, or LlamaIndex, you can build Retrieval-Augmented Generation (RAG) pipelines that retrieve live web data while keeping all processing local.

Architecture Overview

The integration relies on four core components that communicate asynchronously:

  • wigolo server: A subprocess launched via npx wigolo that implements the MCP JSON-RPC protocol and performs actual web crawling and rendering.
  • WigoloMcpClient: An async Python wrapper (implemented in packages/wigolo-langchain/wigolo_langchain/client.py and packages/wigolo-llamaindex/wigolo_llamaindex/client.py) that manages the subprocess lifecycle, connection reuse, and clean shutdown.
  • LangChain integration (wigolo-langchain): Provides WigoloSearchRetriever for retrieval pipelines and WigoloSearchTool/WigoloFetchTool for agent frameworks. CrewAI consumes these LangChain tools natively.
  • LlamaIndex integration (wigolo-llamaindex): Supplies WigoloWebReader and WigoloSearchReader in packages/wigolo-llamaindex/wigolo_llamaindex/reader.py, returning Document objects compatible with VectorStoreIndex.

Because the client maintains a persistent connection to the subprocess, repeated calls avoid the overhead of spawning new processes.

Prerequisites and Installation

First, ensure the wigolo server is available on your system. Then install the appropriate Python SDK for your framework.

  1. Install and run the wigolo server:

# Global installation (recommended)

npm install -g wigolo

# Or run on-demand without installation

npx wigolo
  1. Install the Python integration packages:

# For LangChain or CrewAI

pip install wigolo-langchain

# For LlamaIndex

pip install wigolo-llamaindex

LangChain and CrewAI Integration

The wigolo-langchain package exposes retrievers for RAG pipelines and tools for agentic workflows. CrewAI agents inherit LangChain tool compatibility directly.

Configuring the WigoloSearchRetriever

Use WigoloSearchRetriever when you need to retrieve web search results as LangChain Document objects for downstream processing.

from wigolo_langchain import WigoloMcpClient, WigoloSearchRetriever

async def run_retriever():
    async with WigoloMcpClient() as client:
        retriever = WigoloSearchRetriever(
            client=client,
            max_results=5,
            include_domains=["docs.python.org"],
            category="docs",  # Options: "general", "code", "news", "papers"

        )
        docs = await retriever.ainvoke("Python async/await tutorial")
        for doc in docs:
            print(doc.metadata["title"], "→", doc.metadata["url"])

The WigoloMcpClient context manager defined in packages/wigolo-langchain/wigolo_langchain/client.py handles JSON-RPC request marshaling and ensures the npx wigolo subprocess terminates cleanly on exit.

Building CrewAI Agents with Wigolo Tools

For agentic frameworks like CrewAI, instantiate WigoloSearchTool and WigoloFetchTool and pass them to the Agent constructor.

from wigolo_langchain import WigoloMcpClient, WigoloSearchTool, WigoloFetchTool
from crewai import Agent, Task, Crew

async def crew_agent():
    async with WigoloMcpClient() as client:
        tools = [
            WigoloSearchTool(client=client),
            WigoloFetchTool(client=client),
        ]

        researcher = Agent(
            role="Web Researcher",
            goal="Find up-to-date technical documentation",
            backstory="You have access to the wigolo web-search MCP.",
            tools=tools,
            verbose=True,
        )

        task = Task(
            description="Search for the latest Python 3.12 release notes and fetch the page.",
            expected_output="A short summary and the URL",
            agent=researcher
        )

        crew = Crew(agents=[researcher], tasks=[task])
        await crew.kickoff()

LlamaIndex Integration

The wigolo-llamaindex package provides BaseReader implementations that convert web content into LlamaIndex Document objects for indexing.

Fetching Static Content with WigoloWebReader

Use WigoloWebReader to crawl specific URLs and convert them into documents. The render_js parameter controls JavaScript execution.

from wigolo_llamaindex import WigoloMcpClient, WigoloWebReader
from llama_index.core import VectorStoreIndex

async def build_index():
    async with WigoloMcpClient() as client:
        reader = WigoloWebReader(client=client, render_js="auto")
        docs = await reader.aload_data(urls=[
            "https://docs.python.org/3/library/asyncio.html",
            "https://docs.python.org/3/library/typing.html",
        ])
        index = VectorStoreIndex.from_documents(docs)
        return index

The WigoloWebReader class is implemented in packages/wigolo-llamaindex/wigolo_llamaindex/reader.py and accepts standard LlamaIndex reader conventions.

Dynamic Search with WigoloSearchReader

For search-driven document generation, use WigoloSearchReader, which queries the wigolo server and returns results as documents ready for indexing.

from wigolo_llamaindex import WigoloMcpClient, WigoloSearchReader

async def search_to_docs():
    async with WigoloMcpClient() as client:
        reader = WigoloSearchReader(
            client=client,
            max_results=5,
            include_domains=["docs.python.org"],
            category="docs",
        )
        docs = await reader.aload_data(query="asyncio best practices")
        # docs are LlamaIndex Document objects ready for VectorStoreIndex

End-to-End RAG Pipeline

Combine WigoloWebReader with LlamaIndex’s VectorStoreIndex and an LLM to complete the RAG loop.

import os
from wigolo_llamaindex import WigoloMcpClient, WigoloWebReader
from llama_index.core import VectorStoreIndex, ServiceContext
from llama_index.llms.openai import OpenAI

async def rag_pipeline():
    # Fetch web pages via wigolo

    async with WigoloMcpClient() as client:
        web_reader = WigoloWebReader(client=client)
        docs = await web_reader.aload_data(urls=[
            "https://react.dev/learn",
            "https://react.dev/reference/react",
        ])

    # Build vector index

    llm = OpenAI(model="gpt-4o-mini", api_key=os.getenv("OPENAI_API_KEY"))
    service_ctx = ServiceContext.from_defaults(llm=llm)
    index = VectorStoreIndex.from_documents(docs, service_context=service_ctx)

    # Query with retrieved context

    query_engine = index.as_query_engine()
    response = query_engine.query("How do I use React useEffect with async functions?")
    print(response)

Summary

Frequently Asked Questions

Does wigolo require API keys or external services?

No API keys are required for wigolo itself. The server runs locally on your machine using npx wigolo. You only need API keys for the LLM providers (such as OpenAI) if you are integrating with LangChain or LlamaIndex completion models.

How does the MCP client handle multiple concurrent requests?

The WigoloMcpClient opens the npx wigolo subprocess once upon entering the async context manager and reuses the same JSON-RPC connection for all subsequent calls. This minimizes process spawn overhead and supports concurrent async operations until the context manager exits and cleanly shuts down the subprocess.

Can I restrict searches to specific domains or content types?

Yes. Both the LangChain retrievers/tools and LlamaIndex readers accept an include_domains list parameter to whitelist specific sites. They also support a category parameter (values: "general", "code", "news", "papers", "docs") to narrow result types.

Is CrewAI integration separate from LangChain?

No. CrewAI builds directly on top of LangChain’s tool and agent architecture. Therefore, you use the wigolo-langchain package and pass WigoloSearchTool or WigoloFetchTool instances to your CrewAI Agent objects exactly as you would with standard LangChain agents.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →