How to Set Up wigolo with LangChain, CrewAI, or LlamaIndex for RAG Applications
To set up wigolo for RAG applications, install the wigolo-langchain or wigolo-llamaindex package, initialize the WigoloMcpClient to spawn the local MCP server, and use the provided retrievers, tools, or readers to ingest web content into your vector index.
wigolo is a local-first web-search MCP (Model Context Protocol) server that exposes search and fetch capabilities without requiring external API keys. By integrating wigolo with LangChain, CrewAI, or LlamaIndex, you can build Retrieval-Augmented Generation (RAG) pipelines that retrieve live web data while keeping all processing local.
Architecture Overview
The integration relies on four core components that communicate asynchronously:
- wigolo server: A subprocess launched via
npx wigolothat implements the MCP JSON-RPC protocol and performs actual web crawling and rendering. WigoloMcpClient: An async Python wrapper (implemented inpackages/wigolo-langchain/wigolo_langchain/client.pyandpackages/wigolo-llamaindex/wigolo_llamaindex/client.py) that manages the subprocess lifecycle, connection reuse, and clean shutdown.- LangChain integration (
wigolo-langchain): ProvidesWigoloSearchRetrieverfor retrieval pipelines andWigoloSearchTool/WigoloFetchToolfor agent frameworks. CrewAI consumes these LangChain tools natively. - LlamaIndex integration (
wigolo-llamaindex): SuppliesWigoloWebReaderandWigoloSearchReaderinpackages/wigolo-llamaindex/wigolo_llamaindex/reader.py, returningDocumentobjects compatible withVectorStoreIndex.
Because the client maintains a persistent connection to the subprocess, repeated calls avoid the overhead of spawning new processes.
Prerequisites and Installation
First, ensure the wigolo server is available on your system. Then install the appropriate Python SDK for your framework.
- Install and run the wigolo server:
# Global installation (recommended)
npm install -g wigolo
# Or run on-demand without installation
npx wigolo
- Install the Python integration packages:
# For LangChain or CrewAI
pip install wigolo-langchain
# For LlamaIndex
pip install wigolo-llamaindex
LangChain and CrewAI Integration
The wigolo-langchain package exposes retrievers for RAG pipelines and tools for agentic workflows. CrewAI agents inherit LangChain tool compatibility directly.
Configuring the WigoloSearchRetriever
Use WigoloSearchRetriever when you need to retrieve web search results as LangChain Document objects for downstream processing.
from wigolo_langchain import WigoloMcpClient, WigoloSearchRetriever
async def run_retriever():
async with WigoloMcpClient() as client:
retriever = WigoloSearchRetriever(
client=client,
max_results=5,
include_domains=["docs.python.org"],
category="docs", # Options: "general", "code", "news", "papers"
)
docs = await retriever.ainvoke("Python async/await tutorial")
for doc in docs:
print(doc.metadata["title"], "→", doc.metadata["url"])
The WigoloMcpClient context manager defined in packages/wigolo-langchain/wigolo_langchain/client.py handles JSON-RPC request marshaling and ensures the npx wigolo subprocess terminates cleanly on exit.
Building CrewAI Agents with Wigolo Tools
For agentic frameworks like CrewAI, instantiate WigoloSearchTool and WigoloFetchTool and pass them to the Agent constructor.
from wigolo_langchain import WigoloMcpClient, WigoloSearchTool, WigoloFetchTool
from crewai import Agent, Task, Crew
async def crew_agent():
async with WigoloMcpClient() as client:
tools = [
WigoloSearchTool(client=client),
WigoloFetchTool(client=client),
]
researcher = Agent(
role="Web Researcher",
goal="Find up-to-date technical documentation",
backstory="You have access to the wigolo web-search MCP.",
tools=tools,
verbose=True,
)
task = Task(
description="Search for the latest Python 3.12 release notes and fetch the page.",
expected_output="A short summary and the URL",
agent=researcher
)
crew = Crew(agents=[researcher], tasks=[task])
await crew.kickoff()
LlamaIndex Integration
The wigolo-llamaindex package provides BaseReader implementations that convert web content into LlamaIndex Document objects for indexing.
Fetching Static Content with WigoloWebReader
Use WigoloWebReader to crawl specific URLs and convert them into documents. The render_js parameter controls JavaScript execution.
from wigolo_llamaindex import WigoloMcpClient, WigoloWebReader
from llama_index.core import VectorStoreIndex
async def build_index():
async with WigoloMcpClient() as client:
reader = WigoloWebReader(client=client, render_js="auto")
docs = await reader.aload_data(urls=[
"https://docs.python.org/3/library/asyncio.html",
"https://docs.python.org/3/library/typing.html",
])
index = VectorStoreIndex.from_documents(docs)
return index
The WigoloWebReader class is implemented in packages/wigolo-llamaindex/wigolo_llamaindex/reader.py and accepts standard LlamaIndex reader conventions.
Dynamic Search with WigoloSearchReader
For search-driven document generation, use WigoloSearchReader, which queries the wigolo server and returns results as documents ready for indexing.
from wigolo_llamaindex import WigoloMcpClient, WigoloSearchReader
async def search_to_docs():
async with WigoloMcpClient() as client:
reader = WigoloSearchReader(
client=client,
max_results=5,
include_domains=["docs.python.org"],
category="docs",
)
docs = await reader.aload_data(query="asyncio best practices")
# docs are LlamaIndex Document objects ready for VectorStoreIndex
End-to-End RAG Pipeline
Combine WigoloWebReader with LlamaIndex’s VectorStoreIndex and an LLM to complete the RAG loop.
import os
from wigolo_llamaindex import WigoloMcpClient, WigoloWebReader
from llama_index.core import VectorStoreIndex, ServiceContext
from llama_index.llms.openai import OpenAI
async def rag_pipeline():
# Fetch web pages via wigolo
async with WigoloMcpClient() as client:
web_reader = WigoloWebReader(client=client)
docs = await web_reader.aload_data(urls=[
"https://react.dev/learn",
"https://react.dev/reference/react",
])
# Build vector index
llm = OpenAI(model="gpt-4o-mini", api_key=os.getenv("OPENAI_API_KEY"))
service_ctx = ServiceContext.from_defaults(llm=llm)
index = VectorStoreIndex.from_documents(docs, service_context=service_ctx)
# Query with retrieved context
query_engine = index.as_query_engine()
response = query_engine.query("How do I use React useEffect with async functions?")
print(response)
Summary
- wigolo runs entirely locally via
npx wigolo, eliminating the need for external search API keys while maintaining privacy. - The
WigoloMcpClientclass (found in bothpackages/wigolo-langchain/wigolo_langchain/client.pyandpackages/wigolo-llamaindex/wigolo_llamaindex/client.py) manages the async subprocess connection and protocol handling. - LangChain users should utilize
WigoloSearchRetrieverfor retrieval chains orWigoloSearchTool/WigoloFetchToolfor agents consumed by CrewAI. - LlamaIndex users should import
WigoloWebReaderandWigoloSearchReaderfrompackages/wigolo-llamaindex/wigolo_llamaindex/reader.pyto generateDocumentobjects forVectorStoreIndex. - All integrations support filtering via
include_domains,category, andmax_resultsparameters.
Frequently Asked Questions
Does wigolo require API keys or external services?
No API keys are required for wigolo itself. The server runs locally on your machine using npx wigolo. You only need API keys for the LLM providers (such as OpenAI) if you are integrating with LangChain or LlamaIndex completion models.
How does the MCP client handle multiple concurrent requests?
The WigoloMcpClient opens the npx wigolo subprocess once upon entering the async context manager and reuses the same JSON-RPC connection for all subsequent calls. This minimizes process spawn overhead and supports concurrent async operations until the context manager exits and cleanly shuts down the subprocess.
Can I restrict searches to specific domains or content types?
Yes. Both the LangChain retrievers/tools and LlamaIndex readers accept an include_domains list parameter to whitelist specific sites. They also support a category parameter (values: "general", "code", "news", "papers", "docs") to narrow result types.
Is CrewAI integration separate from LangChain?
No. CrewAI builds directly on top of LangChain’s tool and agent architecture. Therefore, you use the wigolo-langchain package and pass WigoloSearchTool or WigoloFetchTool instances to your CrewAI Agent objects exactly as you would with standard LangChain agents.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →