How to Add a New Web Search Provider to WeKnora: Exa or Metaso Integration Guide
To add a new web search provider like Exa or Metaso to WeKnora, implement a provider class with a search() method that returns standardized results, register the class in internal/websearch/registry.py, and reference it via the retriever_engine_type field when creating a tenant.
WeKnora supports multi-engine retrieval through pluggable web search providers that integrate with tenant-based knowledge systems. Adding a new web search provider to WeKnora requires implementing three components: the API client wrapper, the registry entry, and the tenant configuration mapping.
Understanding WeKnora's Retriever Architecture
WeKnora distinguishes between retriever_type (the logical role such as keywords, vector, or web) and retriever_engine_type (the concrete implementation such as postgres, exa, or metaso). When a tenant is created via WeKnoraClient.create_tenant in mcp-server/weknora_mcp_server.py【mcp-server/weknora_mcp_server.py†L188-L195】, the client sends a retriever_engines list containing descriptors that map these types. The server-side logic then instantiates the appropriate provider class from an internal registry to build the hybrid search pipeline.
Step 1: Implement the Provider Class
Create a new module at internal/websearch/providers/exa.py that wraps the external API. The class must expose a search(self, query: str, limit: int) -> List[Dict] method that returns results in the WeKnora format.
# internal/websearch/providers/exa.py
import os
import requests
from typing import List, Dict
class ExaProvider:
"""Thin wrapper around the Exa search API."""
def __init__(self, api_key: str = None):
self.api_key = api_key or os.getenv("EXA_API_KEY")
if not self.api_key:
raise ValueError("EXA_API_KEY must be set for ExaProvider")
def search(self, query: str, limit: int = 10) -> List[Dict]:
url = "https://api.exa.ai/search"
headers = {"Authorization": f"Bearer {self.api_key}"}
resp = requests.get(url, params={"q": query, "size": limit}, headers=headers)
resp.raise_for_status()
data = resp.json()
# Convert Exa's result format into the generic WeKnora schema
return [
{"title": r["title"], "url": r["url"], "snippet": r["text"]}
for r in data.get("results", [])
]
Key implementation details:
- The constructor reads the API key from the
EXA_API_KEYenvironment variable. - The
searchmethod returns a list of dictionaries containing title, url, and snippet keys, which the downstream hybrid pipeline expects. - Handle authentication errors and API failures with
resp.raise_for_status()to ensure proper error propagation.
Step 2: Register the Provider in the Registry
Add your provider class to the _PROVIDER_REGISTRY dictionary in internal/websearch/registry.py. This follows the same pattern used for parser engines in docreader/parser/registry.py【docreader/parser/registry.py†L35-L48】, where string keys map to concrete classes.
# internal/websearch/registry.py
from .providers.exa import ExaProvider
from .providers.metaso import MetasoProvider # future addition
_PROVIDER_REGISTRY = {
"postgres": PostgresProvider, # existing keyword/vector engine
"exa": ExaProvider, # new web search provider
"metaso": MetasoProvider, # additional provider
}
By registering the mapping, the server can resolve the string "exa" from a tenant's retriever_engine_type to the ExaProvider class at runtime. If your provider requires availability checks at startup (similar to parser engines), implement a check_available method following the ParserEngineRegistry pattern.
Step 3: Configure the Tenant to Use the New Provider
When creating a tenant, include the web search provider in the retriever_engines payload. According to the create_tenant implementation in mcp-server/weknora_mcp_server.py【mcp-server/weknora_mcp_server.py†L188-L195】, the client sends a JSON array mapping retriever types to engine types.
from weknora_client import WeKnoraClient
client = WeKnoraClient(base_url="https://api.weknora.com/v1", api_key="YOUR_MCP_KEY")
client.create_tenant(
name="exa_demo",
description="Tenant with Exa web search integration",
business="default",
retriever_engines=[
{"retriever_type": "keywords", "retriever_engine_type": "postgres"},
{"retriever_type": "vector", "retriever_engine_type": "postgres"},
{"retriever_type": "web", "retriever_engine_type": "exa"},
],
)
The retriever_type: "web" indicates the logical search category, while retriever_engine_type: "exa" tells the server to instantiate your registered provider class.
Step 4: Integrate into the Hybrid Search Pipeline
The server-side service that builds the hybrid search pipeline—typically found in internal/application/service/knowledge_search_service.py or internal/application/service/tenant_skill_verify.py—dispatches queries to the appropriate engine using the registry.
def _build_hybrid_pipeline(tenant_cfg):
pipelines = []
for engine in tenant_cfg["retriever_engines"]:
if engine["retriever_type"] == "web":
provider_cls = _PROVIDER_REGISTRY[engine["retriever_engine_type"]]
web_provider = provider_cls() #instantiates ExaProvider
pipelines.append(web_provider.search)
# ... handling for keywords and vector retrievers ...
return pipelines
If the repository contains an existing web-search adapter (similar to the WikiSearch implementation), you only need to ensure your new provider is registered. Otherwise, create a new adapter following the pattern of existing search implementations.
Testing Your Web Search Integration
Unit testing: Add test cases under mcp-server/tests/ that create a tenant with the Exa engine and mock the API response to assert that search results contain the expected title, url, and snippet fields.
End-to-end testing: Run the server locally using mcp-server/run.py with EXA_API_KEY set in your environment. Call the hybrid search endpoint POST /knowledge-bases/{kb_id}/hybrid-search and verify that responses include items sourced from the Exa API alongside results from other retriever engines.
Summary
- Implement a provider class in
internal/websearch/providers/with a standardizedsearch()method that returns dictionaries containingtitle,url, andsnippet. - Register the class in
internal/websearch/registry.pyby adding an entry to_PROVIDER_REGISTRYthat maps the engine name to your provider class. - Configure the tenant via
WeKnoraClient.create_tenantby settingretriever_typeto"web"andretriever_engine_typeto your engine name (e.g.,"exa"). - Set environment variables (such as
EXA_API_KEY) before starting the server to provide API credentials to your provider.
Frequently Asked Questions
What interface must a web search provider implement in WeKnora?
A web search provider must implement a search(self, query: str, limit: int) -> List[Dict] method that returns a list of dictionaries, each containing title, url, and snippet keys. The constructor typically reads API credentials from environment variables, as demonstrated in the ExaProvider implementation pattern.
Where is the provider registry located in WeKnora?
The provider registry is defined in internal/websearch/registry.py as a dictionary named _PROVIDER_REGISTRY. This follows the architectural pattern established in docreader/parser/registry.py【docreader/parser/registry.py†L35-L48】, where string keys map to concrete provider classes that the server instantiates at runtime.
How do I configure a tenant to use Exa instead of the default search?
When calling WeKnoraClient.create_tenant in mcp-server/weknora_mcp_server.py【mcp-server/weknora_mcp_server.py†L188-L195】, include an additional entry in the retriever_engines list with "retriever_type": "web" and "retriever_engine_type": "exa". The server will automatically route web search queries to your ExaProvider class.
Can I add multiple web search providers to the same WeKnora tenant?
Yes. You can register multiple providers (such as both Exa and Metaso) in internal/websearch/registry.py and reference different engine types in separate retriever_engines entries, or implement logic to select between them based on query characteristics. Each provider operates as an independent engine within the hybrid search pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →