# How to Use the MCP Query Server for External Tool Integration with OpenViking

> Integrate external tools with OpenViking using the MCP query server. Expose RAG capabilities via HTTP or stdio for semantic search and knowledge ingestion.

- Repository: [Volcengine/OpenViking](https://github.com/volcengine/OpenViking)
- Tags: how-to-guide
- Published: 2026-03-08

---

**OpenViking's MCP query server exposes RAG capabilities as standardized Model Context Protocol tools, enabling external agents like Claude to perform semantic search and knowledge ingestion via HTTP or stdio transport.**

The OpenViking repository provides a production-ready Model Context Protocol (MCP) query server that transforms its core retrieval-augmented generation (RAG) engine into standardized tools for external tool integration. By implementing the MCP specification in [`examples/mcp-query/server.py`](https://github.com/volcengine/OpenViking/blob/main/examples/mcp-query/server.py), OpenViking allows any MCP-compatible client—such as Claude Desktop, custom Python scripts, or orchestration systems—to query knowledge bases and ingest resources without requiring custom API wrappers or deep integration with the underlying OpenViking Python API.

## Architecture of the OpenViking MCP Query Server

### Core Components and FastMCP Framework

The server architecture relies on **FastMCP**, the MCP framework that hosts the JSON-RPC endpoint and provides the `@mcp.tool()` decorator used to expose Python callables as standardized tools. In [`examples/mcp-query/server.py`](https://github.com/volcengine/OpenViking/blob/main/examples/mcp-query/server.py) (lines 26-28), FastMCP is imported and instantiated to create the server instance.

The actual search and generation logic delegates to the **Recipe** class from OpenViking's core pipeline. The `_get_recipe()` function (lines 41-47 in [`server.py`](https://github.com/volcengine/OpenViking/blob/main/server.py)) initializes a `Recipe` instance configured with the specified data directory and configuration file, providing the semantic search capabilities that the MCP tools wrap.

### Transport Layer Options

The server supports two transport modes configured via command-line arguments in `parse_args()` (lines 76-82 of [`server.py`](https://github.com/volcengine/OpenViking/blob/main/server.py)):

- **streamable-http**: Default mode exposing an HTTP endpoint on port 2033, suitable for remote integration and microservices architectures.
- **stdio**: Reads JSON-RPC messages from stdin and writes to stdout, enabling direct integration with Claude Desktop and other stdio-based MCP clients.

### Available MCP Tools and Resources

The server exposes three primary tools, each defined as an async function decorated with `@mcp.tool()`:

- **`query`** (lines 64-71): Performs full RAG pipeline queries, accepting parameters including `question`, `top_k`, `temperature`, `max_tokens`, `score_threshold`, and `system_prompt`.
- **`search`** (lines 122-130): Executes semantic search against the vector index, returning relevant chunks without LLM generation.
- **`add_resource`** (lines 162-170): Ingests new documents (local paths or URLs) into the knowledge base, handling chunking and indexing automatically.

Additionally, the server provides a read-only **MCP Resource** at `openviking://status` (lines 144-152), exposing runtime metadata including the configuration path, data directory, and server status for health monitoring.

## Setting Up the MCP Query Server

To begin using the MCP query server for external tool integration with OpenViking, install the dependencies and start the server:

```bash

# Clone the repository and install dependencies

git clone https://github.com/volcengine/OpenViking.git
cd OpenViking
uv venv && source .venv/bin/activate
uv pip install -e .

# Start the server with HTTP transport (default port 2033)

uv run examples/mcp-query/server.py

```

The server supports several configuration flags defined in the `parse_args()` function:

| Flag | Description | Default |
|------|-------------|---------|
| `--config` | Path to OpenViking configuration file ([`ov.conf`](https://github.com/volcengine/OpenViking/blob/main/ov.conf)) | [`./ov.conf`](https://github.com/volcengine/OpenViking/blob/main/./ov.conf) |
| `--data` | Directory containing indexed vector data | `./data` |
| `--host` | HTTP bind address | `127.0.0.1` |
| `--port` | HTTP listening port | `2033` |
| `--transport` | Transport mode: `streamable-http` or `stdio` | `streamable-http` |

When `--transport stdio` is used, the server writes JSON-RPC messages to stdout and reads from stdin, which matches the expectations of the Claude Desktop client.

## Integrating External Tools with the MCP Server

Once running, the MCP query server exposes standardized endpoints that any compatible client can consume. Here are three common integration patterns.

### HTTP JSON-RPC Integration

For custom applications or remote services, communicate directly via HTTP POST requests to the JSON-RPC endpoint:

```python
import requests
import json

url = "http://127.0.0.1:2033/mcp"

payload = {
    "jsonrpc": "2.0",
    "id": 1,
    "method": "query",
    "params": {
        "question": "Explain OpenViking's skill system.",
        "top_k": 3,
        "temperature": 0.7
    }
}

response = requests.post(url, json=payload)
result = response.json()["result"]
print(result)

```

The response includes the generated answer, source citations, and optional timing metadata as implemented in lines 101-118 of [`server.py`](https://github.com/volcengine/OpenViking/blob/main/server.py).

### Stdio Transport for Claude Desktop

For direct integration with Claude Desktop, launch the server with stdio transport:

```bash
python examples/mcp-query/server.py --transport stdio

```

Then add the following configuration to your Claude Desktop settings:

```json
{
  "mcpServers": {
    "openviking": {
      "command": "python",
      "args": [
        "/path/to/OpenViking/examples/mcp-query/server.py",
        "--transport",
        "stdio"
      ]
    }
  }
}

```

This enables Claude to discover the `query`, `search`, and `add_resource` tools automatically and invoke them as native functions.

### Programmatic Python Client

For embedded workflows, spawn the server as a subprocess and communicate via HTTP:

```python
import subprocess
import time
import requests
import json

# Launch server subprocess

proc = subprocess.Popen(
    ["python", "examples/mcp-query/server.py"],
    stdout=subprocess.DEVNULL,
    stderr=subprocess.DEVNULL
)
time.sleep(2)  # Wait for startup

BASE = "http://127.0.0.1:2033/mcp"

def rpc(method, params):
    payload = {
        "jsonrpc": "2.0",
        "id": 1,
        "method": method,
        "params": params
    }
    r = requests.post(BASE, json=payload)
    return r.json()["result"]

# Add a resource

print(rpc("add_resource", {
    "resource_path": "https://example.com/whitepaper.pdf"
}))

# Query the knowledge base

answer = rpc(
    "query",
    {
        "question": "What does OpenViking use for vector embeddings?",
        "top_k": 3
    }
)
print("\n=== Answer ===\n", answer)

# Cleanup

proc.terminate()

```

This pattern is useful for testing or when embedding OpenViking capabilities into larger Python applications without managing separate service deployments.

## Managing Resources Through MCP Tools

The `add_resource` tool enables dynamic knowledge base expansion without server restarts. According to the implementation in [`examples/mcp-query/server.py`](https://github.com/volcengine/OpenViking/blob/main/examples/mcp-query/server.py) (lines 174-199), the tool handles document loading, chunking, and indexing:

```bash

# Using Claude CLI with MCP

claude mcp invoke openviking.add_resource --args '{"resource_path":"./my_docs"}'

# Or via HTTP JSON-RPC

payload = {
    "jsonrpc": "2.0",
    "id": 2,
    "method": "add_resource",
    "params": {
        "resource_path": "https://example.com/report.pdf"
    }
}
requests.post("http://127.0.0.1:2033/mcp", json=payload)

```

The server confirms ingestion by returning processing status, after which the new content is immediately available to the `query` and `search` tools.

## Server Monitoring and Introspection

For operational monitoring, the server exposes the `openviking://status` resource (lines 144-152 in [`server.py`](https://github.com/volcengine/OpenViking/blob/main/server.py)). Retrieve runtime metadata via:

```bash
curl http://127.0.0.1:2033/mcp/resource/openviking://status

```

Response format:

```json
{
  "config_path": "./ov.conf",
  "data_path": "./data",
  "status": "running"
}

```

This endpoint supports health checks in Kubernetes or Docker Compose environments, verifying that the configuration and data paths are correctly loaded according to [`openviking_cli/utils/config/open_viking_config.py`](https://github.com/volcengine/OpenViking/blob/main/openviking_cli/utils/config/open_viking_config.py).

## Summary

- The **MCP query server for external tool integration with OpenViking** exposes RAG capabilities through standardized Model Context Protocol tools, enabling interoperability with Claude Desktop and other MCP clients.
- The architecture combines **FastMCP** for transport, the **Recipe** class for core search logic, and three primary tools: `query`, `search`, and `add_resource`.
- Transport options include **HTTP** (port 2033) for remote integration and **stdio** for direct Claude Desktop compatibility.
- External tools can invoke the server via JSON-RPC over HTTP, stdio pipes, or MCP client libraries, with support for dynamic resource ingestion through the `add_resource` tool and runtime monitoring via the `openviking://status` resource.

## Frequently Asked Questions

### What is the Model Context Protocol (MCP) and why does OpenViking use it?

The **Model Context Protocol (MCP)** is an open standard developed by Anthropic that enables AI systems to discover and invoke external tools through a standardized JSON-RPC interface. OpenViking implements an MCP query server to allow external agents like Claude Desktop to access its RAG capabilities—semantic search, document ingestion, and answer generation—without requiring custom API wrappers or deep integration with the underlying OpenViking Python API.

### How do I configure the MCP query server for Claude Desktop integration?

To integrate with Claude Desktop, launch the server with the `--transport stdio` flag as implemented in [`examples/mcp-query/server.py`](https://github.com/volcengine/OpenViking/blob/main/examples/mcp-query/server.py) (lines 76-82). Then configure Claude Desktop by adding the server command to your MCP settings, specifying the path to [`server.py`](https://github.com/volcengine/OpenViking/blob/main/server.py) and the `stdio` transport argument. This enables Claude to discover the `query`, `search`, and `add_resource` tools automatically and invoke them as native functions.

### Can I use the MCP query server with custom agents instead of Claude?

Yes, the MCP query server supports any client capable of speaking JSON-RPC over HTTP or stdio. Custom agents can connect via HTTP POST requests to `http://localhost:2033/mcp` using standard JSON-RPC 2.0 payloads, or spawn the server as a subprocess with `--transport stdio` and communicate over pipes. The `openviking://status` resource provides runtime health metadata suitable for orchestration systems like Kubernetes.

### What file types can I ingest using the add_resource tool?

The `add_resource` tool accepts both local file paths and remote URLs, delegating to OpenViking's document processing pipeline. According to the implementation in [`examples/mcp-query/server.py`](https://github.com/volcengine/OpenViking/blob/main/examples/mcp-query/server.py) (lines 174-199), the server handles document loading, chunking, and indexing automatically. Supported formats typically include PDF, Markdown, TXT, and other document types compatible with the underlying OpenViking `Recipe` class and AGFS processing pipeline referenced in [`third_party/agfs/agfs-mcp/src/agfs_mcp/server.py`](https://github.com/volcengine/OpenViking/blob/main/third_party/agfs/agfs-mcp/src/agfs_mcp/server.py).