How to Integrate OpenViking with Claude Desktop Using the Memory Plugin: Complete Implementation Guide

The OpenViking memory plugin is a Python wrapper that enables Claude Desktop to capture, store, and recall conversation turns as persistent memories through automated backend detection, session management, and vector search capabilities.

OpenViking by volcengine provides an open-source memory and RAG framework for LLM applications. Integrating OpenViking with Claude Desktop using the memory plugin transforms transient AI conversations into a searchable, long-term knowledge base that persists across sessions and improves contextual relevance over time.

Architecture of the Memory Plugin

The memory plugin (examples/claude-memory-plugin/scripts/ov_memory.py) acts as a bridge between Claude Desktop and the OpenViking backend. It abstracts storage complexity while managing four core operations: backend detection, session persistence, memory ingestion, and vector recall.

Backend Detection and Client Initialization

When invoked, the plugin automatically configures its connection through the detect_backend() function (lines 74-92 in ov_memory.py). This function reads the ov.conf file from your Claude project directory and performs health checks to determine whether OpenViking runs as an HTTP server or operates in local filesystem mode.

Based on this detection, the OVClient.__enter__() method (lines 104-122) instantiates the appropriate synchronous client—either SyncOpenViking for local storage or SyncHTTPClient for HTTP endpoints. This client selection logic allows the same plugin code to function across both deployment modes without manual configuration changes.

The Memory Ingestion Pipeline

The plugin captures conversation context through a coordinated sequence of function calls:

  • cmd_ingest_stop() (lines 86-136): Parses the Claude transcript file (transcript.jsonl) using extract_last_turn(), generates a semantic summary via summarize_turn(), and writes both the user prompt and assistant response to the active OpenViking session using cli.add_message().

  • cmd_session_end() (lines 146-176): Persists queued messages by calling cli.commit_session(), which triggers OpenViking’s extraction pipeline implemented in openviking/storage/viking_vector_index_backend.py. This pipeline performs LLM-driven summarization, semantic chunking, embedding generation, and vector indexing.

  • cmd_recall() (lines 193-242): Executes vector searches against user and agent memory roots using cli.find(), then deduplicates, ranks by relevance score, and returns the top-k memories with associated snippets.

The plugin maintains runtime state—including session IDs, backend mode, and last processed turn timestamps—in a JSON state file specified by the --state-file parameter. This state management enables idempotent ingestion and seamless continuity across Claude Desktop restarts.

Step-by-Step Integration

1. Start a Memory Session

Initialize a new OpenViking session from your Claude project directory:

python examples/claude-memory-plugin/scripts/ov_memory.py \
  session-start \
  --project-dir "$(pwd)" \
  --state-file ./ov_memory_state.json

This command invokes detect_backend() to parse ov.conf, creates a new session via the appropriate synchronous client, and stores the session_id in ov_memory_state.json for subsequent operations.

2. Ingest Conversation Turns

After completing a conversation turn, capture it to the memory store:

python examples/claude-memory-plugin/scripts/ov_memory.py \
  ingest-stop \
  --project-dir "$(pwd)" \
  --state-file ./ov_memory_state.json \
  --transcript-path ./transcript.jsonl

The script extracts the last user-assistant exchange from transcript.jsonl, generates a summary using the logic in openviking/utils/summarizer.py (or Claude itself when available), and appends both messages to your active session queue.

3. Commit the Session

Trigger the memory extraction pipeline to persist ingested data to the vector store:

python examples/claude-memory-plugin/scripts/ov_memory.py \
  session-end \
  --project-dir "$(pwd)" \
  --state-file ./ov_memory_state.json

This executes cli.commit_session(), processing all queued messages through OpenViking’s embedding pipeline and indexing them for semantic retrieval.

4. Recall Relevant Memories

Retrieve contextual memories during active conversations:

python examples/claude-memory-plugin/scripts/ov_memory.py \
  recall \
  --project-dir "$(pwd)" \
  --state-file ./ov_memory_state.json \
  --query "How does my project handle retries?" \
  --top-k 5

This command searches both user and agent memory roots, returning deduplicated results ranked by semantic similarity with URIs, scores, and text snippets.

Configuring Claude Desktop Hooks

To automate memory operations without manual CLI invocation, configure Claude Desktop hooks that trigger the plugin at specific workflow stages:

{
  "name": "OpenVikingMemory",
  "executable": "/usr/bin/python3",
  "args": [
    "examples/claude-memory-plugin/scripts/ov_memory.py",
    "session-start",
    "--project-dir", "{{project_dir}}",
    "--state-file", "{{project_dir}}/.ov_memory_state.json"
  ]
}

Create separate hook entries for ingest-stop, session-end, and recall sub-commands, substituting the appropriate command in each configuration. Claude Desktop will execute these scripts automatically at designated lifecycle points, such as session initialization or conversation completion.

Key Implementation Files

Understanding these source files enables customization and debugging of your integration:

Summary

  • The memory plugin provides a thin Python wrapper (ov_memory.py) that abstracts OpenViking’s backend complexity for Claude Desktop integration.
  • Backend detection via detect_backend() automatically configures HTTP or local mode by parsing ov.conf and health-checking endpoints.
  • Session management uses OVClient with synchronous clients (SyncOpenViking or SyncHTTPClient) to maintain persistent connections across Claude restarts.
  • Memory ingestion combines transcript parsing (extract_last_turn()), summarization (summarize_turn()), and message queueing (cli.add_message()) before committing (cli.commit_session()).
  • Vector recall executes semantic search via cli.find() against indexed memory roots with deduplication and ranking.
  • State persistence through --state-file enables idempotent operations and cross-session continuity.

Frequently Asked Questions

What backend modes does the OpenViking memory plugin support?

The plugin supports two backend modes as defined in openviking_cli/utils/config/agfs_config.py: HTTP mode, where OpenViking runs as a separate server process (openviking serve), and local mode, where the plugin operates directly on a filesystem-backed vector store. The detect_backend() function in ov_memory.py (lines 74-92) automatically detects your configuration by reading the ov.conf file and performing health checks on HTTP endpoints.

How does the plugin prevent duplicate memory ingestion?

The plugin maintains a JSON state file specified by the --state-file parameter. This file tracks the session_id, backend mode, and last processed turn identifier. Before ingesting new turns, cmd_ingest_stop() compares the transcript against this state, enabling idempotent operations that skip already-processed conversation segments even if Claude Desktop restarts or hooks re-trigger.

Can I use the memory plugin without running an OpenViking server?

Yes. When detect_backend() identifies local mode in your ov.conf, the OVClient.__enter__() method (lines 104-122) instantiates SyncOpenViking instead of SyncHTTPClient. This local mode operates directly on vector database files without requiring a running HTTP server, though the extraction pipeline (commit_session) still executes LLM-driven processing as implemented in viking_vector_index_backend.py.

What happens during the session commit operation?

When you invoke session-end, the cmd_session_end() function (lines 146-176) calls cli.commit_session(session_id), which triggers OpenViking’s complete memory extraction pipeline. According to the source code in viking_vector_index_backend.py, this pipeline performs semantic chunking of conversation summaries, generates embeddings, and indexes vectors for later retrieval via cmd_recall() and its cli.find() operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →