# How to Migrate from a Vector Database to the OpenViking Filesystem Paradigm

> Easily migrate vector databases to OpenViking. Learn to export embeddings, initialize VikingFS, and upsert records with vectorized metadata synchronization. Get started today.

- Repository: [Volcengine/OpenViking](https://github.com/volcengine/OpenViking)
- Tags: migration-guide
- Published: 2026-03-08

---

**Migrating from an external vector database to OpenViking requires exporting your existing embeddings, initializing the VikingFS layer with a VikingVectorIndexBackend, and upserting records using URIs as primary keys, after which all file operations automatically synchronize vector metadata.**

The OpenViking project (volcengine/OpenViking) eliminates the traditional separation between file storage and vector search by tightly coupling them through the **VikingFS** layer. This migration guide explains how to transition from standalone vector databases like Milvus or Pinecone to OpenViking’s unified filesystem paradigm, where vectors are indexed by `viking://` URIs and kept in sync automatically during file operations.

## Understanding the OpenViking Architecture

OpenViking’s architecture merges file storage and vector-search storage into a single coherent system. The **VikingFS** layer in [`openviking/storage/viking_fs.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/storage/viking_fs.py) provides a POSIX-like API (`read`, `write`, `rm`, `mv`, `ls`) that operates on logical **Viking URIs** (`viking://…`). It translates these URIs to AGFS paths, enforces access control, and synchronizes vector embeddings whenever a file is created, moved, or deleted.

The **VikingVectorIndexBackend** in [`openviking/storage/viking_vector_index_backend.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/storage/viking_vector_index_backend.py) implements a single-collection vector store (VikingDB) using an adapter pattern. All CRUD operations—`upsert`, `delete`, and `search`—are tenant-aware and use the same URI as the key in the vector index. The **AGFS (Abstracted Global File System)** serves as the actual filesystem backend (local disk or object store), accessed via `self.agfs` in VikingFS.

When a file operation occurs, VikingFS executes four steps: permission checking via `_ensure_access`, URI-to-path conversion via `_uri_to_path`, raw file manipulation via AGFS, and vector index updates via `_update_vector_store_uris` (on move) or `_delete_from_vector_store` (on delete). This design ensures the vector store mirrors the filesystem hierarchy exactly.

## Step-by-Step Migration Process

Migrating from an external vector database to OpenViking involves five logical stages: exporting existing vectors, initializing the OpenViking stack, upserting records, reconciling filesystem-vector links, and updating client code.

### 1. Export Existing Vectors

Query your external database (e.g., Milvus, Pinecone) for all records and retrieve the **URI**, **embedding vector**, and custom metadata such as `owner_space` and `context_type`. Export this data to JSON or CSV format matching OpenViking’s upsert payload structure, which requires `id`, `uri`, `vector`, `context_type`, `owner_space`, and `account_id` fields.

```bash

# Example export from Milvus

python export_milvus.py --host $MILVUS_HOST --port $MILVUS_PORT > vectors.json

```

### 2. Initialize the OpenViking Vector Store

Call `init_viking_fs` with an AGFS client, an embedder (or `None` if you already have embeddings), and a `VikingVectorIndexBackend` instance. This establishes the connection to the underlying collection via `create_collection`.

```python
from openviking.storage.viking_fs import init_viking_fs
from openviking.storage.viking_vector_index_backend import VikingVectorIndexBackend
from openviking_cli.utils.config.vectordb_config import VectorDBBackendConfig
from openviking.pyagfs.client import AGFSClient

# Create AGFS client (local filesystem for development)

agfs = AGFSClient(root="/tmp/openviking_agfs")

# Configure vector backend to match your exported dimensions

cfg = VectorDBBackendConfig(
    name="context",               # collection name

    dimension=768,                # dimension of your embeddings

    distance_metric="IP",         # inner product or L2

    sparse_weight=0.0,
)
vector_backend = VikingVectorIndexBackend(cfg)

# Initialize filesystem layer (no embedder needed for migration)

vfs = init_viking_fs(
    agfs=agfs,
    query_embedder=None,
    rerank_config=None,
    vector_store=vector_backend,
)

```

### 3. Create the Collection Schema

Before upserting, ensure the collection exists with the correct schema. The schema must define fields for `id`, `uri`, `vector`, `context_type`, `owner_space`, and `account_id`.

```python
import asyncio

async def setup_collection():
    schema = {
        "Fields": [
            {"FieldName": "id", "FieldType": "string"},
            {"FieldName": "uri", "FieldType": "string"},
            {"FieldName": "vector", "FieldType": "vector", "Dim": 768},
            {"FieldName": "context_type", "FieldType": "string"},
            {"FieldName": "owner_space", "FieldType": "string"},
            {"FieldName": "account_id", "FieldType": "string"},
        ]
    }
    if not await vfs.vector_store.collection_exists():
        await vfs.vector_store.create_collection(name="context", schema=schema)

asyncio.run(setup_collection())

```

### 4. Upsert Exported Records

Iterate over your exported data and call `vector_store.upsert(payload)`. The payload must contain at least `uri` and `vector`; OpenViking filters unknown fields automatically via `_filter_known_fields`.

```python
import json

async def import_vectors():
    with open("vectors.json") as f:
        records = json.load(f)
    
    for rec in records:
        payload = {
            "uri": rec["uri"],
            "vector": rec["embedding"],    # list[float]

            "context_type": rec.get("type", "resource"),
            "owner_space": rec.get("owner_space", ""),
            "account_id": rec.get("account_id", "default"),
        }
        await vfs.vector_store.upsert(payload)

asyncio.run(import_vectors())

```

### 5. Re-link Filesystem and Vectors (Optional)

If your external database contained vectors for files that no longer exist, run a consistency check to remove stale entries. Use `delete_account_data` or `delete_uris` to purge orphaned vectors.

```python
async def cleanup_stale_vectors():
    # Remove vectors for deleted accounts

    await vfs.vector_store.delete_account_data("old_account_id")
    # Or remove specific URIs

    await vfs.vector_store.delete_uris(["viking://user/old/file.txt"])

asyncio.run(cleanup_stale_vectors())

```

### 6. Update Client Code

Replace legacy SDK calls with OpenViking’s unified API. The `viking_fs.search` and `viking_fs.find` methods accept query strings, internally embed them via `query_embedder`, and route to the vector backend through `HierarchicalRetriever`.

```python
async def demo_search():
    result = await vfs.search(
        query="how to reset password",
        target_uri="viking://user/myspace/",
        limit=5,
    )
    print(result)

asyncio.run(demo_search())

```

## Complete Migration Implementation

Combine all steps into a single executable script for production migrations:

```python
import asyncio
import json
from openviking.storage.viking_fs import init_viking_fs
from openviking.storage.viking_vector_index_backend import VikingVectorIndexBackend
from openviking_cli.utils.config.vectordb_config import VectorDBBackendConfig
from openviking.pyagfs.client import AGFSClient

async def migrate_from_vectordb(export_path: str):
    # Initialize clients

    agfs = AGFSClient(root="/tmp/viking_agfs")
    cfg = VectorDBBackendConfig(name="context", dimension=768, distance_metric="IP")
    backend = VikingVectorIndexBackend(cfg)
    vfs = init_viking_fs(agfs=agfs, query_embedder=None, vector_store=backend)
    
    # Setup collection

    if not await vfs.vector_store.collection_exists():
        await vfs.vector_store.create_collection(
            name="context",
            schema={"Fields": [
                {"FieldName": "id", "FieldType": "string"},
                {"FieldName": "uri", "FieldType": "string"},
                {"FieldName": "vector", "FieldType": "vector", "Dim": 768},
                {"FieldName": "context_type", "FieldType": "string"},
                {"FieldName": "owner_space", "FieldType": "string"},
                {"FieldName": "account_id", "FieldType": "string"},
            ]}
        )
    
    # Bulk import

    with open(export_path) as f:
        docs = json.load(f)
    
    for doc in docs:
        await vfs.vector_store.upsert({
            "uri": doc["uri"],
            "vector": doc["embedding"],
            "context_type": doc.get("type", "resource"),
            "owner_space": doc.get("owner_space", ""),
            "account_id": doc.get("account_id", "default"),
        })
    
    # Verify

    stats = await vfs.vector_store.get_stats()
    print(f"Migration complete. Vector store stats: {stats}")

if __name__ == "__main__":
    asyncio.run(migrate_from_vectordb("vectors.json"))

```

## Key Implementation Details

**URI as Primary Key** – Both the filesystem and vector store use `viking://` URIs as logical identifiers. The conversion helpers `_uri_to_path` and `_path_to_uri` in [`viking_fs.py`](https://github.com/volcengine/OpenViking/blob/main/viking_fs.py) (lines 63-78) guarantee a deterministic bijection between the logical namespace and physical AGFS paths.

**Tenant-Aware Filtering** – The backend automatically injects account-level filters via `_tenant_filter` (lines 509-527 in [`viking_vector_index_backend.py`](https://github.com/volcengine/OpenViking/blob/main/viking_vector_index_backend.py)), ensuring users only access vectors belonging to their space.

**Automatic Synchronization** – Vector updates are embedded directly into `rm`, `mv`, and `write_file` code paths. The helper methods `_delete_from_vector_store` and `_update_vector_store_uris` eliminate the need for separate background synchronization jobs.

## Summary

- **Export** existing vectors with URIs, embeddings, and metadata from your current database (Milvus, Pinecone, etc.).
- **Initialize** `VikingFS` with `VikingVectorIndexBackend` and an AGFS client; use a dummy embedder if you already possess vectors.
- **Upsert** records using the `uri` field as the primary key; OpenViking filters unknown fields automatically.
- **Validate** migration via `get_stats()` and clean stale vectors with `delete_account_data` or `delete_uris` if necessary.
- **Operate** using `viking_fs.search` and file operations (`mv`, `rm`, `write_file`), which automatically keep the vector index synchronized.

## Frequently Asked Questions

### What identifier serves as the primary key in OpenViking's vector store?

OpenViking uses the **Viking URI** (`viking://…`) as the primary key for vector records. The `VikingFS` layer maintains a bijective mapping between these logical URIs and physical AGFS paths through the `_uri_to_path` and `_path_to_uri` methods in [`openviking/storage/viking_fs.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/storage/viking_fs.py).

### Does OpenViking support sparse vectors during migration?

Yes. The `VikingVectorIndexBackend` supports sparse vectors via the `sparse_query_vector` parameter in search operations and accepts sparse representations during upsert. Configure `sparse_weight` in your `VectorDBBackendConfig` to control hybrid search behavior.

### How does OpenViking handle vector consistency when files are moved?

When you call `vfs.mv()`, the `VikingFS` layer automatically invokes `_update_vector_store_uris` to atomically update the URI key in the vector index. Similarly, `vfs.rm()` triggers `_delete_from_vector_store` to remove associated embeddings. This ensures the vector store always mirrors the filesystem state without manual intervention.

### Can I migrate from Pinecone or Milvus without re-embedding documents?

Yes. If your existing vectors are compatible with the dimensionality configured in `VectorDBBackendConfig`, you can migrate the raw embeddings directly. Set `query_embedder=None` when calling `init_viking_fs` to prevent automatic re-embedding, then upsert your existing vectors using the `vector_store.upsert()` method.