How to Migrate from a Vector Database to the OpenViking Filesystem Paradigm
Migrating from an external vector database to OpenViking requires exporting your existing embeddings, initializing the VikingFS layer with a VikingVectorIndexBackend, and upserting records using URIs as primary keys, after which all file operations automatically synchronize vector metadata.
The OpenViking project (volcengine/OpenViking) eliminates the traditional separation between file storage and vector search by tightly coupling them through the VikingFS layer. This migration guide explains how to transition from standalone vector databases like Milvus or Pinecone to OpenViking’s unified filesystem paradigm, where vectors are indexed by viking:// URIs and kept in sync automatically during file operations.
Understanding the OpenViking Architecture
OpenViking’s architecture merges file storage and vector-search storage into a single coherent system. The VikingFS layer in openviking/storage/viking_fs.py provides a POSIX-like API (read, write, rm, mv, ls) that operates on logical Viking URIs (viking://…). It translates these URIs to AGFS paths, enforces access control, and synchronizes vector embeddings whenever a file is created, moved, or deleted.
The VikingVectorIndexBackend in openviking/storage/viking_vector_index_backend.py implements a single-collection vector store (VikingDB) using an adapter pattern. All CRUD operations—upsert, delete, and search—are tenant-aware and use the same URI as the key in the vector index. The AGFS (Abstracted Global File System) serves as the actual filesystem backend (local disk or object store), accessed via self.agfs in VikingFS.
When a file operation occurs, VikingFS executes four steps: permission checking via _ensure_access, URI-to-path conversion via _uri_to_path, raw file manipulation via AGFS, and vector index updates via _update_vector_store_uris (on move) or _delete_from_vector_store (on delete). This design ensures the vector store mirrors the filesystem hierarchy exactly.
Step-by-Step Migration Process
Migrating from an external vector database to OpenViking involves five logical stages: exporting existing vectors, initializing the OpenViking stack, upserting records, reconciling filesystem-vector links, and updating client code.
1. Export Existing Vectors
Query your external database (e.g., Milvus, Pinecone) for all records and retrieve the URI, embedding vector, and custom metadata such as owner_space and context_type. Export this data to JSON or CSV format matching OpenViking’s upsert payload structure, which requires id, uri, vector, context_type, owner_space, and account_id fields.
# Example export from Milvus
python export_milvus.py --host $MILVUS_HOST --port $MILVUS_PORT > vectors.json
2. Initialize the OpenViking Vector Store
Call init_viking_fs with an AGFS client, an embedder (or None if you already have embeddings), and a VikingVectorIndexBackend instance. This establishes the connection to the underlying collection via create_collection.
from openviking.storage.viking_fs import init_viking_fs
from openviking.storage.viking_vector_index_backend import VikingVectorIndexBackend
from openviking_cli.utils.config.vectordb_config import VectorDBBackendConfig
from openviking.pyagfs.client import AGFSClient
# Create AGFS client (local filesystem for development)
agfs = AGFSClient(root="/tmp/openviking_agfs")
# Configure vector backend to match your exported dimensions
cfg = VectorDBBackendConfig(
name="context", # collection name
dimension=768, # dimension of your embeddings
distance_metric="IP", # inner product or L2
sparse_weight=0.0,
)
vector_backend = VikingVectorIndexBackend(cfg)
# Initialize filesystem layer (no embedder needed for migration)
vfs = init_viking_fs(
agfs=agfs,
query_embedder=None,
rerank_config=None,
vector_store=vector_backend,
)
3. Create the Collection Schema
Before upserting, ensure the collection exists with the correct schema. The schema must define fields for id, uri, vector, context_type, owner_space, and account_id.
import asyncio
async def setup_collection():
schema = {
"Fields": [
{"FieldName": "id", "FieldType": "string"},
{"FieldName": "uri", "FieldType": "string"},
{"FieldName": "vector", "FieldType": "vector", "Dim": 768},
{"FieldName": "context_type", "FieldType": "string"},
{"FieldName": "owner_space", "FieldType": "string"},
{"FieldName": "account_id", "FieldType": "string"},
]
}
if not await vfs.vector_store.collection_exists():
await vfs.vector_store.create_collection(name="context", schema=schema)
asyncio.run(setup_collection())
4. Upsert Exported Records
Iterate over your exported data and call vector_store.upsert(payload). The payload must contain at least uri and vector; OpenViking filters unknown fields automatically via _filter_known_fields.
import json
async def import_vectors():
with open("vectors.json") as f:
records = json.load(f)
for rec in records:
payload = {
"uri": rec["uri"],
"vector": rec["embedding"], # list[float]
"context_type": rec.get("type", "resource"),
"owner_space": rec.get("owner_space", ""),
"account_id": rec.get("account_id", "default"),
}
await vfs.vector_store.upsert(payload)
asyncio.run(import_vectors())
5. Re-link Filesystem and Vectors (Optional)
If your external database contained vectors for files that no longer exist, run a consistency check to remove stale entries. Use delete_account_data or delete_uris to purge orphaned vectors.
async def cleanup_stale_vectors():
# Remove vectors for deleted accounts
await vfs.vector_store.delete_account_data("old_account_id")
# Or remove specific URIs
await vfs.vector_store.delete_uris(["viking://user/old/file.txt"])
asyncio.run(cleanup_stale_vectors())
6. Update Client Code
Replace legacy SDK calls with OpenViking’s unified API. The viking_fs.search and viking_fs.find methods accept query strings, internally embed them via query_embedder, and route to the vector backend through HierarchicalRetriever.
async def demo_search():
result = await vfs.search(
query="how to reset password",
target_uri="viking://user/myspace/",
limit=5,
)
print(result)
asyncio.run(demo_search())
Complete Migration Implementation
Combine all steps into a single executable script for production migrations:
import asyncio
import json
from openviking.storage.viking_fs import init_viking_fs
from openviking.storage.viking_vector_index_backend import VikingVectorIndexBackend
from openviking_cli.utils.config.vectordb_config import VectorDBBackendConfig
from openviking.pyagfs.client import AGFSClient
async def migrate_from_vectordb(export_path: str):
# Initialize clients
agfs = AGFSClient(root="/tmp/viking_agfs")
cfg = VectorDBBackendConfig(name="context", dimension=768, distance_metric="IP")
backend = VikingVectorIndexBackend(cfg)
vfs = init_viking_fs(agfs=agfs, query_embedder=None, vector_store=backend)
# Setup collection
if not await vfs.vector_store.collection_exists():
await vfs.vector_store.create_collection(
name="context",
schema={"Fields": [
{"FieldName": "id", "FieldType": "string"},
{"FieldName": "uri", "FieldType": "string"},
{"FieldName": "vector", "FieldType": "vector", "Dim": 768},
{"FieldName": "context_type", "FieldType": "string"},
{"FieldName": "owner_space", "FieldType": "string"},
{"FieldName": "account_id", "FieldType": "string"},
]}
)
# Bulk import
with open(export_path) as f:
docs = json.load(f)
for doc in docs:
await vfs.vector_store.upsert({
"uri": doc["uri"],
"vector": doc["embedding"],
"context_type": doc.get("type", "resource"),
"owner_space": doc.get("owner_space", ""),
"account_id": doc.get("account_id", "default"),
})
# Verify
stats = await vfs.vector_store.get_stats()
print(f"Migration complete. Vector store stats: {stats}")
if __name__ == "__main__":
asyncio.run(migrate_from_vectordb("vectors.json"))
Key Implementation Details
URI as Primary Key – Both the filesystem and vector store use viking:// URIs as logical identifiers. The conversion helpers _uri_to_path and _path_to_uri in viking_fs.py (lines 63-78) guarantee a deterministic bijection between the logical namespace and physical AGFS paths.
Tenant-Aware Filtering – The backend automatically injects account-level filters via _tenant_filter (lines 509-527 in viking_vector_index_backend.py), ensuring users only access vectors belonging to their space.
Automatic Synchronization – Vector updates are embedded directly into rm, mv, and write_file code paths. The helper methods _delete_from_vector_store and _update_vector_store_uris eliminate the need for separate background synchronization jobs.
Summary
- Export existing vectors with URIs, embeddings, and metadata from your current database (Milvus, Pinecone, etc.).
- Initialize
VikingFSwithVikingVectorIndexBackendand an AGFS client; use a dummy embedder if you already possess vectors. - Upsert records using the
urifield as the primary key; OpenViking filters unknown fields automatically. - Validate migration via
get_stats()and clean stale vectors withdelete_account_dataordelete_urisif necessary. - Operate using
viking_fs.searchand file operations (mv,rm,write_file), which automatically keep the vector index synchronized.
Frequently Asked Questions
What identifier serves as the primary key in OpenViking's vector store?
OpenViking uses the Viking URI (viking://…) as the primary key for vector records. The VikingFS layer maintains a bijective mapping between these logical URIs and physical AGFS paths through the _uri_to_path and _path_to_uri methods in openviking/storage/viking_fs.py.
Does OpenViking support sparse vectors during migration?
Yes. The VikingVectorIndexBackend supports sparse vectors via the sparse_query_vector parameter in search operations and accepts sparse representations during upsert. Configure sparse_weight in your VectorDBBackendConfig to control hybrid search behavior.
How does OpenViking handle vector consistency when files are moved?
When you call vfs.mv(), the VikingFS layer automatically invokes _update_vector_store_uris to atomically update the URI key in the vector index. Similarly, vfs.rm() triggers _delete_from_vector_store to remove associated embeddings. This ensures the vector store always mirrors the filesystem state without manual intervention.
Can I migrate from Pinecone or Milvus without re-embedding documents?
Yes. If your existing vectors are compatible with the dimensionality configured in VectorDBBackendConfig, you can migrate the raw embeddings directly. Set query_embedder=None when calling init_viking_fs to prevent automatic re-embedding, then upsert your existing vectors using the vector_store.upsert() method.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →