How to Optimize Token Consumption with Tiered Context Loading in OpenViking
Use OpenViking's three-level knowledge hierarchy (L0 abstract → L1 overview → L2 full content) to load only the minimal context required for each operation, reducing LLM token usage by up to 90% compared to fetching entire knowledge bases.
OpenViking implements a sophisticated tiered context loading system that minimizes token consumption when retrieving knowledge from large repositories. By structuring every piece of knowledge as a three-level hierarchy and deferring full content loads until explicitly requested, the system ensures that LLM context windows are consumed efficiently. This article explains how to leverage OpenViking's VikingFS abstraction and client APIs to optimize token usage in your applications.
Understanding the Three-Level Knowledge Hierarchy
OpenViking stores every piece of knowledge as a three-level hierarchy that progresses from minimal summaries to complete documents:
| Level | File Suffix | Typical Size | Purpose |
|---|---|---|---|
| L0 | .abstract.md |
≤ 128 tokens | Quick summary for thousands of entries at virtually no cost |
| L1 | .overview.md |
≤ 256 tokens | Detailed description when users drill down |
| L2 | Full content files | Unbounded | Complete resource fetched only when explicitly requested |
When a client requests a directory tree, OpenViking loads only L0 abstracts for each node and defers L1/L2 loading until the caller needs them. This tiered approach ensures that listing a directory containing thousands of files consumes only a few thousand tokens rather than the entire knowledge base.
Implementing Tiered Loading with VikingFS
The VikingFS class in openviking/storage/viking_fs.py provides the core filesystem abstraction that implements tiered context loading through three key mechanisms:
Abstract Reading with abstract()
The abstract() method reads the .abstract.md file for a directory, validating the target and returning only the L0 summary:
# From openviking/storage/viking_fs.py
async def abstract(self, uri: str, ctx: RequestContext) -> str:
# Validates target is a directory and reads .abstract.md
return await self._read_abstract_file(uri, ctx=ctx)
Tree Traversal with Batch Abstract Fetching
The _tree_agent() method walks the directory tree while respecting node_limit and level_limit constraints. After traversal, it calls _batch_fetch_abstracts() to fetch abstracts in parallel with a maximum of 6 concurrent fetches:
# From openviking/storage/viking_fs.py
async def _batch_fetch_abstracts(self, nodes: List[Node], abs_limit: int):
semaphore = asyncio.Semaphore(6) # Max 6 concurrent fetches
async def fetch_with_limit(node):
async with semaphore:
content = await self.abstract(node.uri)
# Truncate to abs_limit tokens if needed
if token_count(content) > abs_limit:
content = truncate(content, abs_limit) + "..."
return content
# Fetch all abstracts in parallel
results = await asyncio.gather(*[fetch_with_limit(n) for n in nodes])
Token-Aware Truncation
If an abstract exceeds the abs_limit parameter, the system truncates the content and appends an ellipsis, guaranteeing a predictable token budget per node regardless of the underlying file size.
Configuring Token Limits and Parallelism
The tiered loading system exposes several parameters to tune token consumption according to your LLM's context window and latency requirements:
| Parameter | Location | Description | Typical Values |
|---|---|---|---|
abs_limit |
tree() / _batch_fetch_abstracts |
Maximum tokens per abstract (truncates if exceeded) | 128 – 256 |
node_limit |
tree() / _tree_agent |
Upper bound on nodes returned (prevents O(N) token blow-up) | 500 – 2000 |
level_limit |
tree() / _tree_agent |
Maximum depth traversed (deeper levels increase payload) | 2 – 3 |
semaphore |
_batch_fetch_abstracts |
Parallelism for abstract fetching (balances latency vs. API rate) | 4 – 8 |
Adjust these values in your tree() calls to match your specific token budget:
# Conservative token usage for large directories
tree = await client.fs.tree(
uri="viking://knowledge_base/",
output="agent",
abs_limit=128, # Short abstracts
node_limit=500, # Limit total entries
level_limit=2, # Shallow traversal
ctx=ctx,
)
Client-Side Lazy Loading Patterns
A typical client workflow demonstrates how tiered loading minimizes token consumption through progressive disclosure:
from openviking import OpenVikingClient, RequestContext
async def optimized_workflow():
client = await OpenVikingClient.create()
ctx = RequestContext(user="alice")
# Step 1: List directory - only L0 abstracts fetched (~256 tokens each)
tree = await client.fs.tree(
uri="viking://my_knowledge/",
output="agent", # Enables tiered loading
abs_limit=256,
node_limit=1000,
ctx=ctx,
)
for entry in tree:
print(f"{entry['rel_path']}: {entry.get('abstract', '')}")
# Step 2: User selects specific entry - load L1 overview
selected = "viking://my_knowledge/specific_topic/"
overview = await client.fs.overview(selected, ctx=ctx) # Loads .overview.md
print("\nOverview:\n", overview)
# Step 3: User requests full document - load L2 content
full_content = await client.fs.read(
f"{selected}/deep_doc.md",
ctx=ctx
) # Loads complete file
print("\nFull content:\n", full_content)
import asyncio
asyncio.run(optimized_workflow())
Because each tier loads on-demand, the number of tokens sent to the LLM equals the sum of the abstracts actually needed, not the size of the entire knowledge base.
Summary
- OpenViking structures knowledge as three tiers: L0 abstracts (≤ 128 tokens), L1 overviews (≤ 256 tokens), and L2 full content (unbounded).
VikingFS._batch_fetch_abstracts()fetches only L0 summaries in parallel (max 6 concurrent) and truncates them toabs_limittokens.- Client workflows should use progressive disclosure: call
tree()withoutput="agent"for L0,overview()for L1, andread()for L2. - Tune
abs_limit,node_limit, andlevel_limitto match your LLM's context window and prevent token budget overruns.
Frequently Asked Questions
What is the maximum token size for each tier in OpenViking?
L0 abstracts are designed to be ≤ 128 tokens, L1 overviews are typically ≤ 256 tokens, and L2 full content has no fixed limit. The abs_limit parameter in tree() calls enforces truncation at your specified threshold, ensuring predictable token budgets regardless of underlying file sizes.
How does OpenViking prevent loading too many nodes at once?
The VikingFS._tree_agent() method respects both node_limit (total entries returned) and level_limit (directory depth). These constraints prevent O(N) token blow-up when traversing large knowledge bases, ensuring that even directories with thousands of files return only a bounded token payload.
Can I adjust the parallelism of abstract fetching?
Yes. The _batch_fetch_abstracts() method uses an asyncio.Semaphore with a default value of 6 concurrent fetches. You can modify this semaphore value in the source code or implement custom batching logic to balance between latency (higher parallelism) and API rate limits (lower parallelism).
When should I use L1 overview versus L2 full content?
Use L1 overviews (client.fs.overview()) when users need more context than the abstract provides but don't require complete documents—ideal for search results previews or category browsing. Use L2 full content (client.fs.read()) only when the user explicitly opens a specific document or asks a detailed question requiring the complete resource.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →