# How to Optimize Token Consumption with Tiered Context Loading in OpenViking

> Optimize token consumption with OpenViking's tiered context loading. Reduce LLM usage up to 90% by loading only necessary context using a three-level knowledge hierarchy.

- Repository: [Volcengine/OpenViking](https://github.com/volcengine/OpenViking)
- Tags: how-to-guide
- Published: 2026-03-08

---

**Use OpenViking's three-level knowledge hierarchy (L0 abstract → L1 overview → L2 full content) to load only the minimal context required for each operation, reducing LLM token usage by up to 90% compared to fetching entire knowledge bases.**

OpenViking implements a sophisticated tiered context loading system that minimizes token consumption when retrieving knowledge from large repositories. By structuring every piece of knowledge as a three-level hierarchy and deferring full content loads until explicitly requested, the system ensures that LLM context windows are consumed efficiently. This article explains how to leverage OpenViking's `VikingFS` abstraction and client APIs to optimize token usage in your applications.

## Understanding the Three-Level Knowledge Hierarchy

OpenViking stores every piece of knowledge as a **three-level hierarchy** that progresses from minimal summaries to complete documents:

| Level | File Suffix | Typical Size | Purpose |
|-------|-------------|--------------|---------|
| **L0** | [`.abstract.md`](https://github.com/volcengine/OpenViking/blob/main/.abstract.md) | ≤ 128 tokens | Quick summary for thousands of entries at virtually no cost |
| **L1** | [`.overview.md`](https://github.com/volcengine/OpenViking/blob/main/.overview.md) | ≤ 256 tokens | Detailed description when users drill down |
| **L2** | Full content files | Unbounded | Complete resource fetched only when explicitly requested |

When a client requests a directory tree, OpenViking **loads only L0 abstracts** for each node and **defers L1/L2 loading** until the caller needs them. This tiered approach ensures that listing a directory containing thousands of files consumes only a few thousand tokens rather than the entire knowledge base.

## Implementing Tiered Loading with VikingFS

The `VikingFS` class in [`openviking/storage/viking_fs.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/storage/viking_fs.py) provides the core filesystem abstraction that implements tiered context loading through three key mechanisms:

### Abstract Reading with `abstract()`

The `abstract()` method reads the [`.abstract.md`](https://github.com/volcengine/OpenViking/blob/main/.abstract.md) file for a directory, validating the target and returning only the L0 summary:

```python

# From openviking/storage/viking_fs.py

async def abstract(self, uri: str, ctx: RequestContext) -> str:
    # Validates target is a directory and reads .abstract.md

    return await self._read_abstract_file(uri, ctx=ctx)

```

### Tree Traversal with Batch Abstract Fetching

The `_tree_agent()` method walks the directory tree while respecting `node_limit` and `level_limit` constraints. After traversal, it calls `_batch_fetch_abstracts()` to **fetch abstracts in parallel** with a maximum of 6 concurrent fetches:

```python

# From openviking/storage/viking_fs.py

async def _batch_fetch_abstracts(self, nodes: List[Node], abs_limit: int):
    semaphore = asyncio.Semaphore(6)  # Max 6 concurrent fetches

    
    async def fetch_with_limit(node):
        async with semaphore:
            content = await self.abstract(node.uri)
            # Truncate to abs_limit tokens if needed

            if token_count(content) > abs_limit:
                content = truncate(content, abs_limit) + "..."
            return content
    
    # Fetch all abstracts in parallel

    results = await asyncio.gather(*[fetch_with_limit(n) for n in nodes])

```

### Token-Aware Truncation

If an abstract exceeds the `abs_limit` parameter, the system truncates the content and appends an ellipsis, guaranteeing a predictable token budget per node regardless of the underlying file size.

## Configuring Token Limits and Parallelism

The tiered loading system exposes several parameters to tune token consumption according to your LLM's context window and latency requirements:

| Parameter | Location | Description | Typical Values |
|-----------|----------|-------------|----------------|
| `abs_limit` | `tree()` / `_batch_fetch_abstracts` | Maximum tokens per abstract (truncates if exceeded) | 128 – 256 |
| `node_limit` | `tree()` / `_tree_agent` | Upper bound on nodes returned (prevents O(N) token blow-up) | 500 – 2000 |
| `level_limit` | `tree()` / `_tree_agent` | Maximum depth traversed (deeper levels increase payload) | 2 – 3 |
| `semaphore` | `_batch_fetch_abstracts` | Parallelism for abstract fetching (balances latency vs. API rate) | 4 – 8 |

Adjust these values in your `tree()` calls to match your specific token budget:

```python

# Conservative token usage for large directories

tree = await client.fs.tree(
    uri="viking://knowledge_base/",
    output="agent",
    abs_limit=128,      # Short abstracts

    node_limit=500,     # Limit total entries

    level_limit=2,      # Shallow traversal

    ctx=ctx,
)

```

## Client-Side Lazy Loading Patterns

A typical client workflow demonstrates how tiered loading minimizes token consumption through progressive disclosure:

```python
from openviking import OpenVikingClient, RequestContext

async def optimized_workflow():
    client = await OpenVikingClient.create()
    ctx = RequestContext(user="alice")

    # Step 1: List directory - only L0 abstracts fetched (~256 tokens each)

    tree = await client.fs.tree(
        uri="viking://my_knowledge/",
        output="agent",          # Enables tiered loading

        abs_limit=256,
        node_limit=1000,
        ctx=ctx,
    )
    
    for entry in tree:
        print(f"{entry['rel_path']}: {entry.get('abstract', '')}")

    # Step 2: User selects specific entry - load L1 overview

    selected = "viking://my_knowledge/specific_topic/"
    overview = await client.fs.overview(selected, ctx=ctx)  # Loads .overview.md

    print("\nOverview:\n", overview)

    # Step 3: User requests full document - load L2 content

    full_content = await client.fs.read(
        f"{selected}/deep_doc.md", 
        ctx=ctx
    )  # Loads complete file

    print("\nFull content:\n", full_content)

import asyncio
asyncio.run(optimized_workflow())

```

Because each tier loads on-demand, the number of tokens sent to the LLM equals **the sum of the abstracts actually needed**, not the size of the entire knowledge base.

## Summary

- **OpenViking structures knowledge as three tiers**: L0 abstracts (≤ 128 tokens), L1 overviews (≤ 256 tokens), and L2 full content (unbounded).
- **`VikingFS._batch_fetch_abstracts()`** fetches only L0 summaries in parallel (max 6 concurrent) and truncates them to `abs_limit` tokens.
- **Client workflows should use progressive disclosure**: call `tree()` with `output="agent"` for L0, `overview()` for L1, and `read()` for L2.
- **Tune `abs_limit`, `node_limit`, and `level_limit`** to match your LLM's context window and prevent token budget overruns.

## Frequently Asked Questions

### What is the maximum token size for each tier in OpenViking?

L0 abstracts are designed to be ≤ 128 tokens, L1 overviews are typically ≤ 256 tokens, and L2 full content has no fixed limit. The `abs_limit` parameter in `tree()` calls enforces truncation at your specified threshold, ensuring predictable token budgets regardless of underlying file sizes.

### How does OpenViking prevent loading too many nodes at once?

The `VikingFS._tree_agent()` method respects both `node_limit` (total entries returned) and `level_limit` (directory depth). These constraints prevent O(N) token blow-up when traversing large knowledge bases, ensuring that even directories with thousands of files return only a bounded token payload.

### Can I adjust the parallelism of abstract fetching?

Yes. The `_batch_fetch_abstracts()` method uses an `asyncio.Semaphore` with a default value of 6 concurrent fetches. You can modify this semaphore value in the source code or implement custom batching logic to balance between latency (higher parallelism) and API rate limits (lower parallelism).

### When should I use L1 overview versus L2 full content?

Use **L1 overviews** (`client.fs.overview()`) when users need more context than the abstract provides but don't require complete documents—ideal for search results previews or category browsing. Use **L2 full content** (`client.fs.read()`) only when the user explicitly opens a specific document or asks a detailed question requiring the complete resource.