# How to Implement Hotness Scoring for Time-Sensitive Retrieval in OpenViking

> Learn how to implement hotness scoring in OpenViking to improve retrieval relevance. Combine document access frequency and recency for dynamic, time-sensitive results.

- Repository: [Volcengine/OpenViking](https://github.com/volcengine/OpenViking)
- Tags: tutorial
- Published: 2026-03-08

---

**OpenViking boosts retrieval relevance by computing a hotness score that blends document access frequency with recency-based exponential decay, then mixes this score with vector similarity using a configurable weight to prioritize both popular and fresh content.**

OpenViking is an open-source knowledge retrieval system that surfaces time-sensitive information by applying hotness scoring to semantic search results. This mechanism ensures that frequently accessed and recently updated documents rise to the top of retrieval rankings, even when their pure vector similarity might be lower than older alternatives.

## Core Hotness Scoring Algorithm

The pure hotness calculation lives in [`openviking/retrieve/memory_lifecycle.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/retrieve/memory_lifecycle.py). The `hotness_score` function combines two independent signals into a normalized value between `0.0` and `1.0`.

### Frequency Component via Sigmoid

The frequency component applies a sigmoid function to the natural logarithm of the access count. This saturates the influence of extremely popular documents while still rewarding active usage.

The formula uses `log1p(active_count)` to handle zero counts gracefully, passing the result through a standard sigmoid:

```python
freq = 1.0 / (1.0 + math.exp(-math.log1p(active_count)))

```

### Recency Component via Exponential Decay

The recency component calculates the age of the document in days and applies exponential decay based on a configurable half-life. By default, `DEFAULT_HALF_LIFE_DAYS` is set to `7.0`, meaning the recency contribution halves every week.

The decay rate derives from the half-life using the natural logarithm of 2:

```python
age_days = max((now - updated_at).total_seconds() / 86400.0, 0.0)
decay_rate = math.log(2) / half_life_days
recency = math.exp(-decay_rate * age_days)

```

### Final Score Calculation

The function returns the product of frequency and recency, yielding a float in `[0.0, 1.0]`:

```python
return freq * recency

```

The implementation accepts an optional `now` argument to enable deterministic unit testing without side effects.

## Blending Hotness with Semantic Retrieval

The `HierarchicalRetriever` class in [`openviking/retrieve/hierarchical_retriever.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/retrieve/hierarchical_retriever.py) integrates hotness scoring into the retrieval pipeline through three distinct steps.

### Computing the Final Blended Score

In the `_convert_to_matched_contexts` method, the retriever parses the ISO 8601 `updated_at` timestamp using utilities from [`openviking/utils/time_utils.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/utils/time_utils.py), invokes `hotness_score(active_count, updated_at)`, and blends it with the vector similarity using the formula:

```python
final_score = (1 - alpha) * semantic_score + alpha * h_score

```

Here, `semantic_score` represents the vector similarity from the embedding search, while `h_score` is the hotness value computed for that document.

### Configuring the Hotness Influence

The `HOTNESS_ALPHA` class attribute controls the weight of hotness in the final ranking. The default value of `0.2` allocates 20% influence to hotness and 80% to semantic similarity. Setting this to `0.0` disables hotness boosting entirely, while higher values up to `1.0` make recency and frequency dominate the ranking.

### Re-sorting by Blended Rankings

After computing `final_score` for all candidates, the retriever sorts the results in descending order. This re-ranking occurs in [`openviking/retrieve/hierarchical_retriever.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/retrieve/hierarchical_retriever.py) and ensures that a document with moderate vector similarity but high hotness can outrank a semantically closer but stale alternative.

## Practical Implementation Examples

### Calculating Standalone Hotness Scores

To compute hotness outside the retrieval pipeline for analytics or debugging:

```python
from datetime import datetime, timezone, timedelta
from openviking.retrieve.memory_lifecycle import hotness_score

# Simulate a document accessed 5 times, last updated 2 days ago

active_count = 5
updated_at = datetime.now(timezone.utc) - timedelta(days=2)

score = hotness_score(active_count, updated_at)
print(f"Hotness score: {score:.3f}")

```

### Adjusting Decay Rates for Time-Critical Content

For rapidly changing information like news or stock prices, reduce the half-life to increase the penalty for older documents:

```python

# Use a faster decay (half-life = 1 day) for very time-critical data

fast_score = hotness_score(active_count, updated_at, half_life_days=1.0)
print(f"Fast decay hotness: {fast_score:.3f}")

```

### Customizing the Retriever Weight

To increase the influence of hotness in your specific deployment, subclass `HierarchicalRetriever` and override `HOTNESS_ALPHA`:

```python
from openviking.retrieve.hierarchical_retriever import HierarchicalRetriever

class MyRetriever(HierarchicalRetriever):
    HOTNESS_ALPHA = 0.5   # Give hotness a 50% impact

# Use the custom retriever as usual

retriever = MyRetriever(storage=backend, embedder=my_embedder)
results = await retriever.retrieve(query, ctx)

```

### End-to-End Retrieval with Hotness Boost

Complete implementation showing automatic hotness application during query execution:

```python
from openviking_cli.retrieve.types import TypedQuery
from openviking.retrieve.hierarchical_retriever import HierarchicalRetriever

# Build a query object

query = TypedQuery(query="How to reset my password?", context_type=None)

# Initialize with your storage backend and embedder

retriever = HierarchicalRetriever(storage=backend, embedder=embedder)

# Retrieve top-5 results; hotness applied automatically via HOTNESS_ALPHA

matched = await retriever.retrieve(query, ctx, limit=5)

for ctx in matched.matched_contexts:
    print(f"[{ctx.score:.2f}] {ctx.uri} – {ctx.abstract[:80]}...")

```

The `matched_contexts` list returns already sorted by the blended `final_score`, combining semantic relevance with time-sensitive popularity.

## Summary

- **Pure function design**: The `hotness_score` function in [`openviking/retrieve/memory_lifecycle.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/retrieve/memory_lifecycle.py) is stateless and testable, accepting `active_count`, `updated_at`, and optional `now` parameters.
- **Dual-signal scoring**: Hotness combines a sigmoid-transformed access frequency with exponential recency decay based on a configurable half-life (default 7 days).
- **Seamless integration**: `HierarchicalRetriever` automatically blends hotness with vector similarity using the `HOTNESS_ALPHA` weight (default 0.2) and re-sorts results accordingly.
- **Flexible configuration**: Adjust `half_life_days` per-call for different content types, or subclass the retriever to modify `HOTNESS_ALPHA` globally.
- **Zero overhead when disabled**: Setting `HOTNESS_ALPHA = 0.0` removes hotness calculations from the retrieval path entirely.

## Frequently Asked Questions

### How do I completely disable hotness scoring in OpenViking?

Set `HOTNESS_ALPHA = 0.0` in your retriever configuration. According to the source code in [`openviking/retrieve/hierarchical_retriever.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/retrieve/hierarchical_retriever.py), this eliminates the hotness weight from the blending formula `final_score = (1-alpha) * semantic_score + alpha * h_score`, effectively reducing the computation to pure vector similarity ranking.

### What half-life setting should I use for different content types?

Use shorter half-lives for rapidly evolving content and longer periods for stable reference material. The default `DEFAULT_HALF_LIFE_DAYS = 7.0` suits general documentation, but news articles benefit from `half_life_days=1.0` while archival content may use `30.0` or higher. Pass the override directly to `hotness_score(active_count, updated_at, half_life_days=your_value)`.

### Can I test hotness scoring deterministically?

Yes. The `hotness_score` function accepts an optional `now` parameter accepting a `datetime` object. By passing a fixed timestamp in your unit tests rather than relying on the default `datetime.now(timezone.utc)`, you achieve deterministic, reproducible scores regardless of when the test executes.

### Does hotness scoring add significant latency to retrieval?

The overhead is minimal. The calculation involves only basic arithmetic operations and a single `math.exp` call per candidate document. In [`openviking/retrieve/hierarchical_retriever.py`](https://github.com/volcengine/OpenViking/blob/main/openviking/retrieve/hierarchical_retriever.py), the hotness computation occurs during the `_convert_to_matched_contexts` phase, adding microseconds per item compared to the millisecond-scale vector search operations.