How Litho's Caching Mechanism Speeds Up Documentation Generation in deepwiki-rs
Litho caches LLM prompts and compression results using MD5-hashed keys stored as JSON files, eliminating redundant API calls and reducing generation time from seconds to milliseconds on cache hits.
Litho, the engine powering the deepwiki-rs documentation generator, implements a sophisticated file-based caching system to accelerate repeated documentation builds. By storing the results of expensive LLM operations and content compression tasks, Litho's caching mechanism avoids redundant network requests and token consumption, delivering deterministic output with minimal latency.
Architecture of Litho's Caching Mechanism
Core Components
The caching system is built around several specialized structures:
| Component | Responsibility | Key Source |
|---|---|---|
| CacheManager | Central manager that creates hash keys, reads/writes JSON files, enforces expiration, and records statistics. | src/cache/mod.rs |
| CacheEntry | Serialized record holding the cached data, a timestamp, the prompt's MD5 hash, optional token usage and model name. | src/cache/mod.rs#L21-L33 |
| CachePerformanceMonitor | Tracks hits, misses, writes, and errors for each cache category (e.g., prompt, prompt_compression). |
src/cache/performance_monitor.rs |
| GeneratorContext | Holds a shared Arc<RwLock<CacheManager>> that all agents can read/write. |
src/generator/context.rs |
| CacheConfig | User-configurable options: cache directory, enable flag, expiration (hours). | src/config.rs |
Prompt Hashing Strategy
Litho generates a deterministic MD5 hash from the raw prompt string (system + user part) to create unique cache keys.
// src/cache/mod.rs
pub fn hash_prompt(&self, prompt: &str) -> String {
let mut hasher = Md5::new();
hasher.update(prompt.as_bytes());
format!("{:x}", hasher.finalize())
}
The hash becomes the cache key – identical prompts always map to the same file path, ensuring consistent retrieval across runs.
File-Based Storage Structure
A cache entry is saved under the following directory structure:
<cache_dir>/<category>/<hash>.json
The CacheEntry JSON structure includes:
{
"data": "...", // the LLM reply or compressed text
"timestamp": 1700000000, // seconds since UNIX epoch
"prompt_hash": "abc123…",
"token_usage": { … }, // optional, for accurate stats
"model_name": null
}
The set and set_with_tokens methods serialize this structure and write it asynchronously using tokio::fs::write.
Cache Expiration Logic
Before returning a cached value, CacheManager::is_expired checks whether the stored timestamp exceeds the user-defined TTL (config.expire_hours). Expired entries are removed and treated as cache misses, preventing stale data from contaminating new generation runs.
Performance Monitoring
Every cache interaction updates the CachePerformanceMonitor:
record_cache_hit– counts hits and records estimated inference time derived from content lengthrecord_cache_miss– counts missesrecord_cache_write– counts writes
These metrics feed into the Performance Report displayed after a generation run completes.
How Agents Use the Cache
Prompt-Level Caching in AgentExecutor
Every LLM request flows through helper functions in src/generator/agent_executor.rs.
// src/generator/agent_executor.rs
let prompt_key = format!("{}|{}|reply-prompt", prompt_sys, prompt_user);
// Try to get from cache …
if let Some(cached_reply) = context
.cache_manager
.read()
.await
.get::<serde_json::Value>(cache_scope, &prompt_key)
.await?
{
// Cache hit → return cached data
}
If a cached reply exists, Litho prints a localized "cache hit" message and returns it immediately. Otherwise, it calls the LLM, estimates token usage, and writes the result back using set_with_tokens.
Compression Result Caching
Large source files undergo intelligent compression via the Prompt Compressor (src/utils/prompt_compressor.rs).
// src/utils/prompt_compressor.rs
let cache_manager = context.cache_manager.read().await;
if let Ok(Some(cached_result)) = cache_manager
.get_compression_cache(content, content_type)
.await
{
// Return cached compressed content
}
When new compression is performed, the result is stored with set_compression_cache. This prevents repeated LLM-driven compression of identical files across generation runs.
Shared GeneratorContext
All agents receive the same GeneratorContext, which carries the Arc<RwLock<CacheManager>>. The Arc guarantees a single cache instance per process, while the RwLock allows concurrent reads and exclusive writes—essential for high-throughput async pipelines.
Performance Benefits
| Benefit | Explanation |
|---|---|
| Eliminated network round-trip | Cache hits bypass the remote LLM service, removing latency (often > 200 ms per request). |
| Reduced token consumption | Re-using prior answers means fewer tokens billed to the provider, lowering cost. |
| Deterministic output | Cached answers are immutable, giving reproducible documentation across runs. |
| Parallel safety | The asynchronous RwLock lets many agents read simultaneously, while writes lock only briefly. |
| Observability | The performance monitor reports hit-rate, enabling users to tune cache size or expiration. |
Implementation Examples
Direct Cache Access
For manual cache operations outside the standard agent flow:
use deepwiki_rs::cache::CacheManager;
use deepwiki_rs::config::CacheConfig;
// Create a manager (normally done by Litho)
let cfg = CacheConfig {
enabled: true,
cache_dir: std::path::PathBuf::from("./cache"),
expire_hours: 24,
..Default::default()
};
let manager = CacheManager::new(cfg, deepwiki_rs::i18n::TargetLanguage::En);
// Store an arbitrary value
manager
.set("custom_category", "my prompt key", "my cached response".to_string())
.await?;
// Retrieve it later
if let Some(resp) = manager.get::<String>("custom_category", "my prompt key").await? {
println!("Cache hit: {}", resp);
}
See the underlying set/get implementations in [src/cache/mod.rs](https://github.com/sopaco/deepwiki-rs/blob/main/src/cache/mod.rs).
Prompt Caching via Agent Helper
Standard usage through the high-level API:
use deepwiki_rs::generator::agent_executor::{prompt, AgentExecuteParams};
use deepwiki_rs::generator::context::GeneratorContext;
// Assume `ctx` is the shared GeneratorContext already built by Litho
let params = AgentExecuteParams {
prompt_sys: "You are a Rust expert.".into(),
prompt_user: "Explain lifetimes in Rust.".into(),
cache_scope: "prompt".into(),
log_tag: "Rust Lifetimes".into(),
};
let answer = prompt(&ctx, params).await?;
println!("LLM answer (cached if possible): {}", answer);
Internally this uses CacheManager::get → CacheManager::set_with_tokens (see lines 22‑30 & 48‑55 in [src/generator/agent_executor.rs](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/agent_executor.rs)).
Compression Result Caching
When working with large content that requires compression:
let compressed = compressor
.compress(context.clone(), large_markdown, "markdown")
.await?; // The function internally checks cache (see lines 104‑108 in `prompt_compressor.rs`)
println!("Compressed size: {}", compressed.compressed_content.len());
If the same markdown is seen again, the compressor returns the cached result without calling the LLM.
Key Source Files
| File | Role | Link |
|---|---|---|
src/cache/mod.rs |
Core cache manager, hash generation, get/set logic, expiration. | view |
src/cache/performance_monitor.rs |
Tracks hits/misses/writes and builds performance reports. | view |
src/generator/context.rs |
Holds shared Arc<RwLock<CacheManager>> used by every agent. |
view |
src/generator/agent_executor.rs |
High-level LLM prompt helpers that automatically cache responses. | view |
src/utils/prompt_compressor.rs |
Caches the result of intelligent content compression. | view |
src/config.rs |
Defines CacheConfig (enable flag, directory, TTL). |
view |
Summary
- Litho's caching mechanism uses MD5 hashes of raw prompt strings as deterministic cache keys, storing responses as JSON files under categorized directories.
- The
CacheManagerinsrc/cache/mod.rsorchestrates all operations, including asynchronous read/write, expiration checks against configurable TTLs, and token usage bookkeeping. - Agents automatically leverage the cache through
AgentExecuteParamsinsrc/generator/agent_executor.rs, bypassing LLM calls when valid cached entries exist. - Thread-safe concurrent access is ensured via
Arc<RwLock<CacheManager>>shared throughGeneratorContext, supporting high-throughput parallel generation pipelines. - The
CachePerformanceMonitorprovides observability into hit rates and estimated time saved, enabling users to optimize cache expiration settings.
Frequently Asked Questions
What types of operations does Litho cache?
Litho primarily caches LLM prompt responses and content compression results. The system generates an MD5 hash from the raw prompt string (combining system and user instructions) to create deterministic cache keys. These entries are stored as JSON files categorized by operation type—such as prompt for standard LLM calls or prompt_compression for content reduction tasks—ensuring identical inputs always retrieve the same cached output.
How does Litho handle cache expiration?
The CacheManager validates cache entries using the is_expired method before returning data. Each CacheEntry stores a Unix timestamp recorded at creation time, which is compared against the user-defined expire_hours value from CacheConfig. If the elapsed time exceeds the TTL, the entry is treated as a miss, removed from disk, and a fresh LLM call is initiated to populate the cache with updated content.
Is Litho's cache thread-safe for parallel documentation generation?
Yes, the architecture ensures thread-safe concurrent access through Arc<RwLock<CacheManager>>. The GeneratorContext holds this shared pointer, allowing multiple async agents to acquire read locks simultaneously for cache lookups while enforcing exclusive write access only during cache updates. This design supports high-throughput parallel pipelines where numerous files are processed concurrently without cache corruption or race conditions.
How can I monitor cache performance in Litho?
The CachePerformanceMonitor in src/cache/performance_monitor.rs automatically tracks hits, misses, writes, and errors for each cache category. These metrics aggregate into a final Performance Report displayed after generation completes, showing hit rates and estimated inference time saved. Users can leverage these statistics to tune the expire_hours configuration or analyze cache directory growth to optimize storage efficiency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →