Cache-Aware Compaction Performance Considerations in Reasonix
Cache-aware compaction in Reasonix optimizes LLM costs by deferring summarization when cache hit rates are high, reducing unnecessary token regeneration and compute overhead.
Reasonix uses cache-aware compaction to balance context window limits against the efficiency gains of token caching. By tracking cache hit/miss statistics, the system decides when compaction is truly necessary—preventing expensive summarization runs when cached tokens already cover most of the user's input. This article examines the architectural components, performance trade-offs, and configuration options that make this optimization possible.
How Cache Metrics Drive Compaction Decisions
The compaction pipeline in Reasonix is event-driven and telemetry-informed. Four key components work together to implement cache-aware behavior.
Event Definitions: CompactionStarted and CompactionDone
All compaction activity is tracked through structured events defined in the core event system. These events carry a Trigger field distinguishing automatic from manual compaction:
// internal/event/event.go
// CompactionStarted marks the start of a context-compaction pass.
// CompactionDone reports a finished compaction pass.
type EventKind int
const (
CompactionStarted EventKind = iota
CompactionDone
)
internal/event/event.go#L64-L72
The Trigger value ("auto" or "manual") allows downstream components to differentiate compaction patterns and optimize accordingly.
Wire Format Serialization
When compaction events propagate through the system, cache context travels with them. The wire format preserves message metadata needed for cache-hit analysis:
// internal/eventwire/wire.go
// Compaction is the JSON form of an event.Compaction.
type Compaction struct {
Trigger string `json:"trigger,omitempty"` // "auto" or "manual"
Messages []Message `json:"messages,omitempty"` // messages replaced by the summary
Summary string `json:"summary,omitempty"` // the new compacted text
Archive []Message `json:"archive,omitempty"` // optional raw archive
}
internal/eventwire/wire.go#L330-L350
Telemetry Bucketing for Cache Correlation
The telemetry sink aggregates compaction frequency alongside cache performance metrics, enabling operators to identify patterns:
// internal/telemetry/sink.go
case event.CompactionStarted:
add(s.counts, "compaction",
enumBucket(e.Compaction.Trigger, "auto", "manual"), 1)
internal/telemetry/sink.go#L164-L168
By bucketing compaction counts by trigger type, the system surfaces whether automatic compaction is firing too frequently relative to cache efficiency.
Guardian Trigger Logic
The guardian component implements the core cache-aware decision logic. Before emitting a CompactionStarted event, it evaluates the ratio of CacheHitTokens to total tokens. When this ratio exceeds a configurable threshold, compaction is deferred even if the raw message count exceeds the window limit.
Payload Structures for Pre- and Post-Summarization
Two payload types carry compaction state through the extension dispatch system, preserving cache statistics for downstream summarizers:
// internal/extension/dispatch/payloads.go
type CompactionPreparePayload struct {
Messages []protocol.ProviderMessage `json:"messages"`
Guidance string `json:"guidance,omitempty"`
}
type CompactionCompletePayload struct {
Summary string `json:"summary"`
}
internal/extension/dispatch/payloads.go#L200-L230
Performance Trade-Offs: When to Compact
Cache-aware compaction involves balancing five key performance factors:
| Consideration | Performance Impact | Implementation Detail |
|---|---|---|
| Token re-send reduction | Lower per-turn cost | CacheHitTokens tracked in sdk/go/types_generated.go and aggregated in telemetry/sink.go |
| Inference latency | Faster model responses | Guardian defers compaction when hit-rate exceeds threshold |
| Summarization overhead | Extra CPU and tokens consumed | CompactionPreparePayload emitted only when justified by low cache coverage |
| Storage I/O | Minimal disk impact | internal/usagecatalog/catalog.go uses SQLite with fallback to in-memory when CacheDir() unavailable |
| Operational tuning | Data-driven threshold adjustment | Telemetry buckets enable correlation of compaction frequency with hit-rate trends |
The Cache Hit-Rate Threshold
The critical optimization occurs when the cache hit-rate is high. Reasonix may skip compaction entirely—even with 100+ messages in context—if cached tokens comprise 90%+ of the effective input. This avoids:
- Summarization latency: 200-500ms additional model call
- Token costs: 2-4x tokens for summary generation
- Context loss: Potential information degradation from summarization
Conversely, low hit-rates trigger automatic compaction to prevent model context overflow and unpredictable latency spikes.
Code Examples
Trigger Manual Compaction via Go SDK
For scenarios requiring explicit compaction control, the SDK exposes event-driven compaction:
import "github.com/esengine/DeepSeek-Reasonix/sdk/go"
// Assume `client` is an authenticated Reasonix client.
resp, err := client.Intercept(
sdk.EventCompactionPrepare, // tell the server we want a compaction
sdk.CompactionPreparePayload{
Messages: []sdk.ProviderMessage{
{Role: sdk.RoleUser, Content: "Long conversation …"},
},
Guidance: "Summarise the dialogue",
},
)
if err != nil { /* handle error */ }
// `resp` will contain a CompactionCompletePayload with the new summary.
fmt.Println("Compacted summary:", resp.Summary)
Key files: sdk/go/types_ext.go (event constants), internal/extension/dispatch/payloads.go (payload definitions), internal/eventwire/wire.go (serialization).
Inspect Session Cache Statistics
Monitor cache efficiency to validate compaction decisions:
stats, err := client.Stats()
if err != nil { /* handle */ }
fmt.Printf("Cache hit: %d tokens, miss: %d tokens\n",
stats.CacheHitTokens, stats.CacheMissTokens)
Key files: sdk/go/types_generated.go (token fields), internal/telemetry/sink.go (aggregation logic).
Configure Automatic Compaction Threshold
Adjust the balance point for automatic compaction via configuration:
compaction:
autoThreshold: 0.8 # trigger auto-compaction if cache hit-rate < 80%
Key file: internal/config/config.go (configuration parsing).
Key Source Files
| File | Responsibility |
|---|---|
internal/event/event.go |
Core compaction event definitions (CompactionStarted, CompactionDone) |
internal/eventwire/wire.go |
JSON serialization for compaction data including trigger metadata |
internal/telemetry/sink.go |
Metric aggregation for compaction frequency and cache performance |
internal/guardian/guardian.go |
Cache-aware trigger logic for automatic compaction |
internal/extension/dispatch/payloads.go |
Payload structures for compaction prepare and complete phases |
sdk/go/types_generated.go |
Public API for cache token statistics |
sdk/go/types_ext.go |
SDK event constants for manual compaction triggers |
internal/usagecatalog/catalog.go |
SQLite-based cache usage persistence |
Summary
- Cache-aware compaction defers summarization when cache hit-rates are high, avoiding unnecessary compute and token costs
- Four components collaborate: event definitions, wire serialization, telemetry aggregation, and guardian trigger logic
- Performance gains include reduced per-turn latency, lower token consumption, and minimized I/O overhead
- Operational visibility comes from telemetry buckets correlating compaction frequency with cache efficiency
- Tunable thresholds allow workload-specific optimization via configuration
Frequently Asked Questions
What happens when the cache hit-rate exceeds the auto-compaction threshold?
Reasonix postpones compaction and continues appending messages to context. The guardian re-evaluates the hit-rate on each turn, triggering compaction only when cached coverage drops below the threshold or the absolute token limit is reached.
Does manual compaction bypass cache awareness?
No. Manual compaction via EventCompactionPrepare still operates on the same payload structures and respects the underlying cache state. However, it forces the compaction workflow regardless of the hit-rate threshold, useful for explicit context reset scenarios.
How can I monitor whether compaction is firing too frequently?
Query the telemetry metrics for the compaction bucket grouped by trigger. A high auto count combined with high cache_hit metrics indicates the threshold may be set too aggressively. Adjust compaction.autoThreshold upward to reduce automatic compaction frequency.
What storage overhead does the cache catalog introduce?
The cache catalog uses a SQLite database at the system cache directory location, with automatic fallback to in-memory operation when persistent storage is unavailable. Compaction itself writes only the new summary without modifying cache entries, keeping I/O minimal.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →