# Cache-Aware Compaction Performance Considerations in Reasonix

> Learn about cache-aware compaction performance in Reasonix. Optimize LLM costs by reducing token regeneration and compute overhead with high cache hit rates.

- Repository: [YHH/DeepSeek-Reasonix](https://github.com/esengine/DeepSeek-Reasonix)
- Tags: performance
- Published: 2026-08-11

---

**Cache-aware compaction in Reasonix optimizes LLM costs by deferring summarization when cache hit rates are high, reducing unnecessary token regeneration and compute overhead.**

Reasonix uses **cache-aware compaction** to balance context window limits against the efficiency gains of token caching. By tracking cache hit/miss statistics, the system decides when compaction is truly necessary—preventing expensive summarization runs when cached tokens already cover most of the user's input. This article examines the architectural components, performance trade-offs, and configuration options that make this optimization possible.

## How Cache Metrics Drive Compaction Decisions

The compaction pipeline in Reasonix is event-driven and telemetry-informed. Four key components work together to implement cache-aware behavior.

### Event Definitions: CompactionStarted and CompactionDone

All compaction activity is tracked through structured events defined in the core event system. These events carry a `Trigger` field distinguishing automatic from manual compaction:

```go
// internal/event/event.go
// CompactionStarted marks the start of a context-compaction pass.
// CompactionDone reports a finished compaction pass.
type EventKind int
const (
    CompactionStarted EventKind = iota
    CompactionDone
)

```

[internal/event/event.go#L64-L72](https://github.com/esengine/DeepSeek-Reasonix/blob/main-v2/internal/event/event.go#L64-L72)

The `Trigger` value ("auto" or "manual") allows downstream components to differentiate compaction patterns and optimize accordingly.

### Wire Format Serialization

When compaction events propagate through the system, cache context travels with them. The wire format preserves message metadata needed for cache-hit analysis:

```go
// internal/eventwire/wire.go
// Compaction is the JSON form of an event.Compaction.
type Compaction struct {
    Trigger   string   `json:"trigger,omitempty"`   // "auto" or "manual"
    Messages  []Message `json:"messages,omitempty"` // messages replaced by the summary
    Summary   string   `json:"summary,omitempty"`   // the new compacted text
    Archive   []Message `json:"archive,omitempty"`   // optional raw archive
}

```

[internal/eventwire/wire.go#L330-L350](https://github.com/esengine/DeepSeek-Reasonix/blob/main-v2/internal/eventwire/wire.go#L330-L350)

### Telemetry Bucketing for Cache Correlation

The telemetry sink aggregates compaction frequency alongside cache performance metrics, enabling operators to identify patterns:

```go
// internal/telemetry/sink.go
case event.CompactionStarted:
    add(s.counts, "compaction",
        enumBucket(e.Compaction.Trigger, "auto", "manual"), 1)

```

[internal/telemetry/sink.go#L164-L168](https://github.com/esengine/DeepSeek-Reasonix/blob/main-v2/internal/telemetry/sink.go#L164-L168)

By bucketing compaction counts by trigger type, the system surfaces whether automatic compaction is firing too frequently relative to cache efficiency.

### Guardian Trigger Logic

The **guardian** component implements the core cache-aware decision logic. Before emitting a `CompactionStarted` event, it evaluates the ratio of `CacheHitTokens` to total tokens. When this ratio exceeds a configurable threshold, compaction is deferred even if the raw message count exceeds the window limit.

[internal/guardian/guardian.go](https://github.com/esengine/DeepSeek-Reasonix/blob/main-v2/internal/guardian/guardian.go)

### Payload Structures for Pre- and Post-Summarization

Two payload types carry compaction state through the extension dispatch system, preserving cache statistics for downstream summarizers:

```go
// internal/extension/dispatch/payloads.go
type CompactionPreparePayload struct {
    Messages []protocol.ProviderMessage `json:"messages"`
    Guidance string                     `json:"guidance,omitempty"`
}
type CompactionCompletePayload struct {
    Summary string `json:"summary"`
}

```

[internal/extension/dispatch/payloads.go#L200-L230](https://github.com/esengine/DeepSeek-Reasonix/blob/main-v2/internal/extension/dispatch/payloads.go#L200-L230)

## Performance Trade-Offs: When to Compact

Cache-aware compaction involves balancing five key performance factors:

| Consideration | Performance Impact | Implementation Detail |
|-------------|------------------|----------------------|
| **Token re-send reduction** | Lower per-turn cost | `CacheHitTokens` tracked in [`sdk/go/types_generated.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/sdk/go/types_generated.go) and aggregated in [`telemetry/sink.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/telemetry/sink.go) |
| **Inference latency** | Faster model responses | Guardian defers compaction when hit-rate exceeds threshold |
| **Summarization overhead** | Extra CPU and tokens consumed | `CompactionPreparePayload` emitted only when justified by low cache coverage |
| **Storage I/O** | Minimal disk impact | [`internal/usagecatalog/catalog.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/usagecatalog/catalog.go) uses SQLite with fallback to in-memory when `CacheDir()` unavailable |
| **Operational tuning** | Data-driven threshold adjustment | Telemetry buckets enable correlation of compaction frequency with hit-rate trends |

### The Cache Hit-Rate Threshold

The critical optimization occurs when the **cache hit-rate** is high. Reasonix may skip compaction entirely—even with 100+ messages in context—if cached tokens comprise 90%+ of the effective input. This avoids:

- **Summarization latency**: 200-500ms additional model call
- **Token costs**: 2-4x tokens for summary generation
- **Context loss**: Potential information degradation from summarization

Conversely, **low hit-rates** trigger automatic compaction to prevent model context overflow and unpredictable latency spikes.

## Code Examples

### Trigger Manual Compaction via Go SDK

For scenarios requiring explicit compaction control, the SDK exposes event-driven compaction:

```go
import "github.com/esengine/DeepSeek-Reasonix/sdk/go"

// Assume `client` is an authenticated Reasonix client.
resp, err := client.Intercept(
    sdk.EventCompactionPrepare, // tell the server we want a compaction
    sdk.CompactionPreparePayload{
        Messages: []sdk.ProviderMessage{
            {Role: sdk.RoleUser, Content: "Long conversation …"},
        },
        Guidance: "Summarise the dialogue",
    },
)
if err != nil { /* handle error */ }

// `resp` will contain a CompactionCompletePayload with the new summary.
fmt.Println("Compacted summary:", resp.Summary)

```

**Key files**: [`sdk/go/types_ext.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/sdk/go/types_ext.go) (event constants), [`internal/extension/dispatch/payloads.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/extension/dispatch/payloads.go) (payload definitions), [`internal/eventwire/wire.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/eventwire/wire.go) (serialization).

### Inspect Session Cache Statistics

Monitor cache efficiency to validate compaction decisions:

```go
stats, err := client.Stats()
if err != nil { /* handle */ }
fmt.Printf("Cache hit: %d tokens, miss: %d tokens\n",
    stats.CacheHitTokens, stats.CacheMissTokens)

```

**Key files**: [`sdk/go/types_generated.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/sdk/go/types_generated.go) (token fields), [`internal/telemetry/sink.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/telemetry/sink.go) (aggregation logic).

### Configure Automatic Compaction Threshold

Adjust the balance point for automatic compaction via configuration:

```yaml
compaction:
  autoThreshold: 0.8   # trigger auto-compaction if cache hit-rate < 80%

```

**Key file**: [`internal/config/config.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/config/config.go) (configuration parsing).

## Key Source Files

| File | Responsibility |
|------|---------------|
| [`internal/event/event.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/event/event.go) | Core compaction event definitions (`CompactionStarted`, `CompactionDone`) |
| [`internal/eventwire/wire.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/eventwire/wire.go) | JSON serialization for compaction data including trigger metadata |
| [`internal/telemetry/sink.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/telemetry/sink.go) | Metric aggregation for compaction frequency and cache performance |
| [`internal/guardian/guardian.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/guardian/guardian.go) | Cache-aware trigger logic for automatic compaction |
| [`internal/extension/dispatch/payloads.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/extension/dispatch/payloads.go) | Payload structures for compaction prepare and complete phases |
| [`sdk/go/types_generated.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/sdk/go/types_generated.go) | Public API for cache token statistics |
| [`sdk/go/types_ext.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/sdk/go/types_ext.go) | SDK event constants for manual compaction triggers |
| [`internal/usagecatalog/catalog.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/usagecatalog/catalog.go) | SQLite-based cache usage persistence |

## Summary

- **Cache-aware compaction defers summarization** when cache hit-rates are high, avoiding unnecessary compute and token costs
- **Four components collaborate**: event definitions, wire serialization, telemetry aggregation, and guardian trigger logic
- **Performance gains** include reduced per-turn latency, lower token consumption, and minimized I/O overhead
- **Operational visibility** comes from telemetry buckets correlating compaction frequency with cache efficiency
- **Tunable thresholds** allow workload-specific optimization via configuration

## Frequently Asked Questions

### What happens when the cache hit-rate exceeds the auto-compaction threshold?

Reasonix postpones compaction and continues appending messages to context. The guardian re-evaluates the hit-rate on each turn, triggering compaction only when cached coverage drops below the threshold or the absolute token limit is reached.

### Does manual compaction bypass cache awareness?

No. Manual compaction via `EventCompactionPrepare` still operates on the same payload structures and respects the underlying cache state. However, it forces the compaction workflow regardless of the hit-rate threshold, useful for explicit context reset scenarios.

### How can I monitor whether compaction is firing too frequently?

Query the telemetry metrics for the `compaction` bucket grouped by `trigger`. A high `auto` count combined with high `cache_hit` metrics indicates the threshold may be set too aggressively. Adjust `compaction.autoThreshold` upward to reduce automatic compaction frequency.

### What storage overhead does the cache catalog introduce?

The cache catalog uses a SQLite database at the system cache directory location, with automatic fallback to in-memory operation when persistent storage is unavailable. Compaction itself writes only the new summary without modifying cache entries, keeping I/O minimal.