# How DSpark Prevents KV Cache Contamination with Request-Stable Slot Mapping

> DSpark prevents KV cache contamination using request-stable slot mapping. Learn how its _row_to_slot algorithm ensures persistent slot indexes for stable decoding.

- Repository: [Mia's AI Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark)
- Tags: internals
- Published: 2026-09-09

---

**DSpark prevents KV cache contamination by assigning each request a persistent slot index that remains stable across all decoding steps, decoupling draft model state from volatile batch-row positions through the `_row_to_slot` mapping algorithm.**

The `MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark` repository implements DSpark, a speculative decoding system that maintains isolated per-request state in its draft model's sliding-window cache. Unlike standard vLLM speculative decoding where batch reordering can cause state leakage, DSpark introduces **request-stable KV slots** that bind request IDs to fixed cache indices regardless of batch composition changes.

## Why Volatile Batch-Row Positions Risk State Contamination

In standard speculative decoding implementations, the draft model's KV cache is indexed by batch-row position. When requests complete and the batch is reordered, a request's KV data may shift to different physical indices. This volatility creates contamination risks where stale KV data from a finished request could persist in slots later assigned to new requests. DSpark eliminates this risk by decoupling logical request identity from physical batch position.

## The Request-Stable KV Slot Architecture

### Persistent Slot Assignment

Each incoming request receives a permanent KV slot that persists for the entire decoding lifecycle. This slot assignment survives batch restructuring, ensuring the draft model always accesses consistent memory locations for a given request ID.

### Mapping Container (`_req_id_to_slot`)

The core data structure `self._req_id_to_slot: dict[str, int]` maintains the bidirectional mapping between request UUIDs (strings) and their assigned integer slot indices. According to the source code in [`recipe/overlay/vllm/v1/spec_decode/dspark_proposer.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/recipe/overlay/vllm/v1/spec_decode/dspark_proposer.py) (lines 30-49), this dictionary ensures O(1) lookup of stable slots during proposal generation.

### Free-Slot Pool (`_free_slots`)

The `self._free_slots` list maintains a pool of available slot indices. When requests finish, their slots return to this pool for subsequent reuse. The implementation (lines 538-540) demonstrates how finished requests trigger immediate slot reclamation to prevent memory leaks.

## Slot Allocation Algorithm (`_row_to_slot`)

The `_row_to_slot(self, req_ids: list[str])` method implements the deterministic mapping logic (lines 43-49). This function executes every proposal step to ensure slot stability:

- **Garbage collection**: Removes entries for finished requests not present in the current `req_ids` list, returning their slots to `self._free_slots`
- **New request allocation**: Assigns the lowest-available free slot from the pool to any previously unseen request IDs
- **Batch ordering**: Returns a list of slots corresponding exactly to the input `req_ids` batch order

```python

# Example: obtaining stable KV slots for the current batch

# `req_ids` is the list of request IDs in batch order

slots = dspark_proposer._row_to_slot(req_ids)

# `slots[i]` is the persistent KV slot that request `req_ids[i]` will use

```

The method guarantees that even when vLLM reorders the batch as requests complete, the draft model's cross-step state follows the request ID rather than shifting batch positions.

## Slot Reclamation and Lifecycle Management

When requests terminate, DSpark immediately reclaims their dedicated slots to maintain pool availability. The reclamation logic (lines 538-540) identifies stale mappings by checking set membership against active requests:

```python

# Internally, DSpark reclaims slots when a request disappears:

live = set(req_ids)                               # currently active requests

for stale in [r for r in self._req_id_to_slot   # find IDs no longer live

              if r not in live]:
    self._free_slots.append(self._req_id_to_slot.pop(stale))

```

New requests receive the lowest-indexed available slot through a deterministic lowest-first allocation strategy:

```python

# Assign a new slot to a fresh request:

if slot is None:
    slot = self._free_slots.pop(0)               # lowest-first reuse

    self._req_id_to_slot[req_id] = slot

```

## Benefits of Deterministic Slot Assignment

The lowest-first reuse policy provides specific advantages for deployment scenarios. Single-request servers always use slot 0, creating an identity permutation that eliminates mapping overhead while maintaining the contamination-prevention guarantee. This deterministic approach ensures that no stale KV data can leak into new request contexts, as each request ID maps to a stable physical cache location throughout its lifetime.

## Summary

- DSpark implements **request-stable KV slots** by mapping each request ID to a persistent cache index via `self._req_id_to_slot`
- The `_row_to_slot` method in [`dspark_proposer.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/dspark_proposer.py) (lines 43-49) executes every proposal step to maintain stable mappings regardless of batch reordering
- Finished requests trigger immediate slot reclamation (lines 538-540), returning indices to `self._free_slots` for reuse
- Lowest-first allocation ensures single-request deployments use slot 0 exclusively, optimizing the common case while preventing cross-request contamination
- The draft model's sliding-window cache remains isolated per request because state access follows request ID rather than volatile batch-row position

## Frequently Asked Questions

### What is the purpose of request-stable KV slots in DSpark?

Request-stable KV slots prevent state contamination by ensuring each request accesses the same physical cache indices throughout its entire decoding lifecycle. This isolation prevents draft model KV data from leaking between requests when vLLM reorders batch positions as other requests complete.

### How does the `_row_to_slot` method ensure stable KV cache access?

The `_row_to_slot` method maintains a dictionary mapping request IDs to slot integers and executes three operations every proposal step: reclaiming slots from finished requests, assigning lowest-available free slots to new requests, and returning slots in exact batch-row order. This deterministic algorithm decouples cache access from volatile batch positions.

### What happens to KV slots when a request finishes?

When a request terminates, the system identifies its entry in `self._req_id_to_slot` and moves the slot index back to `self._free_slots` for subsequent allocation. This reclamation occurs in lines 538-540 of [`dspark_proposer.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/dspark_proposer.py), ensuring no slot leaks and preventing memory exhaustion during long-running server operations.

### Why does DSpark use a lowest-first slot allocation strategy?

The lowest-first strategy (popping index 0 from `self._free_slots`) optimizes single-request scenarios where the server consistently uses slot 0, eliminating mapping overhead. This approach also promotes cache locality and simplifies debugging by making slot assignments predictable while maintaining the contamination-prevention invariant.