How DSpark Prevents KV Cache Contamination with Request-Stable Slot Mapping

DSpark prevents KV cache contamination by assigning each request a persistent slot index that remains stable across all decoding steps, decoupling draft model state from volatile batch-row positions through the _row_to_slot mapping algorithm.

The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository implements DSpark, a speculative decoding system that maintains isolated per-request state in its draft model's sliding-window cache. Unlike standard vLLM speculative decoding where batch reordering can cause state leakage, DSpark introduces request-stable KV slots that bind request IDs to fixed cache indices regardless of batch composition changes.

Why Volatile Batch-Row Positions Risk State Contamination

In standard speculative decoding implementations, the draft model's KV cache is indexed by batch-row position. When requests complete and the batch is reordered, a request's KV data may shift to different physical indices. This volatility creates contamination risks where stale KV data from a finished request could persist in slots later assigned to new requests. DSpark eliminates this risk by decoupling logical request identity from physical batch position.

The Request-Stable KV Slot Architecture

Persistent Slot Assignment

Each incoming request receives a permanent KV slot that persists for the entire decoding lifecycle. This slot assignment survives batch restructuring, ensuring the draft model always accesses consistent memory locations for a given request ID.

Mapping Container (_req_id_to_slot)

The core data structure self._req_id_to_slot: dict[str, int] maintains the bidirectional mapping between request UUIDs (strings) and their assigned integer slot indices. According to the source code in recipe/overlay/vllm/v1/spec_decode/dspark_proposer.py (lines 30-49), this dictionary ensures O(1) lookup of stable slots during proposal generation.

Free-Slot Pool (_free_slots)

The self._free_slots list maintains a pool of available slot indices. When requests finish, their slots return to this pool for subsequent reuse. The implementation (lines 538-540) demonstrates how finished requests trigger immediate slot reclamation to prevent memory leaks.

Slot Allocation Algorithm (_row_to_slot)

The _row_to_slot(self, req_ids: list[str]) method implements the deterministic mapping logic (lines 43-49). This function executes every proposal step to ensure slot stability:

  • Garbage collection: Removes entries for finished requests not present in the current req_ids list, returning their slots to self._free_slots
  • New request allocation: Assigns the lowest-available free slot from the pool to any previously unseen request IDs
  • Batch ordering: Returns a list of slots corresponding exactly to the input req_ids batch order

# Example: obtaining stable KV slots for the current batch

# `req_ids` is the list of request IDs in batch order

slots = dspark_proposer._row_to_slot(req_ids)

# `slots[i]` is the persistent KV slot that request `req_ids[i]` will use

The method guarantees that even when vLLM reorders the batch as requests complete, the draft model's cross-step state follows the request ID rather than shifting batch positions.

Slot Reclamation and Lifecycle Management

When requests terminate, DSpark immediately reclaims their dedicated slots to maintain pool availability. The reclamation logic (lines 538-540) identifies stale mappings by checking set membership against active requests:


# Internally, DSpark reclaims slots when a request disappears:

live = set(req_ids)                               # currently active requests

for stale in [r for r in self._req_id_to_slot   # find IDs no longer live

              if r not in live]:
    self._free_slots.append(self._req_id_to_slot.pop(stale))

New requests receive the lowest-indexed available slot through a deterministic lowest-first allocation strategy:


# Assign a new slot to a fresh request:

if slot is None:
    slot = self._free_slots.pop(0)               # lowest-first reuse

    self._req_id_to_slot[req_id] = slot

Benefits of Deterministic Slot Assignment

The lowest-first reuse policy provides specific advantages for deployment scenarios. Single-request servers always use slot 0, creating an identity permutation that eliminates mapping overhead while maintaining the contamination-prevention guarantee. This deterministic approach ensures that no stale KV data can leak into new request contexts, as each request ID maps to a stable physical cache location throughout its lifetime.

Summary

  • DSpark implements request-stable KV slots by mapping each request ID to a persistent cache index via self._req_id_to_slot
  • The _row_to_slot method in dspark_proposer.py (lines 43-49) executes every proposal step to maintain stable mappings regardless of batch reordering
  • Finished requests trigger immediate slot reclamation (lines 538-540), returning indices to self._free_slots for reuse
  • Lowest-first allocation ensures single-request deployments use slot 0 exclusively, optimizing the common case while preventing cross-request contamination
  • The draft model's sliding-window cache remains isolated per request because state access follows request ID rather than volatile batch-row position

Frequently Asked Questions

What is the purpose of request-stable KV slots in DSpark?

Request-stable KV slots prevent state contamination by ensuring each request accesses the same physical cache indices throughout its entire decoding lifecycle. This isolation prevents draft model KV data from leaking between requests when vLLM reorders batch positions as other requests complete.

How does the _row_to_slot method ensure stable KV cache access?

The _row_to_slot method maintains a dictionary mapping request IDs to slot integers and executes three operations every proposal step: reclaiming slots from finished requests, assigning lowest-available free slots to new requests, and returning slots in exact batch-row order. This deterministic algorithm decouples cache access from volatile batch positions.

What happens to KV slots when a request finishes?

When a request terminates, the system identifies its entry in self._req_id_to_slot and moves the slot index back to self._free_slots for subsequent allocation. This reclamation occurs in lines 538-540 of dspark_proposer.py, ensuring no slot leaks and preventing memory exhaustion during long-running server operations.

Why does DSpark use a lowest-first slot allocation strategy?

The lowest-first strategy (popping index 0 from self._free_slots) optimizes single-request scenarios where the server consistently uses slot 0, eliminating mapping overhead. This approach also promotes cache locality and simplifies debugging by making slot assignments predictable while maintaining the contamination-prevention invariant.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →