# Partial-Prefill Admission Hotfix Explained: Restoring Decode Lane Balance in DeepSeek-v4

> Understand the partial-prefill admission hotfix restoring decode lane balance in DeepSeek-v4. Learn how it prevents token starvation and ensures fair allocation for active decode requests.

- Repository: [Mia's AI Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark)
- Tags: internals
- Published: 2026-09-09

---

**The partial-prefill admission hotfix reintroduces missing concurrency gates in the vLLM v1 scheduler to cap simultaneous partial prefills, ensuring decode-active requests receive token budget allocation instead of being starved.**

The `MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark` repository implements this critical patch to resolve Issue #27, where unconstrained partial prefill admission monopolized the batch token budget and silently skipped decoding requests. By restoring the intended `max_num_partial_prefills` admission logic, the hotfix preserves balanced throughput between prompt processing and token generation phases.

## Why the Hotfix Is Needed

Upstream vLLM defines the configuration fields `max_num_partial_prefills` and `max_long_partial_prefills` within `SchedulerConfig`, but the v1 scheduler’s admission loop never consults these values. Instead, it only validates against `max_num_seqs` and the aggregate token budget (`max_num_batched_tokens`).

When multiple prefilling requests occupy the front of the `self.running` list, each consumes the full batch token budget per scheduling step. Decode-active requests positioned later in the queue receive `num_new_tokens == 0` and are silently skipped via a `continue` statement. This creates **zero-preemption decode lane starvation** that intensifies with longer prompts, effectively halting generation throughput for active sequences.

## How the Partial-Prefill Admission Hotfix Works

The patch in [`patches/hotfix-dsv4-issue27-partial-prefill-concurrency.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/patches/hotfix-dsv4-issue27-partial-prefill-concurrency.py) injects initialization logic and an admission guard into the scheduler core.

### Parsing the Admission Cap

During `Scheduler` construction, the hotfix reads the environment variable `DSPARK_MAX_INFLIGHT_PREFILLS` and falls back to `SchedulerConfig.max_num_partial_prefills` when the variable is unset or invalid. The parsed value is stored in `self._dspark_max_inflight_prefills` and hard-capped at 3 to prevent excessive concurrency:

```python
from os import environ

# Parse environment override with config fallback

pp_cap_raw = environ.get('DSPARK_MAX_INFLIGHT_PREFILLS', '').strip()
pp_cap = int(pp_cap_raw) if pp_cap_raw else 0
if pp_cap <= 0:
    pp_cap = scheduler_config.max_num_partial_prefills
self._dspark_max_inflight_prefills = min(pp_cap, 3)

```

### Injecting the Admission Guard

At the top of the waiting-admission loop, the hotfix counts in-flight partial prefills directly from `self.running`. It distinguishes between previously admitted prefills and newly scheduled ones by comparing `num_computed_tokens` against `num_prompt_tokens`:

```python
_pp_running = 0
for i, req in enumerate(self.running):
    if i < _pp_old:  # Already admitted prefills

        if req.num_computed_tokens < req.num_prompt_tokens:
            _pp_running += 1
    elif req.num_computed_tokens + num_scheduled_tokens.get(req.request_id, 0) < req.num_tokens:
        _pp_running += 1  # Newly admitted prefills

    
    # Enforce the admission cap

    if _pp_running >= self._dspark_max_inflight_prefills:
        break  # Stop admitting additional prefill requests this step

```

When the count reaches the configured cap (default 1), the loop breaks immediately, preventing additional prefill admission regardless of remaining token budget.

### Diagnostic Support

The hotfix includes optional telemetry activated via the `DSPARK_ISSUE43_SCHED_DIAG` environment variable. This logging provides per-step metrics tracking in-flight prefill counts and triggers warnings when under-counting is detected, enabling operators to verify the admission gate is functioning correctly.

## Effect on Decode Lanes

By capping concurrent partial prefills, the hotfix guarantees that at most one prefill chunk consumes the batch token budget in any scheduling iteration. This preservation of budget capacity allows decode-active requests following the prefill slots in `self.running` to receive non-zero `num_new_tokens` allocations.

The result is **restored preemption for decode lanes**: generation requests are no longer skipped with zero token allocations, and the scheduler resumes balanced token distribution between prefill and decode phases. This eliminates the starvation condition documented in Issue #27, maintaining consistent latency for ongoing generation tasks even under heavy prompt-ingestion loads.

## Summary

- **Missing Logic**: The upstream v1 scheduler ignored `max_num_partial_prefills`, allowing unlimited partial prefills to monopolize token budgets.
- **Admission Gate**: The hotfix parses `DSPARK_MAX_INFLIGHT_PREFILLS` (capped at 3) and enforces it within the scheduling loop.
- **Decode Protection**: Limiting partial prefills ensures decode requests receive token allocations instead of being skipped.
- **Source Location**: Implementation resides in [`patches/hotfix-dsv4-issue27-partial-prefill-concurrency.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/patches/hotfix-dsv4-issue27-partial-prefill-concurrency.py) with diagnostics controlled by `DSPARK_ISSUE43_SCHED_DIAG`.

## Frequently Asked Questions

### What triggers the partial-prefill admission hotfix to break the admission loop?

The hotfix breaks the admission loop when the count of in-flight partial prefills (`_pp_running`) reaches `self._dspark_max_inflight_prefills`, which defaults to 1 unless overridden by the `DSPARK_MAX_INFLIGHT_PREFILLS` environment variable or `SchedulerConfig` settings.

### How does the hotfix differ from standard vLLM scheduler behavior?

Standard vLLM only checks `max_num_seqs` and token budgets during admission, ignoring the `max_num_partial_prefills` configuration. The hotfix adds an explicit concurrency cap specifically for partial prefills, preventing them from exhausting the token budget before decode requests are considered.

### Why is the admission cap hard-limited to 3?

The cap of 3 prevents operator misconfiguration from destabilizing the scheduler with excessive partial prefill concurrency, which would recreate the starvation conditions the hotfix is designed to resolve. Values above 3 risk reintroducing decode lane starvation even with the patch active.

### Where can I verify the hotfix is working correctly?

Enable diagnostic logging by setting `DSPARK_ISSUE43_SCHED_DIAG=1` to view per-step in-flight prefill counts in the scheduler logs. Additionally, the test suite in [`tests/test_issue27_inflight_cap.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/tests/test_issue27_inflight_cap.py) provides behavioral validation that the cap limits prefills and prevents decode starvation under synthetic load.