Truncating Tool Calls with `finish_reason: "length"`: Implications for Client-Side Retries

When a request hits the max_tokens limit mid-tool-call, the finish_reason: "length" signal unambiguously indicates token exhaustion, allowing clients to retry safely without parsing malformed JSON fragments because the server strips invalid partial tool calls from the response.

In the MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository, a critical hot-fix resolves a stock vLLM bug where token limit truncation during tool generation was erroneously reported as finish_reason: "tool_calls". This correction ensures that clients receive a canonical finish_reason: "length" indicator when the engine stops generating due to token limits, enabling reliable retry logic and eliminating spurious JSON parsing errors.

The Stock vLLM Bug and Token Limit Truncation

In the standard vLLM implementation, when the engine exhausts the max_tokens budget while generating a tool call, it emits FinishReason.LENGTH internally. However, the original code mistakenly mapped this internal signal to finish_reason: "tool_calls" in the API response. According to the analysis in patches/hotfix-dsv4-issue55-tool-truncation.py (lines 4‑6), this mapping caused the arguments field of truncated tool calls to contain partial, non‑JSON strings.

Clients attempting to parse these malformed arguments encountered validation errors such as HTTP 400 “Unterminated string …”. This misclassification prevented clients from distinguishing between a successfully completed tool call and one cut short by token limits, breaking automated retry logic.

How the Hot-Fix Corrects Finish Reason Mapping

The hot‑fix in patches/hotfix-dsv4-issue55-tool-truncation.py implements two critical changes to serve as a canonical source of truth for truncation states.

Streaming Mode Correction

In streaming responses, the fix ensures that finish_reason: "tool_calls" is only reported when the engine did not terminate due to length constraints. As shown in lines 15‑18 of the patch file, the logic checks str(output.finish_reason) != "length" before setting the tool_calls finish reason. If the engine reports FinishReason.LENGTH, the response preserves finish_reason: "length".

Non-Streaming Path Protection

For non‑streaming requests, lines 99‑102 of the patch prevent the server from claiming a successful tool call when truncation occurred. The code explicitly checks if str(output.finish_reason) == "length" and sets is_finish_reason_tool_calls = False, ensuring the API returns "length" rather than a misleading "tool_calls" status.

Removal of Invalid Tool Calls

Both streaming and non‑streaming paths strip any partially‑generated tool_calls whose arguments do not parse as valid JSON. This prevents clients from receiving broken tool call objects that would cause downstream parsing failures.

Implementing Safe Client-Side Retry Logic

With the corrected finish_reason: "length" signal, clients can implement robust retry strategies that detect token exhaustion and adjust parameters accordingly.

Detecting Truncation Unambiguously

Clients should inspect the finish_reason field in the API response. A value of "length" definitively indicates that the max_tokens limit was reached during generation, not that the model completed its reasoning or tool selection.

Retry Logic with Exponential Backoff

When truncation is detected, clients should retry the request with an increased max_tokens budget. The following Python implementation demonstrates this pattern:

import json
import urllib.request

def call_llm(messages, max_tokens=800):
    body = {
        "model": "deepseek-v4-flash-0731",
        "messages": messages,
        "max_tokens": max_tokens,
        "temperature": 0
    }
    req = urllib.request.Request(
        "http://127.0.0.1:8888/v1/chat/completions",
        data=json.dumps(body).encode(),
        headers={"Content-Type": "application/json"},
    )
    with urllib.request.urlopen(req, timeout=30) as r:
        return json.load(r)

def safe_tool_call(messages, max_tokens=800):
    resp = call_llm(messages, max_tokens)
    choice = resp["choices"][0]
    
    # Detect truncation via finish_reason

    if choice["finish_reason"] == "length":
        print("Request truncated – retrying with larger max_tokens")
        return safe_tool_call(messages, max_tokens * 2)
    
    # Process valid tool calls only

    tool_calls = choice["message"].get("tool_calls") or []
    for tc in tool_calls:
        args = json.loads(tc["function"]["arguments"])
        # Handle tool execution...

    return resp

This pattern ensures that truncated requests are automatically retried with double the token budget, while valid responses proceed to tool execution without JSON parsing risks.

Verification Through Test Harnesses

The repository includes dedicated test scripts to verify the truncation behavior. In scripts/tool-battery.py (lines 99‑106), the test harness validates that truncated calls report finish_reason: "length" and contain no broken JSON arguments. Similarly, scripts/test-issue55-tool-truncation.py (lines 8‑10) provides unit tests confirming that the hot‑fix eliminates invalid tool call objects from responses when truncation occurs.

Summary

  • finish_reason: "length" unambiguously signals that the request was cut short by token limits, not by tool completion.
  • Partial tool calls with invalid JSON arguments are automatically stripped from responses in both streaming and non‑streaming modes.
  • Safe retry logic can be implemented by detecting "length" and resubmitting with an increased max_tokens budget.
  • The fix applies to both streaming and non‑streaming inference paths in the DeepSeek-v4-Flash deployment.

Frequently Asked Questions

Why did stock vLLM report finish_reason: "tool_calls" when truncated?

Stock vLLM incorrectly mapped the internal FinishReason.LENGTH signal to the API response type tool_calls whenever a tool call was being generated, regardless of whether it completed. This occurred because the completion logic did not distinguish between successful tool generation and token limit exhaustion during tool generation.

How can my client detect if a tool call was truncated mid-generation?

Your client should check if choice["finish_reason"] == "length" in the API response. This specific value indicates that the max_tokens limit was reached during generation. Unlike the stock behavior, you will not see finish_reason: "tool_calls" paired with partial JSON arguments.

What happens to partial tool calls when truncation occurs?

The server automatically removes any tool calls whose arguments field contains invalid JSON before sending the response to your client. This means you will either receive a clean finish_reason: "length" with no tool calls, or valid tool calls only, eliminating the risk of parsing errors from truncated strings.

Does this fix work for streaming responses as well?

Yes, the hot‑fix applies to both streaming and non‑streaming modes. In streaming responses, the server only emits finish_reason: "tool_calls" when the engine did not stop due to length constraints, and it filters out invalid tool call chunks when finish_reason: "length" is reported.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →