How Prompt Cache Reuse Works in Bonsai-Demo and When It’s Disabled in Speculative Mode

Prompt cache reuse caches tokenized prompts—including vision encodings—across independent API calls when using the llama.cpp backend, but is intentionally disabled when speculative decoding is enabled to force full conversation history reprocessing.

The Bonsai-demo repository leverages the llama.cpp backend to eliminate redundant computation in multi-turn conversations. By persisting the prompt cache across requests, the server avoids re-encoding images or re-tokenizing text for follow-up queries. However, this optimization is deliberately incompatible with speculative decoding mode, which requires fresh context for each generation pass.

How Prompt Cache Reuse Works

The prompt cache mechanism operates exclusively through the llama.cpp backend, storing computed token representations in memory between API calls.

Prompt Caching in llama.cpp

According to VISION.md, the Bonsai-demo uses llama.cpp to cache the tokenized prompt—including any vision-token encoding—across independent API calls. This means that after an image is uploaded once, the model does not need to re-encode the image for subsequent turns in the same conversation. As noted in the documentation, this design makes follow-up questions essentially instantaneous by retrieving pre-computed tokens from memory rather than reprocessing the raw input.

Vision Token Optimization

Vision inputs particularly benefit from this caching strategy. When you send an image in the first turn of a conversation, the base64-encoded image data gets transformed into vision tokens through the CLIP or equivalent vision encoder. Without caching, every subsequent question about that image would require re-running this expensive encoding step. With prompt cache reuse enabled, the server stores these encoded representations, reducing latency to near-zero for follow-up queries referencing the same image data.

MLX Backend Limitations

The cache works only when the server runs the normal llama.cpp path. The MLX backend has no cross-request prompt cache, so every turn re-processes the entire conversation, including image tokens. If you start the server using MLX instead of llama.cpp, you lose the performance benefits of prompt caching entirely, regardless of whether speculative decoding is active.

Why Speculative Decoding Disables the Cache

When speculative decoding is enabled via the experimental draft-drafter mode, the system enforces constraints that fundamentally conflict with cached prompt states.

Full History Reprocessing Requirement

According to SPECULATIVE.md, cross-request prompt-cache reuse is explicitly disabled when speculative decoding is active. The speculative workflow forces the server to re-process the full conversation history for each request so that the drafter can draft tokens against the exact same prompt each time. This ensures consistency between the draft model and the target model, but eliminates the latency benefits of cached prompts.

Single-Slot Configuration

Speculative decoding also forces a single-slot configuration (-np 1), meaning the server cannot parallelize contexts across multiple slots. This architectural constraint further prevents the reuse of cached states between turns, as each request must start with a clean context window to maintain the strict synchronization required between the drafter and main model.

Detecting Speculative Mode and Cache Status

You can verify whether speculative decoding is active—and consequently whether prompt cache reuse is disabled—by examining the response timings from the API.

Check for the presence of draft_n in the response:

curl -s http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"messages":[{"role":"user","content":"Implement quicksort in Python."}],"max_tokens":400}' \
  | jq '.timings | {draft_n, draft_n_accepted}'

If draft_n is present and non-zero, speculative decoding is engaged. This confirms that prompt cache reuse is disabled and the server is reprocessing the full conversation history for each request.

Configuration Examples

Starting with Prompt Cache Enabled (Default)

To run the server with normal prompt-cache reuse active, use the standard llama.cpp startup script:

./scripts/start_llama_server.sh

This configuration maintains cached prompts across requests, optimizing for multi-turn chat and vision workloads.

Starting with Speculative Decoding (Cache Disabled)

To enable speculative decoding—which automatically disables prompt cache reuse—set the environment variable before launching:

BONSAI_SPECULATIVE=1 ./scripts/start_llama_server.sh

This mode trades caching efficiency for generation speed through draft-token speculation.

Vision Request Example

A typical request that benefits from prompt-cache reuse when not in speculative mode:

curl -s http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "messages":[
          {"role":"user","content":[{"type":"text","text":"Describe this photo."},
                                    {"type":"image_url","image_url":{"url":"data:image/png;base64,..."}}]
          }
        ],
        "max_tokens":200
      }' | jq .

After the first call, subsequent turns referencing the same image will be near-instant with caching enabled, but will incur full reprocessing delays if speculative mode is active.

Summary

  • Prompt cache reuse stores tokenized prompts and vision encodings across API calls when using the llama.cpp backend in Bonsai-demo.
  • The MLX backend does not support cross-request prompt caching, requiring full reprocessing every turn.
  • Speculative decoding (BONSAI_SPECULATIVE=1) intentionally disables prompt cache reuse to force full history reprocessing for draft-model consistency.
  • When speculative mode is active, the server runs in single-slot configuration (-np 1) and cannot reuse cached contexts.
  • Check for draft_n in API response timings to confirm whether speculative decoding—and therefore cache disabling—is active.

Frequently Asked Questions

What is prompt cache reuse in Bonsai-demo?

Prompt cache reuse is a performance optimization in the Bonsai-demo that stores computed token representations—including vision encodings—in memory across independent API calls. When using the llama.cpp backend, this prevents the server from re-tokenizing text or re-encoding images for follow-up queries in the same conversation, reducing latency to near-zero for subsequent turns.

Why is prompt cache reuse disabled during speculative decoding?

Speculative decoding disables prompt cache reuse to ensure the draft model processes the exact same full conversation history as the target model for each token generation pass. As documented in SPECULATIVE.md, this constraint prevents cache synchronization issues between the drafter and main model, though it forces complete reprocessing of prompts and increases time-to-first-token for multi-turn chats.

How can I tell if prompt cache reuse is active?

You can verify speculative decoding status by checking API response timings for the draft_n field. If draft_n exists and is non-zero, speculative decoding is engaged and prompt cache reuse is necessarily disabled. Additionally, if you are using the MLX backend rather than llama.cpp, prompt caching is never available regardless of other settings.

Does the MLX backend support prompt cache reuse?

No. According to VISION.md, the MLX backend has no cross-request prompt cache, meaning every API call re-processes the entire conversation history including any image tokens. Only the llama.cpp backend supports prompt cache reuse in the Bonsai-demo architecture.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →