# How to Configure vLLM with Tool Calling for the Chat-Completions Backend in Speech-to-Speech

> Configure vLLM with tool calling for the chat-completions backend in speech-to-speech. Use ChatCompletionsApiModelHandler with environment variables and the chat-completions flag.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-11

---

**Use the `ChatCompletionsApiModelHandler` with environment variables pointing to your vLLM server, ensuring the `--backend chat-completions` flag is set and tools are defined in OpenAI-compatible format.**

The Hugging Face `speech-to-speech` repository provides a fully OpenAI-compatible inference layer that allows you to swap in any server exposing the `/v1/chat/completions` endpoint—including locally-hosted vLLM instances. When configured correctly, the library automatically translates your tool definitions into the Chat-Completions format and streams partial tool-call deltas into complete `ToolCall` events.

## Understanding the Chat-Completions Backend Architecture

The speech-to-speech library implements tool calling through a dedicated handler that mirrors the OpenAI Chat-Completions protocol. This design allows seamless integration with vLLM without requiring custom adapters.

### ChatCompletionsApiModelHandler Implementation

The core logic resides in [`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py), where the `ChatCompletionsApiModelHandler` class manages the entire lifecycle of a chat request. This handler builds the JSON payload, manages HTTP streaming, and extracts tool invocations from delta responses.

Key methods include:
- **`_chat_messages`** – Transforms internal message lists into Chat-Completions-compatible dictionaries, ensuring tool arguments are JSON strings and multimodal content follows the `text`/`image_url` schema.
- **`_iter_stream_events`** – Aggregates partial tool-call deltas from the HTTP stream, yielding `TextDelta` objects for content and `ToolCall` objects once function invocations are complete.
- **`_to_chat_tools`** – Converts generic function definitions into the nested Chat-Completions format (`{"type": "function", "function": {...}}`).
- **`_to_chat_tool_choice`** – Maps user-provided tool selection preferences to the `tool_choice` parameter expected by the API.

### Configuration via Arguments Classes

The handler receives its settings from `ChatCompletionsLanguageModelHandlerArguments`, defined in [`src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py). This dataclass extends the base Responses-API arguments and adds the `reasoning_effort` parameter for controlling model-specific reasoning depth.

## Setting Up vLLM for OpenAI-Compatible Tool Calling

Before connecting speech-to-speech, you must expose your model through vLLM's OpenAI-compatible server with the Chat-Completions endpoint enabled.

Launch the vLLM server with the following command:

```bash
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-2-7b-chat-hf \
    --port 8000 \
    --api-key dummy \
    --disable-log-requests \
    --chat-completions

```

This starts a server at `http://localhost:8000/v1/chat/completions` that understands standard OpenAI function-calling schemas.

## Configuring the Speech-to-Speech Pipeline

Once vLLM is running, configure the speech-to-speech pipeline to route requests to your local endpoint.

### Environment Variables

The library reads connection details from environment variables that map to the arguments class:

```bash
export RESPONSES_API_BASE_URL="http://localhost:8000/v1"
export RESPONSES_API_API_KEY="dummy"
export RESPONSES_API_STREAM="true"
export RESPONSES_API_REASONING_EFFORT="none"

```

Set `RESPONSES_API_BASE_URL` to the root path of your vLLM server (including `/v1`). The `RESPONSES_API_STREAM` flag must be `"true"` to enable the streaming parser that handles partial tool-call deltas.

### CLI Configuration

Explicitly select the Chat-Completions backend to ensure the pipeline instantiates the correct handler:

```bash
python -m speech_to_speech.main \
    --backend chat-completions \
    --model-id meta-llama/Llama-2-7b-chat-hf \
    --system-prompt "You are a helpful assistant with tool access."

```

The `--backend chat-completions` flag forces the pipeline to use `ChatCompletionsApiModelHandler` rather than the default Responses-API handler.

### Defining Tools for Function Calling

Tools must be defined as dictionaries conforming to the OpenAI function schema. The backend automatically converts these using `_to_chat_tools` before sending them to vLLM.

Example tool definition:

```python
tools = [
    {
        "type": "function",
        "name": "search_web",
        "description": "Search the web for a query and return the first result.",
        "parameters": {
            "type": "object",
            "properties": {
                "query": {"type": "string", "description": "Search query"},
            },
            "required": ["query"],
        },
    }
]

```

Pass these through the `--tools` CLI flag as JSON, or provide them directly when constructing the `Chat` object in Python.

## Handling Tool Calls in Practice

When the model decides to invoke a function, the `_iter_stream_events` method aggregates the streaming deltas and emits a `ToolCall` event containing the function name and arguments.

Here is a complete working example:

```python
import os
from speech_to_speech.LLM.chat import Chat
from speech_to_speech.pipeline import s2s_pipeline

# Configure vLLM endpoint

os.environ["RESPONSES_API_BASE_URL"] = "http://localhost:8000/v1"
os.environ["RESPONSES_API_API_KEY"] = "dummy"

# Define available tools

tools = [
    {
        "type": "function",
        "name": "get_time",
        "description": "Return the current UTC time.",
        "parameters": {"type": "object", "properties": {}, "required": []},
    }
]

# Initialize chat context

chat = Chat(
    system="You may call tools when necessary.",
    temperature=0.2,
)

# Build pipeline with explicit backend selection

pipeline = s2s_pipeline(
    chat=chat,
    backend="chat-completions",
    tools=tools,
)

# Optional: Register callback to handle tool calls

def on_tool_call(event):
    print(f"Executing: {event.item.name} with args {event.item.arguments}")

pipeline.register_callback("tool", on_tool_call)

# Start conversation

pipeline.run()

```

The pipeline handles the complexity of streaming responses, reassembling partial JSON tool arguments, and notifying your callbacks when complete `ToolCall` objects are ready.

## Summary

- **Use `ChatCompletionsApiModelHandler`** located in [`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py) to communicate with vLLM's OpenAI-compatible endpoint.
- **Set environment variables** `RESPONSES_API_BASE_URL` and `RESPONSES_API_API_KEY` to point to your local vLLM instance.
- **Enable streaming** via `RESPONSES_API_STREAM=true` to support incremental tool-call parsing through `_iter_stream_events`.
- **Define tools** using standard OpenAI function schemas; the backend automatically converts them via `_to_chat_tools`.
- **Select the backend** explicitly using `--backend chat-completions` or the `S2S_BACKEND` environment variable.

## Frequently Asked Questions

### Does vLLM need special configuration to support tool calling with speech-to-speech?

No special configuration is required beyond enabling the Chat-Completions endpoint with `--chat-completions`. Ensure your model supports function calling (e.g., Llama-2-Chat, Mistral-Instruct, or Qwen-Chat variants). The speech-to-speech library handles the protocol translation via `_to_chat_tools` and `_to_chat_tool_choice`, making vLLM appear identical to the OpenAI API.

### How does the streaming parser handle partial tool-call deltas?

The `_iter_stream_events` method in [`chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model.py) maintains an internal buffer of tool-call fragments as they arrive from vLLM's SSE stream. When the stream indicates a tool call is complete, it assembles the fragments into a coherent `ToolCall` object and yields it to the pipeline. This allows real-time audio feedback while waiting for tool arguments to finish streaming.

### Can I use the Chat-Completions backend with cloud providers instead of vLLM?

Yes. The `ChatCompletionsApiModelHandler` is provider-agnostic. Set `RESPONSES_API_BASE_URL` to any OpenAI-compatible endpoint (such as OpenRouter, Together AI, or local LM Studio instances). The same tool-calling logic applies, though the specific models available and their tool-calling capabilities will vary by provider.

### What is the difference between the Chat-Completions backend and the Responses-API backend?

The Chat-Completions backend uses the `/v1/chat/completions` endpoint and the `ChatCompletionsApiModelHandler`, while the Responses-API backend targets OpenAI's newer `/v1/responses` endpoint. For vLLM and most open-source inference servers, you must use the Chat-Completions backend because vLLM implements the older, more widely adopted specification. The Chat-Completions handler also supports the `reasoning_effort` parameter through the arguments class defined in [`chat_completions_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model_arguments.py).