How to Configure vLLM with Tool Calling for the Chat-Completions Backend in Speech-to-Speech
Use the ChatCompletionsApiModelHandler with environment variables pointing to your vLLM server, ensuring the --backend chat-completions flag is set and tools are defined in OpenAI-compatible format.
The Hugging Face speech-to-speech repository provides a fully OpenAI-compatible inference layer that allows you to swap in any server exposing the /v1/chat/completions endpoint—including locally-hosted vLLM instances. When configured correctly, the library automatically translates your tool definitions into the Chat-Completions format and streams partial tool-call deltas into complete ToolCall events.
Understanding the Chat-Completions Backend Architecture
The speech-to-speech library implements tool calling through a dedicated handler that mirrors the OpenAI Chat-Completions protocol. This design allows seamless integration with vLLM without requiring custom adapters.
ChatCompletionsApiModelHandler Implementation
The core logic resides in src/speech_to_speech/LLM/chat_completions_language_model.py, where the ChatCompletionsApiModelHandler class manages the entire lifecycle of a chat request. This handler builds the JSON payload, manages HTTP streaming, and extracts tool invocations from delta responses.
Key methods include:
_chat_messages– Transforms internal message lists into Chat-Completions-compatible dictionaries, ensuring tool arguments are JSON strings and multimodal content follows thetext/image_urlschema._iter_stream_events– Aggregates partial tool-call deltas from the HTTP stream, yieldingTextDeltaobjects for content andToolCallobjects once function invocations are complete._to_chat_tools– Converts generic function definitions into the nested Chat-Completions format ({"type": "function", "function": {...}})._to_chat_tool_choice– Maps user-provided tool selection preferences to thetool_choiceparameter expected by the API.
Configuration via Arguments Classes
The handler receives its settings from ChatCompletionsLanguageModelHandlerArguments, defined in src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py. This dataclass extends the base Responses-API arguments and adds the reasoning_effort parameter for controlling model-specific reasoning depth.
Setting Up vLLM for OpenAI-Compatible Tool Calling
Before connecting speech-to-speech, you must expose your model through vLLM's OpenAI-compatible server with the Chat-Completions endpoint enabled.
Launch the vLLM server with the following command:
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-2-7b-chat-hf \
--port 8000 \
--api-key dummy \
--disable-log-requests \
--chat-completions
This starts a server at http://localhost:8000/v1/chat/completions that understands standard OpenAI function-calling schemas.
Configuring the Speech-to-Speech Pipeline
Once vLLM is running, configure the speech-to-speech pipeline to route requests to your local endpoint.
Environment Variables
The library reads connection details from environment variables that map to the arguments class:
export RESPONSES_API_BASE_URL="http://localhost:8000/v1"
export RESPONSES_API_API_KEY="dummy"
export RESPONSES_API_STREAM="true"
export RESPONSES_API_REASONING_EFFORT="none"
Set RESPONSES_API_BASE_URL to the root path of your vLLM server (including /v1). The RESPONSES_API_STREAM flag must be "true" to enable the streaming parser that handles partial tool-call deltas.
CLI Configuration
Explicitly select the Chat-Completions backend to ensure the pipeline instantiates the correct handler:
python -m speech_to_speech.main \
--backend chat-completions \
--model-id meta-llama/Llama-2-7b-chat-hf \
--system-prompt "You are a helpful assistant with tool access."
The --backend chat-completions flag forces the pipeline to use ChatCompletionsApiModelHandler rather than the default Responses-API handler.
Defining Tools for Function Calling
Tools must be defined as dictionaries conforming to the OpenAI function schema. The backend automatically converts these using _to_chat_tools before sending them to vLLM.
Example tool definition:
tools = [
{
"type": "function",
"name": "search_web",
"description": "Search the web for a query and return the first result.",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "Search query"},
},
"required": ["query"],
},
}
]
Pass these through the --tools CLI flag as JSON, or provide them directly when constructing the Chat object in Python.
Handling Tool Calls in Practice
When the model decides to invoke a function, the _iter_stream_events method aggregates the streaming deltas and emits a ToolCall event containing the function name and arguments.
Here is a complete working example:
import os
from speech_to_speech.LLM.chat import Chat
from speech_to_speech.pipeline import s2s_pipeline
# Configure vLLM endpoint
os.environ["RESPONSES_API_BASE_URL"] = "http://localhost:8000/v1"
os.environ["RESPONSES_API_API_KEY"] = "dummy"
# Define available tools
tools = [
{
"type": "function",
"name": "get_time",
"description": "Return the current UTC time.",
"parameters": {"type": "object", "properties": {}, "required": []},
}
]
# Initialize chat context
chat = Chat(
system="You may call tools when necessary.",
temperature=0.2,
)
# Build pipeline with explicit backend selection
pipeline = s2s_pipeline(
chat=chat,
backend="chat-completions",
tools=tools,
)
# Optional: Register callback to handle tool calls
def on_tool_call(event):
print(f"Executing: {event.item.name} with args {event.item.arguments}")
pipeline.register_callback("tool", on_tool_call)
# Start conversation
pipeline.run()
The pipeline handles the complexity of streaming responses, reassembling partial JSON tool arguments, and notifying your callbacks when complete ToolCall objects are ready.
Summary
- Use
ChatCompletionsApiModelHandlerlocated insrc/speech_to_speech/LLM/chat_completions_language_model.pyto communicate with vLLM's OpenAI-compatible endpoint. - Set environment variables
RESPONSES_API_BASE_URLandRESPONSES_API_API_KEYto point to your local vLLM instance. - Enable streaming via
RESPONSES_API_STREAM=trueto support incremental tool-call parsing through_iter_stream_events. - Define tools using standard OpenAI function schemas; the backend automatically converts them via
_to_chat_tools. - Select the backend explicitly using
--backend chat-completionsor theS2S_BACKENDenvironment variable.
Frequently Asked Questions
Does vLLM need special configuration to support tool calling with speech-to-speech?
No special configuration is required beyond enabling the Chat-Completions endpoint with --chat-completions. Ensure your model supports function calling (e.g., Llama-2-Chat, Mistral-Instruct, or Qwen-Chat variants). The speech-to-speech library handles the protocol translation via _to_chat_tools and _to_chat_tool_choice, making vLLM appear identical to the OpenAI API.
How does the streaming parser handle partial tool-call deltas?
The _iter_stream_events method in chat_completions_language_model.py maintains an internal buffer of tool-call fragments as they arrive from vLLM's SSE stream. When the stream indicates a tool call is complete, it assembles the fragments into a coherent ToolCall object and yields it to the pipeline. This allows real-time audio feedback while waiting for tool arguments to finish streaming.
Can I use the Chat-Completions backend with cloud providers instead of vLLM?
Yes. The ChatCompletionsApiModelHandler is provider-agnostic. Set RESPONSES_API_BASE_URL to any OpenAI-compatible endpoint (such as OpenRouter, Together AI, or local LM Studio instances). The same tool-calling logic applies, though the specific models available and their tool-calling capabilities will vary by provider.
What is the difference between the Chat-Completions backend and the Responses-API backend?
The Chat-Completions backend uses the /v1/chat/completions endpoint and the ChatCompletionsApiModelHandler, while the Responses-API backend targets OpenAI's newer /v1/responses endpoint. For vLLM and most open-source inference servers, you must use the Chat-Completions backend because vLLM implements the older, more widely adopted specification. The Chat-Completions handler also supports the reasoning_effort parameter through the arguments class defined in chat_completions_language_model_arguments.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →