How to Configure HF Inference Providers as LLM Backends in the Speech-to-Speech Library

Configure HF Inference Providers as LLM backends by setting the --llm-provider hf flag alongside --hf-endpoint and --hf-token, which routes requests through the OpenAI-compatible ResponsesApiLanguageModel wrapper to your HF Inference endpoint.

The huggingface/speech-to-speech library treats Hugging Face Inference Endpoints as first-class LLM backends through an OpenAI-compatible abstraction layer. When you configure HF Inference Providers as LLM backends, the pipeline converts standard OpenAI request formats into HF-specific API calls, enabling seamless integration of remote inference services into the speech-to-speech workflow.

Architecture Overview

The integration relies on four core components that handle argument parsing, request normalization, and pipeline orchestration.

Argument Definitions

Two dataclasses manage configuration parameters. LanguageModelHandlerArguments in src/speech_to_speech/arguments_classes/language_model_arguments.py defines generic generation settings including llm_device, llm_torch_dtype, llm_gen_temperature, and llm_gen_max_new_tokens.

ResponsesApiLanguageModelArguments in src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py exposes provider-specific flags such as --hf-endpoint, --hf-token, and provider_disable_reasoning that target the HF Inference service.

OpenAI-Compatible Wrapper

The BaseOpenAICompatibleLanguageModel class in src/speech_to_speech/LLM/base_openai_compatible_language_model.py normalizes all remote LLM interactions. It converts internal request objects into OpenAI-compatible JSON schemas and manages streaming response parsing, ensuring that HF Inference endpoints receive properly formatted payloads regardless of the upstream caller.

Concrete Provider Implementation

ResponsesApiLanguageModel in src/speech_to_speech/LLM/responses_api_language_model.py extends the base wrapper to target HF Inference services specifically. This class constructs HTTP POST requests to the load-balancer URL, injects the Authorization: Bearer <HF_TOKEN> header, and manages provider-specific body fields like disable_reasoning via the extra_body parameter.

Pipeline Wiring

The S2SPipeline in src/speech_to_speech/s2s_pipeline.py orchestrates the end-to-end flow. It instantiates the appropriate LLM handler based on parsed arguments and injects it between the STT (Speech-to-Text) and TTS (Text-to-Speech) components, allowing the HF Inference backend to operate interchangeably with local models.

Configuration Steps

When you configure HF Inference Providers as LLM backends, the library executes the following sequence:

  1. Parse arguments – The argparse layer populates LanguageModelHandlerArguments and ResponsesApiLanguageModelArguments from CLI flags or programmatic inputs.
  2. Instantiate the provider – The pipeline creates a ResponsesApiLanguageModel instance using the supplied endpoint URL and authentication token.
  3. Build extra body – Inside BaseOpenAICompatibleLanguageModel._build_extra_body(), the wrapper adds HF-specific fields such as "disable_reasoning": true to the request payload.
  4. Issue HTTP request – The class POSTs to https://<load-balancer>.us-east-1.aws.endpoints.huggingface.cloud/v1/chat/completions or your custom --hf-endpoint URL.
  5. Parse stream – The wrapper normalizes HF-specific response fields to the standard OpenAI schema for downstream consumption by s2s_pipeline.py.

Implementation Examples

CLI Usage

Use the following command to launch the demo server with an HF Inference backend:

python -m speech_to_speech.demo.server \
    --llm-provider hf \
    --hf-endpoint https://my-lb.us-east-1.aws.endpoints.huggingface.cloud \
    --hf-token $HF_TOKEN \
    --llm-gen-temperature 0.7 \
    --llm-gen-max-new-tokens 512

The --llm-provider hf flag selects ResponsesApiLanguageModel, while --hf-endpoint and --hf-token authenticate against your private inference service. All standard generation flags continue to function as defined in LanguageModelHandlerArguments.

Programmatic Configuration

Instantiate the pipeline directly in Python to configure HF Inference Providers as LLM backends:

from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelHandlerArguments
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import ResponsesApiLanguageModelArguments
from speech_to_speech.s2s_pipeline import S2SPipeline

# Generic LLM generation settings

llm_args = LanguageModelHandlerArguments(
    llm_device="cpu",
    llm_torch_dtype="float32",
    llm_gen_temperature=0.8,
    llm_gen_max_new_tokens=256,
)

# HF Inference specific configuration

hf_args = ResponsesApiLanguageModelArguments(
    hf_endpoint="https://my-lb.us-east-1.aws.endpoints.huggingface.cloud",
    hf_token="hf_XXXXXXXXXXXXXXXXXXXXXXXX",
    provider_disable_reasoning=False,
)

# Build the end-to-end pipeline

pipeline = S2SPipeline(
    llm_handler_args=llm_args,
    hf_llm_args=hf_args,
)

# Execute inference

output = pipeline.run_text_prompt("Explain the principle of quantum tunnelling.")
print(output)

Custom Provider Subclass

For endpoints requiring custom request shapes, subclass ResponsesApiLanguageModel and override _build_extra_body:

from speech_to_speech.LLM.responses_api_language_model import ResponsesApiLanguageModel

class CustomHFModel(ResponsesApiLanguageModel):
    def _build_extra_body(self, request_json):
        extra = super()._build_extra_body(request_json)
        extra["my_custom_flag"] = True
        return extra

Pass llm_class=CustomHFModel when constructing the S2SPipeline to inject your custom logic while retaining the OpenAI-compatible interface.

Summary

  • Use --llm-provider hf to select the Hugging Face Inference backend in the speech-to-speech library.
  • Configure authentication via --hf-token and endpoint routing via --hf-endpoint in ResponsesApiLanguageModelArguments.
  • The BaseOpenAICompatibleLanguageModel wrapper in src/speech_to_speech/LLM/base_openai_compatible_language_model.py handles OpenAI schema compatibility, while ResponsesApiLanguageModel manages HF-specific HTTP headers and body fields.
  • All standard generation parameters (temperature, max tokens, device) remain available through LanguageModelHandlerArguments regardless of backend choice.
  • Override _build_extra_body() in a subclass to support custom HF endpoint features without modifying the core library.

Frequently Asked Questions

Do I need an OpenAI API key to use HF Inference Providers?

No. The ResponsesApiLanguageModel uses an OpenAI-compatible request format internally for normalization purposes, but it targets HF Inference endpoints exclusively. You only need a valid Hugging Face token (HF_TOKEN) for authentication against your inference endpoint.

What is the default endpoint if I do not specify --hf-endpoint?

If omitted, the library defaults to the public Hugging Face Inference endpoint. However, for production workloads, you should explicitly set --hf-endpoint to your dedicated load-balancer URL to ensure consistent performance and avoid rate limits.

Can I use the same generation arguments for local models and HF Inference?

Yes. Parameters like llm_gen_temperature and llm_gen_max_new_tokens are defined in LanguageModelHandlerArguments and apply universally, whether you run a local model or configure HF Inference Providers as LLM backends. The pipeline abstracts the execution differences while preserving the generation interface.

How does streaming work with HF Inference endpoints?

The BaseOpenAICompatibleLanguageModel parses the HF Inference streaming response format and normalizes it to match OpenAI's server-sent events structure. Downstream components in s2s_pipeline.py consume these streams identically to local model outputs, enabling real-time speech-to-speech latency without code changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →