# How to Configure HF Inference Providers as LLM Backends in the Speech-to-Speech Library

> Configure HF Inference Providers as LLM backends in speech-to-speech. Learn to set flags like --llm-provider hf, --hf-endpoint, and --hf-token for seamless integration.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**Configure HF Inference Providers as LLM backends by setting the `--llm-provider hf` flag alongside `--hf-endpoint` and `--hf-token`, which routes requests through the OpenAI-compatible `ResponsesApiLanguageModel` wrapper to your HF Inference endpoint.**

The huggingface/speech-to-speech library treats Hugging Face Inference Endpoints as first-class LLM backends through an OpenAI-compatible abstraction layer. When you configure HF Inference Providers as LLM backends, the pipeline converts standard OpenAI request formats into HF-specific API calls, enabling seamless integration of remote inference services into the speech-to-speech workflow.

## Architecture Overview

The integration relies on four core components that handle argument parsing, request normalization, and pipeline orchestration.

### Argument Definitions

Two dataclasses manage configuration parameters. `LanguageModelHandlerArguments` in [`src/speech_to_speech/arguments_classes/language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/language_model_arguments.py) defines generic generation settings including `llm_device`, `llm_torch_dtype`, `llm_gen_temperature`, and `llm_gen_max_new_tokens`.

`ResponsesApiLanguageModelArguments` in [`src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py) exposes provider-specific flags such as `--hf-endpoint`, `--hf-token`, and `provider_disable_reasoning` that target the HF Inference service.

### OpenAI-Compatible Wrapper

The `BaseOpenAICompatibleLanguageModel` class in [`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py) normalizes all remote LLM interactions. It converts internal request objects into OpenAI-compatible JSON schemas and manages streaming response parsing, ensuring that HF Inference endpoints receive properly formatted payloads regardless of the upstream caller.

### Concrete Provider Implementation

`ResponsesApiLanguageModel` in [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py) extends the base wrapper to target HF Inference services specifically. This class constructs HTTP POST requests to the load-balancer URL, injects the `Authorization: Bearer <HF_TOKEN>` header, and manages provider-specific body fields like `disable_reasoning` via the `extra_body` parameter.

### Pipeline Wiring

The `S2SPipeline` in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) orchestrates the end-to-end flow. It instantiates the appropriate LLM handler based on parsed arguments and injects it between the STT (Speech-to-Text) and TTS (Text-to-Speech) components, allowing the HF Inference backend to operate interchangeably with local models.

## Configuration Steps

When you configure HF Inference Providers as LLM backends, the library executes the following sequence:

1. **Parse arguments** – The `argparse` layer populates `LanguageModelHandlerArguments` and `ResponsesApiLanguageModelArguments` from CLI flags or programmatic inputs.
2. **Instantiate the provider** – The pipeline creates a `ResponsesApiLanguageModel` instance using the supplied endpoint URL and authentication token.
3. **Build extra body** – Inside `BaseOpenAICompatibleLanguageModel._build_extra_body()`, the wrapper adds HF-specific fields such as `"disable_reasoning": true` to the request payload.
4. **Issue HTTP request** – The class POSTs to `https://<load-balancer>.us-east-1.aws.endpoints.huggingface.cloud/v1/chat/completions` or your custom `--hf-endpoint` URL.
5. **Parse stream** – The wrapper normalizes HF-specific response fields to the standard OpenAI schema for downstream consumption by [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).

## Implementation Examples

### CLI Usage

Use the following command to launch the demo server with an HF Inference backend:

```bash
python -m speech_to_speech.demo.server \
    --llm-provider hf \
    --hf-endpoint https://my-lb.us-east-1.aws.endpoints.huggingface.cloud \
    --hf-token $HF_TOKEN \
    --llm-gen-temperature 0.7 \
    --llm-gen-max-new-tokens 512

```

The `--llm-provider hf` flag selects `ResponsesApiLanguageModel`, while `--hf-endpoint` and `--hf-token` authenticate against your private inference service. All standard generation flags continue to function as defined in `LanguageModelHandlerArguments`.

### Programmatic Configuration

Instantiate the pipeline directly in Python to configure HF Inference Providers as LLM backends:

```python
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelHandlerArguments
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import ResponsesApiLanguageModelArguments
from speech_to_speech.s2s_pipeline import S2SPipeline

# Generic LLM generation settings

llm_args = LanguageModelHandlerArguments(
    llm_device="cpu",
    llm_torch_dtype="float32",
    llm_gen_temperature=0.8,
    llm_gen_max_new_tokens=256,
)

# HF Inference specific configuration

hf_args = ResponsesApiLanguageModelArguments(
    hf_endpoint="https://my-lb.us-east-1.aws.endpoints.huggingface.cloud",
    hf_token="hf_XXXXXXXXXXXXXXXXXXXXXXXX",
    provider_disable_reasoning=False,
)

# Build the end-to-end pipeline

pipeline = S2SPipeline(
    llm_handler_args=llm_args,
    hf_llm_args=hf_args,
)

# Execute inference

output = pipeline.run_text_prompt("Explain the principle of quantum tunnelling.")
print(output)

```

### Custom Provider Subclass

For endpoints requiring custom request shapes, subclass `ResponsesApiLanguageModel` and override `_build_extra_body`:

```python
from speech_to_speech.LLM.responses_api_language_model import ResponsesApiLanguageModel

class CustomHFModel(ResponsesApiLanguageModel):
    def _build_extra_body(self, request_json):
        extra = super()._build_extra_body(request_json)
        extra["my_custom_flag"] = True
        return extra

```

Pass `llm_class=CustomHFModel` when constructing the `S2SPipeline` to inject your custom logic while retaining the OpenAI-compatible interface.

## Summary

- Use `--llm-provider hf` to select the Hugging Face Inference backend in the speech-to-speech library.
- Configure authentication via `--hf-token` and endpoint routing via `--hf-endpoint` in `ResponsesApiLanguageModelArguments`.
- The `BaseOpenAICompatibleLanguageModel` wrapper in [`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py) handles OpenAI schema compatibility, while `ResponsesApiLanguageModel` manages HF-specific HTTP headers and body fields.
- All standard generation parameters (temperature, max tokens, device) remain available through `LanguageModelHandlerArguments` regardless of backend choice.
- Override `_build_extra_body()` in a subclass to support custom HF endpoint features without modifying the core library.

## Frequently Asked Questions

### Do I need an OpenAI API key to use HF Inference Providers?

No. The `ResponsesApiLanguageModel` uses an OpenAI-compatible request format internally for normalization purposes, but it targets HF Inference endpoints exclusively. You only need a valid Hugging Face token (`HF_TOKEN`) for authentication against your inference endpoint.

### What is the default endpoint if I do not specify --hf-endpoint?

If omitted, the library defaults to the public Hugging Face Inference endpoint. However, for production workloads, you should explicitly set `--hf-endpoint` to your dedicated load-balancer URL to ensure consistent performance and avoid rate limits.

### Can I use the same generation arguments for local models and HF Inference?

Yes. Parameters like `llm_gen_temperature` and `llm_gen_max_new_tokens` are defined in `LanguageModelHandlerArguments` and apply universally, whether you run a local model or configure HF Inference Providers as LLM backends. The pipeline abstracts the execution differences while preserving the generation interface.

### How does streaming work with HF Inference endpoints?

The `BaseOpenAICompatibleLanguageModel` parses the HF Inference streaming response format and normalizes it to match OpenAI's server-sent events structure. Downstream components in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) consume these streams identically to local model outputs, enabling real-time speech-to-speech latency without code changes.