# FreeLLMAPI response_format Options: Supported Values for Audio, Images, and Speech

> Explore FreeLLMAPI response_format options for audio, image, and speech generation. Discover supported values like json, text, url, and b64_json for seamless integration.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: api-reference
- Published: 2026-09-01

---

**FreeLLMAPI supports `json`, `text`, `verbose_json`, and `vtt` for audio transcriptions, `url` and `b64_json` for image generations, and relies on HTTP `Content-Type` headers rather than a `response_format` parameter for text-to-speech audio delivery.**

FreeLLMAPI (tashfeenahmed/freellmapi) is an OpenAI-compatible proxy that forwards the `response_format` parameter to underlying AI providers, though the supported values vary significantly by endpoint. Understanding these **FreeLLMAPI response_format options** ensures your application receives data in the expected structure, whether processing Whisper transcriptions, DALL-E images, or TTS audio streams.

## Audio Transcription response_format Values

The `/v1/audio/transcriptions` endpoint accepts four distinct output formats defined in the `TRANSCRIPTION_FORMATS` set within [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts).

### Supported Formats for Whisper Endpoints

- **json**: Returns a JSON object containing the transcription text.
- **verbose_json**: Returns detailed JSON including segment-level timestamps and confidence scores.
- **text**: Returns plain text output (internally normalized to JSON for downstream consistency).
- **vtt**: Returns WebVTT subtitle format for video captioning.

### Validation in proxy.ts

According to the FreeLLMAPI source code, the `TRANSCRIPTION_FORMATS` set in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) validates incoming requests. If you specify an unsupported format like `srt`, the proxy returns a **400 Bad Request** with an informative error message before reaching the provider.

### Provider Chain Handling in media.ts

In [`server/src/services/media.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/media.ts), the `runTranscription` function and `resolveTranscriptionChain` logic forward the `responseFormat` string to each provider adapter. Notably, Custom and Groq backends only accept `json` or `verbose_json`; the internal code normalizes `text` requests to the JSON shape to maintain consistent fail-over behavior across the provider chain.

## Image Generation response_format Values

For the `/v1/images/generations` endpoint, FreeLLMAPI supports two mutually exclusive output formats documented in [`server/src/docs/openapi.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts).

### URL vs Base64 Output

- **url**: Returns a publicly accessible URL to the generated image (default behavior).
- **b64_json**: Returns the image as a Base64-encoded JSON string embedded in the response, useful for applications requiring offline data persistence.

### Sampling Parameter Schema

The `response_format` field is defined in the OpenAPI schema within [`server/src/docs/openapi.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts) and processed through [`server/src/lib/sampling-params.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/sampling-params.ts). This generic sampling-parameter handling forwards the format instruction to provider adapters (OpenAI, Google, etc.) without additional transformation.

## Text-to-Speech Format Specification

The `/v1/audio/speech` endpoint operates differently than other media endpoints in the FreeLLMAPI implementation.

### Content-Type Header Negotiation

Unlike transcription or image endpoints, the text-to-speech route **does not expose a `response_format` parameter** in the request body. Instead, the client selects the desired format through HTTP headers, and the server returns raw audio bytes with the corresponding `Content-Type`. Supported MIME types include `audio/mpeg` (MP3), `audio/wav`, `audio/ogg`, and `audio/webm`, as documented in [`server/src/docs/openapi.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts).

## Practical Code Examples

The following examples demonstrate how to structure requests for each endpoint using the supported **response_format** options.

### Requesting JSON Transcription Output

```typescript
await fetch('https://api.example.com/v1/audio/transcriptions', {
  method: 'POST',
  headers: { Authorization: `Bearer ${myKey}` },
  body: new FormData([
    ['file', audioBlob],
    ['model', 'auto'],
    ['response_format', 'json'],           // ← one of json | text | verbose_json | vtt
  ]),
}).then(r => r.json());

```

### Requesting VTT Subtitle Format

```typescript
await fetch('https://api.example.com/v1/audio/transcriptions', {
  method: 'POST',
  headers: { Authorization: `Bearer ${myKey}` },
  body: new FormData([
    ['file', audioBlob],
    ['model', '@cf/openai/whisper'],
    ['response_format', 'vtt'],
  ]),
}).then(r => r.text());  // VTT subtitle text

```

### Requesting Base64 Image Data

```typescript
await fetch('https://api.example.com/v1/images/generations', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${myKey}`,
  },
  body: JSON.stringify({
    prompt: 'A futuristic cityscape at sunset',
    n: 1,
    size: '1024x1024',
    response_format: 'b64_json',           // ← url | b64_json
  }),
}).then(r => r.json());

```

### Requesting MP3 Audio Bytes

```typescript
await fetch('https://api.example.com/v1/audio/speech', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${myKey}`,
  },
  body: JSON.stringify({
    model: 'tts-1',
    input: 'Hello, Free LLMAPI!',
    voice: 'alloy',
    response_format: 'mp3',               // not a request field; the server returns `audio/mpeg`
  }),
}).then(r => r.arrayBuffer());   // raw MP3 bytes

```

## Key Implementation Files

The **FreeLLMAPI response_format** behavior is governed by these specific source files:

- **[`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts)**: Contains the `TRANSCRIPTION_FORMATS` set and validation logic that rejects unsupported transcription formats with 400 errors.
- **[`server/src/services/media.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/media.ts)**: Implements the `TranscriptionParams` interface and `resolveTranscriptionChain` function handling provider-specific format constraints.
- **[`server/src/lib/sampling-params.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/sampling-params.ts)**: Defines the generic `response_format` schema used for image generation parameter forwarding.
- **[`server/src/docs/openapi.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts)**: Documents the OpenAPI component definitions for `ImageRequest`, `SpeechRequest`, and `TranscriptionRequest`, including the response `Content-Type` mappings.

## Summary

- **FreeLLMAPI** supports `json`, `text`, `verbose_json`, and `vtt` for audio transcriptions, rejecting legacy formats like `srt` at the proxy level.
- Image generation accepts `url` or `b64_json` to control whether you receive a hosted link or embedded Base64 data.
- Text-to-speech does not use a `response_format` parameter; instead, specify your desired audio codec through standard HTTP `Accept` headers or receive the default `audio/mpeg` stream.
- All format validations occur in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) for audio and [`server/src/docs/openapi.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts) for image schema definitions.

## Frequently Asked Questions

### What response_format should I use for Whisper transcriptions in FreeLLMAPI?

Use `json` for standard text output, `verbose_json` for timestamp and probability metadata, `text` for plain string responses, or `vtt` for WebVTT subtitle files. Avoid `srt` as the `TRANSCRIPTION_FORMATS` validator in [`proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/proxy.ts) explicitly rejects this legacy format.

### How do I retrieve Base64-encoded images instead of URLs?

Set `response_format: 'b64_json'` in your request body to the `/v1/images/generations` endpoint. This instructs FreeLLMAPI to return the image data as a Base64 string within the JSON response rather than a hosted URL, as defined in the OpenAPI schema in [`server/src/docs/openapi.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts).

### Why does my text-to-speech request ignore the response_format parameter?

The `/v1/audio/speech` endpoint does not accept a `response_format` field in the request body. According to the implementation in [`server/src/docs/openapi.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts), you must rely on the HTTP response `Content-Type` header (e.g., `audio/mpeg`, `audio/wav`) to determine the audio codec, with MP3 being the default return format.

### Which transcription formats work with all providers in FreeLLMAPI?

Only `json` and `verbose_json` are universally supported across all backend providers including Custom and Groq adapters. The `text` format is normalized to JSON internally by the `runTranscription` function in [`server/src/services/media.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/media.ts) to ensure consistent fail-over behavior, while `vtt` availability depends on the specific upstream provider capabilities.