FreeLLMAPI response_format Options: Supported Values for Audio, Images, and Speech

FreeLLMAPI supports json, text, verbose_json, and vtt for audio transcriptions, url and b64_json for image generations, and relies on HTTP Content-Type headers rather than a response_format parameter for text-to-speech audio delivery.

FreeLLMAPI (tashfeenahmed/freellmapi) is an OpenAI-compatible proxy that forwards the response_format parameter to underlying AI providers, though the supported values vary significantly by endpoint. Understanding these FreeLLMAPI response_format options ensures your application receives data in the expected structure, whether processing Whisper transcriptions, DALL-E images, or TTS audio streams.

Audio Transcription response_format Values

The /v1/audio/transcriptions endpoint accepts four distinct output formats defined in the TRANSCRIPTION_FORMATS set within server/src/routes/proxy.ts.

Supported Formats for Whisper Endpoints

  • json: Returns a JSON object containing the transcription text.
  • verbose_json: Returns detailed JSON including segment-level timestamps and confidence scores.
  • text: Returns plain text output (internally normalized to JSON for downstream consistency).
  • vtt: Returns WebVTT subtitle format for video captioning.

Validation in proxy.ts

According to the FreeLLMAPI source code, the TRANSCRIPTION_FORMATS set in server/src/routes/proxy.ts validates incoming requests. If you specify an unsupported format like srt, the proxy returns a 400 Bad Request with an informative error message before reaching the provider.

Provider Chain Handling in media.ts

In server/src/services/media.ts, the runTranscription function and resolveTranscriptionChain logic forward the responseFormat string to each provider adapter. Notably, Custom and Groq backends only accept json or verbose_json; the internal code normalizes text requests to the JSON shape to maintain consistent fail-over behavior across the provider chain.

Image Generation response_format Values

For the /v1/images/generations endpoint, FreeLLMAPI supports two mutually exclusive output formats documented in server/src/docs/openapi.ts.

URL vs Base64 Output

  • url: Returns a publicly accessible URL to the generated image (default behavior).
  • b64_json: Returns the image as a Base64-encoded JSON string embedded in the response, useful for applications requiring offline data persistence.

Sampling Parameter Schema

The response_format field is defined in the OpenAPI schema within server/src/docs/openapi.ts and processed through server/src/lib/sampling-params.ts. This generic sampling-parameter handling forwards the format instruction to provider adapters (OpenAI, Google, etc.) without additional transformation.

Text-to-Speech Format Specification

The /v1/audio/speech endpoint operates differently than other media endpoints in the FreeLLMAPI implementation.

Content-Type Header Negotiation

Unlike transcription or image endpoints, the text-to-speech route does not expose a response_format parameter in the request body. Instead, the client selects the desired format through HTTP headers, and the server returns raw audio bytes with the corresponding Content-Type. Supported MIME types include audio/mpeg (MP3), audio/wav, audio/ogg, and audio/webm, as documented in server/src/docs/openapi.ts.

Practical Code Examples

The following examples demonstrate how to structure requests for each endpoint using the supported response_format options.

Requesting JSON Transcription Output

await fetch('https://api.example.com/v1/audio/transcriptions', {
  method: 'POST',
  headers: { Authorization: `Bearer ${myKey}` },
  body: new FormData([
    ['file', audioBlob],
    ['model', 'auto'],
    ['response_format', 'json'],           // ← one of json | text | verbose_json | vtt
  ]),
}).then(r => r.json());

Requesting VTT Subtitle Format

await fetch('https://api.example.com/v1/audio/transcriptions', {
  method: 'POST',
  headers: { Authorization: `Bearer ${myKey}` },
  body: new FormData([
    ['file', audioBlob],
    ['model', '@cf/openai/whisper'],
    ['response_format', 'vtt'],
  ]),
}).then(r => r.text());  // VTT subtitle text

Requesting Base64 Image Data

await fetch('https://api.example.com/v1/images/generations', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${myKey}`,
  },
  body: JSON.stringify({
    prompt: 'A futuristic cityscape at sunset',
    n: 1,
    size: '1024x1024',
    response_format: 'b64_json',           // ← url | b64_json
  }),
}).then(r => r.json());

Requesting MP3 Audio Bytes

await fetch('https://api.example.com/v1/audio/speech', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${myKey}`,
  },
  body: JSON.stringify({
    model: 'tts-1',
    input: 'Hello, Free LLMAPI!',
    voice: 'alloy',
    response_format: 'mp3',               // not a request field; the server returns `audio/mpeg`
  }),
}).then(r => r.arrayBuffer());   // raw MP3 bytes

Key Implementation Files

The FreeLLMAPI response_format behavior is governed by these specific source files:

  • server/src/routes/proxy.ts: Contains the TRANSCRIPTION_FORMATS set and validation logic that rejects unsupported transcription formats with 400 errors.
  • server/src/services/media.ts: Implements the TranscriptionParams interface and resolveTranscriptionChain function handling provider-specific format constraints.
  • server/src/lib/sampling-params.ts: Defines the generic response_format schema used for image generation parameter forwarding.
  • server/src/docs/openapi.ts: Documents the OpenAPI component definitions for ImageRequest, SpeechRequest, and TranscriptionRequest, including the response Content-Type mappings.

Summary

  • FreeLLMAPI supports json, text, verbose_json, and vtt for audio transcriptions, rejecting legacy formats like srt at the proxy level.
  • Image generation accepts url or b64_json to control whether you receive a hosted link or embedded Base64 data.
  • Text-to-speech does not use a response_format parameter; instead, specify your desired audio codec through standard HTTP Accept headers or receive the default audio/mpeg stream.
  • All format validations occur in server/src/routes/proxy.ts for audio and server/src/docs/openapi.ts for image schema definitions.

Frequently Asked Questions

What response_format should I use for Whisper transcriptions in FreeLLMAPI?

Use json for standard text output, verbose_json for timestamp and probability metadata, text for plain string responses, or vtt for WebVTT subtitle files. Avoid srt as the TRANSCRIPTION_FORMATS validator in proxy.ts explicitly rejects this legacy format.

How do I retrieve Base64-encoded images instead of URLs?

Set response_format: 'b64_json' in your request body to the /v1/images/generations endpoint. This instructs FreeLLMAPI to return the image data as a Base64 string within the JSON response rather than a hosted URL, as defined in the OpenAPI schema in server/src/docs/openapi.ts.

Why does my text-to-speech request ignore the response_format parameter?

The /v1/audio/speech endpoint does not accept a response_format field in the request body. According to the implementation in server/src/docs/openapi.ts, you must rely on the HTTP response Content-Type header (e.g., audio/mpeg, audio/wav) to determine the audio codec, with MP3 being the default return format.

Which transcription formats work with all providers in FreeLLMAPI?

Only json and verbose_json are universally supported across all backend providers including Custom and Groq adapters. The text format is normalized to JSON internally by the runTranscription function in server/src/services/media.ts to ensure consistent fail-over behavior, while vtt availability depends on the specific upstream provider capabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →