FreeLLMAPI response_format Options: Supported Values for Audio, Images, and Speech
FreeLLMAPI supports json, text, verbose_json, and vtt for audio transcriptions, url and b64_json for image generations, and relies on HTTP Content-Type headers rather than a response_format parameter for text-to-speech audio delivery.
FreeLLMAPI (tashfeenahmed/freellmapi) is an OpenAI-compatible proxy that forwards the response_format parameter to underlying AI providers, though the supported values vary significantly by endpoint. Understanding these FreeLLMAPI response_format options ensures your application receives data in the expected structure, whether processing Whisper transcriptions, DALL-E images, or TTS audio streams.
Audio Transcription response_format Values
The /v1/audio/transcriptions endpoint accepts four distinct output formats defined in the TRANSCRIPTION_FORMATS set within server/src/routes/proxy.ts.
Supported Formats for Whisper Endpoints
- json: Returns a JSON object containing the transcription text.
- verbose_json: Returns detailed JSON including segment-level timestamps and confidence scores.
- text: Returns plain text output (internally normalized to JSON for downstream consistency).
- vtt: Returns WebVTT subtitle format for video captioning.
Validation in proxy.ts
According to the FreeLLMAPI source code, the TRANSCRIPTION_FORMATS set in server/src/routes/proxy.ts validates incoming requests. If you specify an unsupported format like srt, the proxy returns a 400 Bad Request with an informative error message before reaching the provider.
Provider Chain Handling in media.ts
In server/src/services/media.ts, the runTranscription function and resolveTranscriptionChain logic forward the responseFormat string to each provider adapter. Notably, Custom and Groq backends only accept json or verbose_json; the internal code normalizes text requests to the JSON shape to maintain consistent fail-over behavior across the provider chain.
Image Generation response_format Values
For the /v1/images/generations endpoint, FreeLLMAPI supports two mutually exclusive output formats documented in server/src/docs/openapi.ts.
URL vs Base64 Output
- url: Returns a publicly accessible URL to the generated image (default behavior).
- b64_json: Returns the image as a Base64-encoded JSON string embedded in the response, useful for applications requiring offline data persistence.
Sampling Parameter Schema
The response_format field is defined in the OpenAPI schema within server/src/docs/openapi.ts and processed through server/src/lib/sampling-params.ts. This generic sampling-parameter handling forwards the format instruction to provider adapters (OpenAI, Google, etc.) without additional transformation.
Text-to-Speech Format Specification
The /v1/audio/speech endpoint operates differently than other media endpoints in the FreeLLMAPI implementation.
Content-Type Header Negotiation
Unlike transcription or image endpoints, the text-to-speech route does not expose a response_format parameter in the request body. Instead, the client selects the desired format through HTTP headers, and the server returns raw audio bytes with the corresponding Content-Type. Supported MIME types include audio/mpeg (MP3), audio/wav, audio/ogg, and audio/webm, as documented in server/src/docs/openapi.ts.
Practical Code Examples
The following examples demonstrate how to structure requests for each endpoint using the supported response_format options.
Requesting JSON Transcription Output
await fetch('https://api.example.com/v1/audio/transcriptions', {
method: 'POST',
headers: { Authorization: `Bearer ${myKey}` },
body: new FormData([
['file', audioBlob],
['model', 'auto'],
['response_format', 'json'], // ← one of json | text | verbose_json | vtt
]),
}).then(r => r.json());
Requesting VTT Subtitle Format
await fetch('https://api.example.com/v1/audio/transcriptions', {
method: 'POST',
headers: { Authorization: `Bearer ${myKey}` },
body: new FormData([
['file', audioBlob],
['model', '@cf/openai/whisper'],
['response_format', 'vtt'],
]),
}).then(r => r.text()); // VTT subtitle text
Requesting Base64 Image Data
await fetch('https://api.example.com/v1/images/generations', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${myKey}`,
},
body: JSON.stringify({
prompt: 'A futuristic cityscape at sunset',
n: 1,
size: '1024x1024',
response_format: 'b64_json', // ← url | b64_json
}),
}).then(r => r.json());
Requesting MP3 Audio Bytes
await fetch('https://api.example.com/v1/audio/speech', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${myKey}`,
},
body: JSON.stringify({
model: 'tts-1',
input: 'Hello, Free LLMAPI!',
voice: 'alloy',
response_format: 'mp3', // not a request field; the server returns `audio/mpeg`
}),
}).then(r => r.arrayBuffer()); // raw MP3 bytes
Key Implementation Files
The FreeLLMAPI response_format behavior is governed by these specific source files:
server/src/routes/proxy.ts: Contains theTRANSCRIPTION_FORMATSset and validation logic that rejects unsupported transcription formats with 400 errors.server/src/services/media.ts: Implements theTranscriptionParamsinterface andresolveTranscriptionChainfunction handling provider-specific format constraints.server/src/lib/sampling-params.ts: Defines the genericresponse_formatschema used for image generation parameter forwarding.server/src/docs/openapi.ts: Documents the OpenAPI component definitions forImageRequest,SpeechRequest, andTranscriptionRequest, including the responseContent-Typemappings.
Summary
- FreeLLMAPI supports
json,text,verbose_json, andvttfor audio transcriptions, rejecting legacy formats likesrtat the proxy level. - Image generation accepts
urlorb64_jsonto control whether you receive a hosted link or embedded Base64 data. - Text-to-speech does not use a
response_formatparameter; instead, specify your desired audio codec through standard HTTPAcceptheaders or receive the defaultaudio/mpegstream. - All format validations occur in
server/src/routes/proxy.tsfor audio andserver/src/docs/openapi.tsfor image schema definitions.
Frequently Asked Questions
What response_format should I use for Whisper transcriptions in FreeLLMAPI?
Use json for standard text output, verbose_json for timestamp and probability metadata, text for plain string responses, or vtt for WebVTT subtitle files. Avoid srt as the TRANSCRIPTION_FORMATS validator in proxy.ts explicitly rejects this legacy format.
How do I retrieve Base64-encoded images instead of URLs?
Set response_format: 'b64_json' in your request body to the /v1/images/generations endpoint. This instructs FreeLLMAPI to return the image data as a Base64 string within the JSON response rather than a hosted URL, as defined in the OpenAPI schema in server/src/docs/openapi.ts.
Why does my text-to-speech request ignore the response_format parameter?
The /v1/audio/speech endpoint does not accept a response_format field in the request body. According to the implementation in server/src/docs/openapi.ts, you must rely on the HTTP response Content-Type header (e.g., audio/mpeg, audio/wav) to determine the audio codec, with MP3 being the default return format.
Which transcription formats work with all providers in FreeLLMAPI?
Only json and verbose_json are universally supported across all backend providers including Custom and Groq adapters. The text format is normalized to JSON internally by the runTranscription function in server/src/services/media.ts to ensure consistent fail-over behavior, while vtt availability depends on the specific upstream provider capabilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →