How to Use the "auto" Model in FreeLLMAPI for Request Routing
Using model: "auto" in FreeLLMAPI routes requests through a dynamic fallback chain instead of a specific provider model, with optional suffixes like :fast or :coding to apply strategies or select named profiles.
The tashfeenahmed/freellmapi repository provides a unified gateway for large language models with an intelligent routing layer. By specifying the "auto" model in your API requests, you delegate model selection to the server's routing logic, which can balance cost, latency, and reliability across multiple providers without changing your client code.
Understanding the "auto" Routing Mechanism
When you send a request with model: "auto", the server intercepts this virtual identifier in server/src/services/router.ts and activates the fallback chain routing system rather than mapping directly to a provider-specific model.
Default Fallback Chain
Specifying "auto" (or omitting the model field entirely) triggers the active fallback chain—the set of models currently enabled on the operator dashboard. The router attempts the request against each model in this chain sequentially until one succeeds, providing automatic failover without client-side retry logic.
According to the source code in server/src/services/router.ts, this behavior is handled within the request preprocessing logic that resolves model strings before forwarding to upstream providers.
Named Profiles and Custom Chains
You can target a specific named fallback chain (profile) by appending a suffix: auto:profilename. For example, auto:coding routes through a curated chain optimized for code generation tasks.
In server/src/services/profile-models.ts, profiles are managed with a flag called auto_include_new_models. When set to 1, newly discovered models are automatically appended to the chain; when 0, the chain remains exactly as configured by the operator.
Routing Strategies for "auto" Requests
FreeLLMAPI supports strategy-based routing when you need to optimize for specific performance characteristics rather than following a static chain order. These strategies reorder all enabled models dynamically.
The following suffixes are supported in server/src/services/router.ts:
auto:fast– Prioritizes models with lowest latencyauto:smart– Optimizes for reasoning quality and context handlingauto:cheap– Routes to the lowest-cost available option firstauto:reliable– Prioritizes providers with highest uptime metricsauto:balanced– Weighs cost, speed, and quality evenly
When a strategy suffix is detected, the router ignores the static chain order and applies the bandit-scoring algorithm (documented in docs/architecture/01-routing-and-bandit-scoring.md) to rank available models.
Configuration and Error Handling
Key Selection Strategy
The router also manages API key selection when multiple keys exist for the same provider. The key_selection_strategy parameter accepts auto (default) or least-remaining. The auto strategy selects keys based on current rate-limit health, while least-remaining targets keys with the lowest remaining quota. This logic resides in server/src/services/router.ts alongside the model routing code.
Invalid Profile Handling
Supplying an unknown profile name (e.g., auto:unknownprofile) results in a 400 Bad Request error with a clear error message, rather than silently falling back to the active chain. This validation occurs early in the request lifecycle within the router service.
Cross-Endpoint Compatibility
The "auto" model works uniformly across all API modalities:
- Chat Completions (
/v1/chat/completions) - Legacy Completions (
/v1/completions) - Embeddings (
/v1/embeddings) – handled inserver/src/services/embeddings.ts - Vision/Image endpoints
Every response includes the X-Routed-Via header indicating the actual <provider>/<model> combination that served the request, enabling observability even when using automatic routing.
Implementation Examples
Basic Auto Routing Request
curl -X POST https://api.freellmapi.local/v1/chat/completions \
-H "Authorization: Bearer $FREELLMAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Explain the law of thermodynamics."}]
}'
Strategy-Based Routing for Low Latency
curl -X POST https://api.freellmapi.local/v1/chat/completions \
-H "Authorization: Bearer $FREELLMAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto:fast",
"messages": [{"role": "user", "content": "Write a quick hello world in Python."}]
}'
Using Named Profiles
curl -X POST https://api.freellmapi.local/v1/chat/completions \
-H "Authorization: Bearer $FREELLMAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto:coding",
"messages": [{"role": "user", "content": "Generate a React component."}]
}'
Embeddings with Auto Selection
curl -X POST https://api.freellmapi.local/v1/embeddings \
-H "Authorization: Bearer $FREELLMAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"input": "FreeLLMAPI provides unified LLM access."
}'
Summary
model: "auto"triggers fallback chain routing through enabled models instead of targeting a specific provider.- Named profiles use the syntax
auto:profilenameto select preconfigured chains managed inserver/src/services/profile-models.ts. - Strategies like
auto:fastorauto:cheapdynamically reorder models based on performance metrics. - The
auto_include_new_modelsflag controls whether new provider models automatically join existing chains. - All endpoints including chat, completions, and embeddings support the auto model with consistent
X-Routed-Viaresponse headers.
Frequently Asked Questions
What is the difference between "auto" and "auto:fast" in FreeLLMAPI?
Using "auto" routes your request through the active fallback chain in the order defined by the operator dashboard. Using "auto:fast" applies the fast routing strategy, which ignores the dashboard order and instead prioritizes all enabled models by lowest latency first. Both options are parsed in server/src/services/router.ts.
How do named profiles work with the auto model?
Named profiles allow you to define specific fallback chains for different use cases (e.g., auto:coding for programming tasks). When you specify auto:profilename, the router looks up that profile's model list and attempts them sequentially. If the profile does not exist, the API returns a 400 error rather than defaulting to the active chain.
Can I use the auto model for embedding requests?
Yes. The auto routing system is model-agnostic and works across chat completions, legacy completions, and embeddings endpoints. The server/src/services/embeddings.ts file handles auto selection specifically for embedding models, applying the same fallback chain logic used for text generation.
What happens when a new model is added to the provider list?
If the active profile has auto_include_new_models set to 1 (managed in server/src/services/profile-models.ts), newly discovered models are automatically appended to the fallback chain. If set to 0, the chain remains static and ignores new models until an operator manually updates the configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →