How to Use the "auto" Model in FreeLLMAPI for Request Routing

Using model: "auto" in FreeLLMAPI routes requests through a dynamic fallback chain instead of a specific provider model, with optional suffixes like :fast or :coding to apply strategies or select named profiles.

The tashfeenahmed/freellmapi repository provides a unified gateway for large language models with an intelligent routing layer. By specifying the "auto" model in your API requests, you delegate model selection to the server's routing logic, which can balance cost, latency, and reliability across multiple providers without changing your client code.

Understanding the "auto" Routing Mechanism

When you send a request with model: "auto", the server intercepts this virtual identifier in server/src/services/router.ts and activates the fallback chain routing system rather than mapping directly to a provider-specific model.

Default Fallback Chain

Specifying "auto" (or omitting the model field entirely) triggers the active fallback chain—the set of models currently enabled on the operator dashboard. The router attempts the request against each model in this chain sequentially until one succeeds, providing automatic failover without client-side retry logic.

According to the source code in server/src/services/router.ts, this behavior is handled within the request preprocessing logic that resolves model strings before forwarding to upstream providers.

Named Profiles and Custom Chains

You can target a specific named fallback chain (profile) by appending a suffix: auto:profilename. For example, auto:coding routes through a curated chain optimized for code generation tasks.

In server/src/services/profile-models.ts, profiles are managed with a flag called auto_include_new_models. When set to 1, newly discovered models are automatically appended to the chain; when 0, the chain remains exactly as configured by the operator.

Routing Strategies for "auto" Requests

FreeLLMAPI supports strategy-based routing when you need to optimize for specific performance characteristics rather than following a static chain order. These strategies reorder all enabled models dynamically.

The following suffixes are supported in server/src/services/router.ts:

  • auto:fast – Prioritizes models with lowest latency
  • auto:smart – Optimizes for reasoning quality and context handling
  • auto:cheap – Routes to the lowest-cost available option first
  • auto:reliable – Prioritizes providers with highest uptime metrics
  • auto:balanced – Weighs cost, speed, and quality evenly

When a strategy suffix is detected, the router ignores the static chain order and applies the bandit-scoring algorithm (documented in docs/architecture/01-routing-and-bandit-scoring.md) to rank available models.

Configuration and Error Handling

Key Selection Strategy

The router also manages API key selection when multiple keys exist for the same provider. The key_selection_strategy parameter accepts auto (default) or least-remaining. The auto strategy selects keys based on current rate-limit health, while least-remaining targets keys with the lowest remaining quota. This logic resides in server/src/services/router.ts alongside the model routing code.

Invalid Profile Handling

Supplying an unknown profile name (e.g., auto:unknownprofile) results in a 400 Bad Request error with a clear error message, rather than silently falling back to the active chain. This validation occurs early in the request lifecycle within the router service.

Cross-Endpoint Compatibility

The "auto" model works uniformly across all API modalities:

  • Chat Completions (/v1/chat/completions)
  • Legacy Completions (/v1/completions)
  • Embeddings (/v1/embeddings) – handled in server/src/services/embeddings.ts
  • Vision/Image endpoints

Every response includes the X-Routed-Via header indicating the actual <provider>/<model> combination that served the request, enabling observability even when using automatic routing.

Implementation Examples

Basic Auto Routing Request

curl -X POST https://api.freellmapi.local/v1/chat/completions \
  -H "Authorization: Bearer $FREELLMAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "messages": [{"role": "user", "content": "Explain the law of thermodynamics."}]
  }'

Strategy-Based Routing for Low Latency

curl -X POST https://api.freellmapi.local/v1/chat/completions \
  -H "Authorization: Bearer $FREELLMAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto:fast",
    "messages": [{"role": "user", "content": "Write a quick hello world in Python."}]
  }'

Using Named Profiles

curl -X POST https://api.freellmapi.local/v1/chat/completions \
  -H "Authorization: Bearer $FREELLMAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto:coding",
    "messages": [{"role": "user", "content": "Generate a React component."}]
  }'

Embeddings with Auto Selection

curl -X POST https://api.freellmapi.local/v1/embeddings \
  -H "Authorization: Bearer $FREELLMAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "input": "FreeLLMAPI provides unified LLM access."
  }'

Summary

  • model: "auto" triggers fallback chain routing through enabled models instead of targeting a specific provider.
  • Named profiles use the syntax auto:profilename to select preconfigured chains managed in server/src/services/profile-models.ts.
  • Strategies like auto:fast or auto:cheap dynamically reorder models based on performance metrics.
  • The auto_include_new_models flag controls whether new provider models automatically join existing chains.
  • All endpoints including chat, completions, and embeddings support the auto model with consistent X-Routed-Via response headers.

Frequently Asked Questions

What is the difference between "auto" and "auto:fast" in FreeLLMAPI?

Using "auto" routes your request through the active fallback chain in the order defined by the operator dashboard. Using "auto:fast" applies the fast routing strategy, which ignores the dashboard order and instead prioritizes all enabled models by lowest latency first. Both options are parsed in server/src/services/router.ts.

How do named profiles work with the auto model?

Named profiles allow you to define specific fallback chains for different use cases (e.g., auto:coding for programming tasks). When you specify auto:profilename, the router looks up that profile's model list and attempts them sequentially. If the profile does not exist, the API returns a 400 error rather than defaulting to the active chain.

Can I use the auto model for embedding requests?

Yes. The auto routing system is model-agnostic and works across chat completions, legacy completions, and embeddings endpoints. The server/src/services/embeddings.ts file handles auto selection specifically for embedding models, applying the same fallback chain logic used for text generation.

What happens when a new model is added to the provider list?

If the active profile has auto_include_new_models set to 1 (managed in server/src/services/profile-models.ts), newly discovered models are automatically appended to the fallback chain. If set to 0, the chain remains static and ignores new models until an operator manually updates the configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →