How to Use Custom Local LLM Endpoints with FreeLLMAPI
FreeLLMAPI supports any OpenAI-compatible local server as a custom endpoint, routing requests through your unified API key while maintaining automatic fail-over and visibility via response headers.
FreeLLMAPI is an open-source proxy layer that unifies access to multiple large language model providers. When you need to integrate self-hosted models running on your own infrastructure, the platform treats these as custom local LLM endpoints, enabling you to leverage the same routing, quota management, and analytics that cloud providers receive.
Architecture of Custom Endpoints
FreeLLMAPI treats local OpenAI-compatible servers—such as llama.cpp, LM Studio, vLLM (with auth disabled), or any self-hosted model exposing standard /v1 endpoints—as first-class providers. The system stores these configurations in the api_keys table with platform = 'custom', allowing the proxy to forward requests to your hardware while maintaining the unified API contract.
Endpoint Registration Logic
In server/src/services/custom-endpoint.ts, the registration system uniquely identifies each custom endpoint by its base_url rather than by generated credentials. This design allows you to rotate secrets without changing the endpoint URL itself【custom-endpoint.ts†L5-L12】.
When you register a local server without authentication, the system stores a placeholder key no-key in the database. If you later add a secret, the service updates the existing row rather than creating a duplicate entry, preventing credential sprawl【custom-endpoint.ts†L14-L22】【custom-endpoint.ts†L84-L92】.
Routing and Fail-Over Integration
The proxy's router (server/src/services/router.ts) resolves incoming model strings to concrete providers. When you specify model="auto" (or any auto:* profile), the router includes enabled custom endpoints in the fallback chain, selecting providers based on request constraints such as vision capability or token limits【router.ts†(routing‑logic)】.
If your local endpoint becomes unavailable, the router automatically fails over to the next enabled model—whether another local instance or a cloud provider—without requiring client-side changes【router.ts†(fallback‑logic)】. Every response includes an X-Routed-Via header revealing the actual serving endpoint, formatted as custom/http://localhost:8000/v1【api.md†L61-L68】.
Configuring Your Local LLM Endpoint
Registering via the Dashboard
Navigate to Keys → Custom endpoint in the web interface. Enter the base URL of your local OpenAI-compatible server (e.g., http://localhost:8000/v1). The system immediately stores this endpoint and begins including it in routing decisions.
Using the CLI
The CLI tool in cli/src/tools.ts provides a terminal interface for endpoint management. To add an unauthenticated local server:
freellmapi keys add --platform custom \
--base-url http://localhost:8000/v1 \
--label "my-local-llama"
For servers requiring authentication, include the --secret flag:
freellmapi keys add --platform custom \
--base-url http://localhost:8000/v1 \
--secret "my-local-api-key" \
--label "my-local-llama-auth"
Client Integration Patterns
Python OpenAI SDK
Use your existing OpenAI-compatible client code. Point the base URL to your FreeLLMAPI proxy and set model="auto" to allow routing to local endpoints:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:3001/v1", # FreeLLMAPI proxy
api_key="freellmapi-your-unified-key", # unified dashboard key
)
resp = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Explain the difference between RAG and fine-tuning."}],
)
print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via")) # e.g. custom/http://localhost:8000/v1
Direct cURL Access
To bypass the proxy and communicate directly with your local server:
curl http://localhost:8000/v1/chat/completions \
-H "Authorization: Bearer <any-key-or-no-key>" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role":"user","content":"Translate \"hola\" to French."}]
}'
Handling Authenticated Local Servers
When your local LLM requires an API key, FreeLLMAPI forwards the stored secret in the Authorization header. Register the endpoint with the --secret parameter as shown above; the proxy handles credential injection transparently, keeping your local keys private from end clients.
Monitoring and Debugging
Inspect the X-Routed-Via response header to verify which endpoint served a specific request. This header reports the provider type and URL (e.g., custom/http://localhost:8000/v1), enabling you to confirm that traffic flows to your local infrastructure during testing or production incidents.
Summary
- FreeLLMAPI treats OpenAI-compatible local servers as first-class providers via
platform = 'custom'entries in theapi_keystable. - The router in
server/src/services/router.tsincludes custom endpoints in automatic fail-over chains when usingmodel="auto". - Registration requires only a
base_url; optional secrets update existing entries rather than creating duplicates. - Response headers reveal the actual serving endpoint via
X-Routed-Via, enabling debugging of routing decisions. - No client code changes are required—use your existing OpenAI-compatible SDKs with the unified FreeLLMAPI base URL.
Frequently Asked Questions
Which local LLM servers are compatible with FreeLLMAPI?
Any server implementing the OpenAI-compatible /v1 REST interface works immediately. This includes llama.cpp (with server mode), LM Studio, vLLM, Ollama (with OpenAI compatibility layer), and custom Python servers using the OpenAI SDK. The only requirement is standard chat completions endpoint support at the base URL you provide.
How does FreeLLMAPI handle authentication for local endpoints?
The system stores optional secrets alongside the base_url in the api_keys table. When a request routes to an authenticated local endpoint, the proxy injects the stored secret into the Authorization header. If you initially register without a secret, the placeholder no-key value is used, and you can update the credential later without reregistering the URL【custom-endpoint.ts†L14-L22】.
Can I use custom endpoints alongside cloud providers in the same request?
Yes. When you set model="auto", the router evaluates all enabled endpoints—including local custom endpoints and cloud providers—against your request constraints. If your local server lacks capacity or returns an error, the proxy automatically fails over to the next available provider in the chain【router.ts†(fallback‑logic)】.
How do I verify that requests are actually routing to my local server?
Check the X-Routed-Via response header. This header contains the provider identifier and base URL (e.g., custom/http://localhost:8000/v1) of the endpoint that generated the response【api.md†L61-L68】. During development, you can also temporarily shut down your local server to trigger fail-over and confirm the behavior through this header's changing values.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →