# How to Use Custom Local LLM Endpoints with FreeLLMAPI

> Connect custom local LLM endpoints to FreeLLMAPI. Use any OpenAI-compatible server with your unified API key for seamless fail-over and control.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-09-01

---

**FreeLLMAPI supports any OpenAI-compatible local server as a custom endpoint, routing requests through your unified API key while maintaining automatic fail-over and visibility via response headers.**

FreeLLMAPI is an open-source proxy layer that unifies access to multiple large language model providers. When you need to integrate self-hosted models running on your own infrastructure, the platform treats these as **custom local LLM endpoints**, enabling you to leverage the same routing, quota management, and analytics that cloud providers receive.

## Architecture of Custom Endpoints

FreeLLMAPI treats local OpenAI-compatible servers—such as **llama.cpp**, **LM Studio**, **vLLM** (with auth disabled), or any self-hosted model exposing standard `/v1` endpoints—as first-class providers. The system stores these configurations in the `api_keys` table with `platform = 'custom'`, allowing the proxy to forward requests to your hardware while maintaining the unified API contract.

### Endpoint Registration Logic

In [`server/src/services/custom-endpoint.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/custom-endpoint.ts), the registration system uniquely identifies each custom endpoint by its **`base_url`** rather than by generated credentials. This design allows you to rotate secrets without changing the endpoint URL itself【custom-endpoint.ts†L5-L12】.

When you register a local server without authentication, the system stores a placeholder key `no-key` in the database. If you later add a secret, the service updates the existing row rather than creating a duplicate entry, preventing credential sprawl【custom-endpoint.ts†L14-L22】【custom-endpoint.ts†L84-L92】.

## Routing and Fail-Over Integration

The proxy's router ([`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)) resolves incoming `model` strings to concrete providers. When you specify `model="auto"` (or any `auto:*` profile), the router includes enabled custom endpoints in the fallback chain, selecting providers based on request constraints such as vision capability or token limits【router.ts†(routing‑logic)】.

If your local endpoint becomes unavailable, the router automatically fails over to the next enabled model—whether another local instance or a cloud provider—without requiring client-side changes【router.ts†(fallback‑logic)】. Every response includes an **`X-Routed-Via`** header revealing the actual serving endpoint, formatted as `custom/http://localhost:8000/v1`【api.md†L61-L68】.

## Configuring Your Local LLM Endpoint

### Registering via the Dashboard

Navigate to **Keys → Custom endpoint** in the web interface. Enter the base URL of your local OpenAI-compatible server (e.g., `http://localhost:8000/v1`). The system immediately stores this endpoint and begins including it in routing decisions.

### Using the CLI

The CLI tool in [`cli/src/tools.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/cli/src/tools.ts) provides a terminal interface for endpoint management. To add an unauthenticated local server:

```bash
freellmapi keys add --platform custom \
  --base-url http://localhost:8000/v1 \
  --label "my-local-llama"

```

For servers requiring authentication, include the `--secret` flag:

```bash
freellmapi keys add --platform custom \
  --base-url http://localhost:8000/v1 \
  --secret "my-local-api-key" \
  --label "my-local-llama-auth"

```

## Client Integration Patterns

### Python OpenAI SDK

Use your existing OpenAI-compatible client code. Point the base URL to your FreeLLMAPI proxy and set `model="auto"` to allow routing to local endpoints:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",          # FreeLLMAPI proxy

    api_key="freellmapi-your-unified-key",        # unified dashboard key

)

resp = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Explain the difference between RAG and fine-tuning."}],
)

print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via"))   # e.g. custom/http://localhost:8000/v1

```

### Direct cURL Access

To bypass the proxy and communicate directly with your local server:

```bash
curl http://localhost:8000/v1/chat/completions \
  -H "Authorization: Bearer <any-key-or-no-key>" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{"role":"user","content":"Translate \"hola\" to French."}]
      }'

```

### Handling Authenticated Local Servers

When your local LLM requires an API key, FreeLLMAPI forwards the stored secret in the `Authorization` header. Register the endpoint with the `--secret` parameter as shown above; the proxy handles credential injection transparently, keeping your local keys private from end clients.

## Monitoring and Debugging

Inspect the **`X-Routed-Via`** response header to verify which endpoint served a specific request. This header reports the provider type and URL (e.g., `custom/http://localhost:8000/v1`), enabling you to confirm that traffic flows to your local infrastructure during testing or production incidents.

## Summary

- FreeLLMAPI treats OpenAI-compatible local servers as first-class providers via `platform = 'custom'` entries in the `api_keys` table.
- The router in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) includes custom endpoints in automatic fail-over chains when using `model="auto"`.
- Registration requires only a `base_url`; optional secrets update existing entries rather than creating duplicates.
- Response headers reveal the actual serving endpoint via `X-Routed-Via`, enabling debugging of routing decisions.
- No client code changes are required—use your existing OpenAI-compatible SDKs with the unified FreeLLMAPI base URL.

## Frequently Asked Questions

### Which local LLM servers are compatible with FreeLLMAPI?

Any server implementing the OpenAI-compatible `/v1` REST interface works immediately. This includes **llama.cpp** (with server mode), **LM Studio**, **vLLM**, **Ollama** (with OpenAI compatibility layer), and custom Python servers using the OpenAI SDK. The only requirement is standard chat completions endpoint support at the base URL you provide.

### How does FreeLLMAPI handle authentication for local endpoints?

The system stores optional secrets alongside the `base_url` in the `api_keys` table. When a request routes to an authenticated local endpoint, the proxy injects the stored secret into the `Authorization` header. If you initially register without a secret, the placeholder `no-key` value is used, and you can update the credential later without reregistering the URL【custom-endpoint.ts†L14-L22】.

### Can I use custom endpoints alongside cloud providers in the same request?

Yes. When you set `model="auto"`, the router evaluates all enabled endpoints—including local custom endpoints and cloud providers—against your request constraints. If your local server lacks capacity or returns an error, the proxy automatically fails over to the next available provider in the chain【router.ts†(fallback‑logic)】.

### How do I verify that requests are actually routing to my local server?

Check the **`X-Routed-Via`** response header. This header contains the provider identifier and base URL (e.g., `custom/http://localhost:8000/v1`) of the endpoint that generated the response【api.md†L61-L68】. During development, you can also temporarily shut down your local server to trigger fail-over and confirm the behavior through this header's changing values.