How FreeLLMAPI Emulates the Ollama API for Full Compatibility
FreeLLMAPI provides an Ollama-compatible shim via the ollamaRouter in server/src/routes/ollama.ts that intercepts standard Ollama HTTP endpoints, validates requests through configurable emulation modes, and translates payloads into the internal FreeLLMAPI service layer.
FreeLLMAPI emulates the Ollama API to allow existing Ollama clients to communicate with the gateway without code modifications. According to the tashfeenahmed/freellmapi source code, this compatibility layer is centralized in the ollamaRouter and leverages the same core services—such as runInboundChat and runEmbeddings—that power FreeLLMAPI's native OpenAI-compatible endpoints.
Ollama Router Architecture and Entry Points
The emulation layer is built around the ollamaRouter defined in [server/src/routes/ollama.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts). This router mounts under the /api/ path and exposes standard Ollama endpoints including /api/tags, /api/chat, /api/generate, /api/embed, and /api/show.
Each endpoint performs three operations:
- Authorization via the
authorizehelper function - Request validation using Zod schemas (e.g.,
embedSchema) - Delegation to core services like
runInboundChat(defined in [server/src/lib/inbound-chat.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/inbound-chat.ts)) orrunEmbeddings(defined in [server/src/services/embeddings.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/embeddings.ts))
This architecture ensures that Ollama clients interact with FreeLLMAPI exactly as they would with a native Olloma server, while the gateway routes traffic through its unified model management, quota, and rate-limiting infrastructure.
Configurable Authorization Modes
FreeLLMAPI supports three distinct emulation modes controlled by the getOllamaEmulationMode function. The active mode is stored in the settings database under the key ollama_emulation and defaults to 'off' via the migration in [server/src/db/migrations/20260727_000001_agent_compat.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/db/migrations/20260727_000001_agent_compat.ts).
| Mode | Behavior |
|---|---|
off |
All Ollama routes return 404; emulation is disabled |
open-loopback |
Only requests from local loopback addresses (127.0.0.1, ::1) are permitted |
key-required |
Requires a valid unified API key in the Authorization: Bearer <key> header |
The authorize helper checks this setting before every request by querying getSetting('ollama_emulation') from [server/src/db/index.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/db/index.ts), ensuring flexible deployment scenarios from fully open local development to secured production environments.
Model Catalog and Inspection Endpoints
Listing Available Models (/api/tags)
When an Ollama client requests the model catalog via /api/tags, FreeLLMAPI calls buildModelListing() from [server/src/services/model-listing.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-listing.ts) to retrieve the available model index. The endpoint returns an "auto" placeholder model plus every configured model, each transformed by the ollamaModel() function into the Ollama-expected shape:
nameandmodelidentifiersmodified_attimestampsizeanddigestfor versioningdetailscontaining architecture family, parameter size, and quantization level
All models report a synthetic format value of freellmapi to identify the source gateway.
Model Details (/api/show)
The /api/show endpoint normalizes incoming model names using normalizeOllamaModel() and retrieves detailed metadata. For the special auto model, it returns generic capabilities including completion and tools support with a standard context window. For concrete models, the response includes modelfile, parameters, template, and model_info (e.g., general.architecture).
Chat and Generate Request Translation
FreeLLMAPI handles Ollama's conversational endpoints through sophisticated payload transformation before delegating to the core chat orchestration service.
Message and Tool Conversion
The ollamaMessages function rewrites Ollama-style message arrays—including multimodal content with vision images and tool call definitions—into the internal ChatMessage format. Similarly, ollamaTools maps Ollama tool definitions to FreeLLMAPI's native ChatToolDefinition structures, enabling function calling compatibility.
Response Wire Formatting
Responses are formatted using ollamaWire (for /api/chat) and generateWire (for legacy /api/generate). These adapters emit newline-delimited JSON (ndjson) streams that match Ollama's streaming protocol exactly:
// Conceptual implementation from ollamaWire
{
model: request.model,
created_at: new Date().toISOString(),
message: { role: 'assistant', content: text },
done: false,
eval_count: tokenCount,
// ... timing metrics via ollamaDurations
}
The adapters translate internal finishReason values to Ollama's done_reason field (values: stop or length) and fabricate timing metrics via ollamaDurations to satisfy clients expecting tokens-per-second statistics.
Embeddings Endpoint Compatibility
Both /api/embed and the legacy /api/embeddings endpoints share the same implementation flow:
- Authorization via the
authorizehelper - Request validation against
embedSchema - Delegation to
runEmbeddingsfor vector generation - Response reshaping to match Ollama's embedding schema
This ensures that vector retrieval works identically whether clients use the modern or legacy endpoint naming convention.
Practical Usage Examples
The following curl commands demonstrate interacting with FreeLLMAPI's Ollama-compatible endpoints. Replace <YOUR_UNIFIED_API_KEY> with your actual API key when using key-required mode.
List available models through the /api/tags endpoint:
curl -X GET http://localhost:3000/api/tags \
-H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>"
Retrieve detailed model information via /api/show:
curl -X POST http://localhost:3000/api/show \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>" \
-d '{"model":"gpt-oss:120b"}'
Send a chat completion request to /api/chat:
curl -X POST http://localhost:3000/api/chat \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>" \
-d '{
"model":"gpt-oss:120b",
"messages":[{"role":"user","content":"Explain quantum tunnelling"}],
"stream":false
}'
Use the legacy generate endpoint at /api/generate:
curl -X POST http://localhost:3000/api/generate \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>" \
-d '{
"model":"gpt-oss:120b",
"prompt":"Write a haiku about winter",
"stream":false
}'
Generate embeddings through /api/embed:
curl -X POST http://localhost:3000/api/embed \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>" \
-d '{"model":"gpt-oss:120b","input":"Free LLMAPI"}'
Summary
- Centralized Router: The
ollamaRouterinserver/src/routes/ollama.tsimplements all Ollama-compatible HTTP endpoints under/api/ - Flexible Security: Three emulation modes (
off,open-loopback,key-required) control access via theollama_emulationsetting stored in the database - Protocol Translation: Functions like
ollamaMessages,ollamaTools, andollamaWireconvert between Ollama and internal formats while preserving streaming semantics - Model Compatibility: The
buildModelListing()andollamaModel()adapters inserver/src/services/model-listing.tspresent FreeLLMAPI's unified model catalog in Ollama-native format - Unified Backend: Despite the Ollama facade, all requests flow through FreeLLMAPI's centralized backend for consistent quota management and rate limiting
Frequently Asked Questions
How do I enable Ollama API emulation in FreeLLMAPI?
Set the ollama_emulation setting to either open-loopback for local-only access or key-required for production use with API key authentication. The setting is stored in the database and defaults to off via the migration 20260727_000001_agent_compat.ts to prevent unauthorized access.
Which Ollama endpoints does FreeLLMAPI support?
FreeLLMAPI supports the complete Ollama endpoint set including /api/tags (model listing), /api/show (model details), /api/chat (conversational AI), /api/generate (legacy completions), and both /api/embed and /api/embeddings (vector generation), all implemented in server/src/routes/ollama.ts.
Does FreeLLMAPI support tool calling through the Ollama API?
Yes. The ollamaTools function translates Ollama tool definitions into FreeLLMAPI's internal ChatToolDefinition format, and ollamaMessages handles tool response messages. This enables function calling capabilities when using the /api/chat endpoint with compatible models.
Can I use the standard Ollama CLI with FreeLLMAPI?
Yes. Point the Ollama CLI to your FreeLLMAPI instance URL instead of the default localhost:11434. When using key-required mode, configure the CLI to include the Bearer token in requests, or use environment variables mapped through server/src/lib/key-parser.ts to authenticate requests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →