How FreeLLMAPI Implements OpenAI-Compatible Tool Calling Across Multiple LLM Providers

FreeLLMAPI normalizes OpenAI's function-calling protocol by probing each model for tool support, filtering routes to only capable providers, and short-circuiting fusion panels when structured tool calls are detected.

FreeLLMAPI is an open-source universal router that aggregates free-tier LLM providers under a single OpenAI-compatible API. To support OpenAI-compatible tool calling seamlessly across disparate backends like Groq, Anthropic, Gemini, and Cloudflare, the project implements a four-layer normalization system spanning type definitions, capability discovery, intelligent routing, and parallel-panel handling.

Unified Type Definitions for OpenAI-Compatible Schemas

All tool-related fields adhere to a canonical OpenAI shape defined in shared/types.ts. This ensures every provider adapter receives and returns identical structures regardless of the underlying API.

The core interfaces include:

  • ChatToolDefinition (lines 59-64): Describes available functions using the standard OpenAI type, function.name, function.description, and function.parameters schema.

  • ChatToolChoice (lines 64-73): Controls invocation behavior with "none", "auto", "required", or specific function selection.

  • ChatToolCall and ChatToolCallFunction (lines 45-50): Carry the model-generated payload containing the target function name and JSON-encoded arguments.

  • ChatMessage (lines 84-90): Extended to include an optional tool_calls array, enabling the proxy to forward structured responses back to the client unchanged.

  • ChatCompletionRequest (lines 96-108): The central request body consumed throughout the server, explicitly declaring optional tools, tool_choice, and parallel_tool_calls fields.

By centralizing these definitions in shared/types.ts, FreeLLMAPI guarantees that client code written for the OpenAI SDK works identically against any configured provider.

Automatic Capability Discovery via Model Probing

Before a model enters the routing pool, server/src/services/model-discovery.ts executes a lightweight probe to determine if it supports structured tool calls. The service sends a dummy get_weather function definition and inspects the response's finish_reason:

// model-discovery.ts – lines 555-572
const toolMessages: ChatMessage[] = [
  { role: 'user', content: 'What is the weather in Paris? Use the get_weather tool.' },
];
const toolRes = await provider.chatCompletion(apiKey, toolMessages, modelId, {
  max_tokens: 64,
  temperature: 0,
  tools: [TOOL_PROBE],
  tool_choice: 'auto',
});
const finishReason = toolRes.choices?.[0]?.finish_reason;
toolCalls = finishReason === 'tool_calls';

If the model returns tool_calls as the finish reason, the discovery service persists this capability to the database:

// model-discovery.ts – line 487
db.prepare(`UPDATE models SET supports_tools = ? WHERE id = ?`).run(toolCalls ? 1 : 0, modelId);

This supports_tools boolean flag becomes the authoritative signal for all subsequent routing decisions, ensuring that tool-capable requests never reach incompatible models.

Routing and Request-Level Filtering

The core router in server/src/services/router.ts constructs a fallback chain of enabled models. When an incoming request contains a non-empty tools array, the router sets requireTools = true and filters the chain accordingly:

// router.ts – lines 2003-2008
if (requireTools && !entry.supports_tools) {
  diag.push(`${label}: no tool-calling support`);
  continue;
}

The same validation occurs within the Fusion service (server/src/services/fusion.ts) when building parallel panels:

// fusion.ts – lines 560-562
const requireTools = (options.tools?.length ?? 0) > 0;
if (requirements.requireTools && !cand.supportsTools) {
  dropped.push(`${id} (no tool-calling support)`);
  continue;
}

This dual-layer filtering ensures that only models with verified supports_tools = 1 ever receive tool-calling payloads, preventing errors from providers that lack function-calling capability.

Fusion Panel Handling of Tool Calls

FreeLLMAPI's Fusion feature executes multiple models in parallel to determine the best response. Because tool calls represent executable actions rather than generative text, the system treats them as high-priority outputs that must not be synthesized or altered.

When processing a Fusion panel, server/src/services/fusion.ts checks for structured tool calls immediately after receiving responses:

// fusion.ts – lines 640-646
const toolCallWinner = survivors.find(a =>
  (a.toolCalls?.length ?? 0) > 0 && a.rawChoice);
if (toolCallWinner) {
  // Return the raw tool call as the final response
}

If any panel member returns a valid tool_calls array, Fusion short-circuits the normal synthesis flow and returns that exact payload to the client with a finish_reason of "tool_calls". This preserves the functional integrity of the tool invocation, ensuring argument fidelity and correct function targeting.

Provider Adapter Implementation

Each provider adapter receives the normalized tools and tool_choice fields from the shared types and forwards them to the upstream API. The generic OpenAI-compatible adapter in server/src/providers/openai-compat.ts demonstrates this passthrough:

// providers/openai-compat.ts – lines 204 (non-stream) & 323 (stream)
await fetch(endpoint, {
  method: 'POST',
  body: JSON.stringify({
    model,
    messages,
    tools: options?.tools,
    tool_choice: options?.tool_choice,
    // …other OpenAI params
  })
});

Specialized adapters for Cohere, Cloudflare, Gemini, and others implement sanitizing helpers (such as sanitizeCohereTools or ollamaTools) that map the canonical OpenAI shapes to provider-specific formats while maintaining the same interface contract.

End-to-End Implementation Example

The following client request works against any free-tier provider configured in FreeLLMAPI, demonstrating the seamless OpenAI-compatible tool calling abstraction:

import fetch from 'node-fetch';

const body = {
  model: 'auto',                // Router selects a tool-capable model
  messages: [{ role: 'user', content: 'What is the weather in Paris?' }],
  tools: [
    {
      type: 'function',
      function: {
        name: 'get_weather',
        description: 'Get the current weather for a city',
        parameters: {
          type: 'object',
          properties: {
            city: { type: 'string', description: 'Name of the city' }
          },
          required: ['city']
        }
      }
    }
  ],
  tool_choice: 'auto'
};

fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify(body)
})
  .then(r => r.json())
  .then(resp => {
    if (resp.choices?.[0]?.message?.tool_calls) {
      console.log('Tool call received:', resp.choices[0].message.tool_calls);
    } else {
      console.log('Plain text answer:', resp.choices[0].message.content);
    }
  });

Under the hood, FreeLLMAPI detects the tools array, filters the model fallback chain to only those entries where supports_tools = 1, dispatches the request, and returns the raw tool call without modification if generated.

Summary

  • Unified schemas in shared/types.ts enforce OpenAI-compatible ChatToolDefinition, ChatToolCall, and ChatCompletionRequest structures across all providers.
  • Capability discovery in model-discovery.ts probes each model with a dummy function and records tool support via the supports_tools database flag.
  • Request filtering in router.ts and fusion.ts automatically excludes non-tool-capable models when the tools array is present.
  • Fusion short-circuiting returns the first valid tool_calls response immediately, bypassing text synthesis to preserve functional accuracy.
  • Provider adapters forward normalized tool payloads to upstream APIs while handling provider-specific sanitization internally.

Frequently Asked Questions

How does FreeLLMAPI determine if a specific model supports tool calling?

The model-discovery.ts service sends a probe request containing a dummy get_weather tool definition when a model first appears in the catalog. If the model responds with a finish_reason of tool_calls, the system sets supports_tools = 1 in the database. This flag is then used by the router to filter eligible models for tool-calling requests.

Can I use tool calling with the Fusion panel feature enabled?

Yes. When Fusion is active, the system runs multiple models in parallel. If any panel member returns a structured tool_calls payload, FreeLLMAPI immediately returns that specific response to the client and skips the normal synthesis process. This ensures tool calls are returned exactly as generated without aggregation or modification.

What happens if no tool-capable models are available for my request?

If the router cannot find any enabled models with supports_tools = 1 when the request includes a tools array, it logs diagnostic messages indicating "no tool-calling support" for each candidate and returns an error to the client. This prevents sending tool schemas to providers that cannot process them, avoiding runtime failures.

Do I need to modify my OpenAI SDK code to work with FreeLLMAPI?

No. Because FreeLLMAPI uses the exact type definitions from shared/types.ts for tools, tool_choice, and tool_calls, standard OpenAI SDK clients work without modification. Simply point the base URL to your FreeLLMAPI instance and ensure your requested model identifier matches or use auto to let the router select a capable backend.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →