How FreeLLMAPI Handles Provider-Specific Nuances Through Its Modular Adapter Architecture

FreeLLMAPI normalizes disparate LLM provider behaviors through a layered adapter system where all providers inherit from BaseProvider and override only vendor-specific behaviors while sharing common utilities for HTTP handling, streaming, and error management.

FreeLLMAPI tackles the fragmentation of modern LLM platforms by abstracting vendor differences behind a clean, modular adapter layer. At the core of this design sits the abstract BaseProvider class defined in [server/src/providers/base.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts), which supplies the foundational plumbing for HTTP calls, timeout handling, retry-after parsing, and stream management. Concrete provider implementations—such as OpenAICompatProvider, GoogleProvider, SailProvider, CohereProvider, and CloudflareProvider—live under server/src/providers/ and selectively override only the behaviors that diverge from the standard contract.

The BaseProvider Contract in server/src/providers/base.ts

The BaseProvider class defines the canonical interface that every LLM adapter must satisfy. It encapsulates the universal concerns that would otherwise be duplicated across dozens of vendor integrations.

Key responsibilities implemented in the base class include:

  • HTTP orchestration – Generic request dispatch and response handling
  • Timeout management – Uniform logic using providerTimeoutMs and timeoutBounds
  • Retry-after extraction – Parsing the Retry-After header via parseStatedRetryMs to harvest back-off hints from response bodies
  • Stream parsing – readSseStream handles Server-Sent Events (SSE) parsing and stall detection using streamStallTimeoutMs
  • Error normalization – providerHttpError extracts HTTP status codes and structures a uniform ProviderHttpError

By centralizing these capabilities, BaseProvider ensures that provider-specific nuances do not leak into the routing layer or quota management systems.

How Concrete Providers Override Provider-Specific Behaviors

Each concrete provider subclasses BaseProvider and overrides targeted methods to accommodate vendor idiosyncrasies. This selective override pattern keeps the codebase modular while respecting platform differences.

Endpoint URL and Path Construction

While BaseProvider builds generic URLs from a platform identifier, individual providers supply their base URLs and specific routes (e.g., /v1/chat/completions, /v1/completions). The [server/src/providers/openai-compat.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/openai-compat.ts) adapter demonstrates this by handling both Azure OpenAI and OpenRouter path variations.

Authentication Schemes

The base class handles generic Authorization header logic, but providers inject their own schemes:

  • Bearer tokens for standard OpenAI-compatible services
  • API-Key headers for platforms requiring x-api-key
  • Custom query parameters for services that authenticate via URL arguments

The [server/src/providers/google.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts) adapter manages Google’s protobuf-style auth schema and retry info parsing, while [server/src/providers/sail.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/sail.ts) demonstrates custom header handling for the Sail platform.

Payload Transformation

Providers receive a standard ChatCompletionRequest matching the OpenAI schema, but may rename fields, add required flags, or drop unsupported keys. For example, some providers reject logprobs or tool_choice parameters, requiring the adapter to sanitize the payload before transmission.

Streaming Adaptations

While BaseProvider.readSseStream handles standard SSE framing, providers can tweak parsing logic for vendor-specific streaming formats. This accommodates variations in how different services chunk their response streams or signal completion.

Error Normalization

Providers may add custom parsing for vendor-specific error structures. Google’s google.rpc.RetryInfo, for instance, requires specialized extraction logic beyond standard HTTP headers, implemented in the Google provider adapter.

The Quirks System for Provider-Specific Adjustments

Vendor peculiarities that affect request parameters rather than transport logic are codified in [server/src/services/quirks.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/quirks.ts). This module maintains a lookup table of adjustments applied automatically by the router:

  • Max-tokens caps – Some providers silently cap max_tokens lower than advertised model limits; the quirk table clamps requests accordingly
  • Token multipliers – Services billing tokens differently (e.g., counting certain model inputs twice) have multipliers applied before quota updates
  • Streaming-first-byte timing – Providers delaying the initial SSE byte trigger adjustments to streamStallTimeoutMs to prevent premature timeout errors

The quirks system integrates with [server/src/services/provider-quota.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts), which consumes these provider-specific caps and cooldowns to maintain accurate rate limiting across heterogeneous backends.

Request Routing and Error Normalization

The [server/src/services/router.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) serves as the central entry point for all LLM requests. When a request arrives, the router:

  1. Resolves the appropriate BaseProvider via getProvider or resolveProvider
  2. Invokes the provider’s complete method, passing CompletionOptions including extended sampling knobs
  3. Relies on shared error-handling utilities (providerHttpError, parseStatedRetryMs) to transform provider-specific HTTP failures into uniform ProviderHttpError instances
  4. Updates quota state using provider-specific information supplied by the quirks module

This architecture allows the router, quota manager, and rate-limiter to treat all LLM backends uniformly while the adapter layer handles provider-specific nuances transparently.

Working with Provider Adapters: Code Examples

The following example demonstrates invoking Google Gemini through the unified API, where provider-specific quirks are applied automatically:

// Example: invoking an LLM through the unified API
import { getProvider } from '@/providers/index.js';
import type { CompletionOptions } from '@/providers/base.js';

async function askGemini(prompt: string) {
  const provider = getProvider('google')!;          // Google‑Gemini provider
  const opts: CompletionOptions = {
    model: 'gemini-1.5-flash',
    temperature: 0.7,
    max_tokens: 1024,
    // provider‑specific quirks (e.g., token multiplier) are applied automatically
  };

  const response = await provider.complete({
    messages: [{ role: 'user', content: prompt }],
    ...opts,
  });

  console.log(response.choices[0].message.content);
}

For error handling with back-off logic, the adapter layer exposes normalized retry information:

// Example: handling a provider‑specific error with back‑off
import { getProvider } from '@/providers/index.js';

async function safeRequest() {
  const provider = getProvider('sail')!;
  try {
    await provider.complete({ messages: [], model: 'sail-7b' });
  } catch (e) {
    if ((e as any).retryAfterMs) {
      // Respect the provider‑stated cooldown before retrying
      await new Promise(r => setTimeout(r, (e as any).retryAfterMs));
    }
    throw e;
  }
}

Summary

  • FreeLLMAPI abstracts provider-specific nuances through a modular adapter architecture centered on the BaseProvider abstract class in server/src/providers/base.ts.
  • Selective overrides allow concrete providers to customize authentication, payload shapes, endpoint routing, and error parsing without duplicating HTTP or streaming logic.
  • The quirks system in server/src/services/quirks.ts codifies parameter-level adjustments like token multipliers and max-token caps, automatically applied during request routing.
  • Uniform error normalization via providerHttpError and parseStatedRetryMs converts vendor-specific failure modes into a standard ProviderHttpError that upstream components can handle consistently.
  • The router (server/src/services/router.ts) orchestrates provider resolution and quota management while remaining agnostic to underlying platform differences.

Frequently Asked Questions

What is the BaseProvider class in FreeLLMAPI?

BaseProvider is the abstract base class defined in server/src/providers/base.ts that establishes the contract all LLM provider adapters must implement. It provides shared functionality for HTTP requests, timeout handling, SSE stream parsing, and error normalization, while allowing subclasses to override specific behaviors for vendor-specific requirements.

How does FreeLLMAPI handle different authentication schemes across providers?

Each concrete provider overrides the base authentication logic to inject its own header scheme. While BaseProvider handles generic Authorization headers, individual providers implement Bearer tokens, API-Key headers, or custom query parameters as needed. For example, the Google provider manages protobuf-style authentication schemas distinct from standard OpenAI-compatible services.

What are provider "quirks" in FreeLLMAPI?

Quirks are provider-specific behavioral adjustments codified in server/src/services/quirks.ts. They address idiosyncrasies such as max-token caps that are lower than advertised limits, token-counting multipliers for billing purposes, and variations in stream timing that require adjusted streamStallTimeoutMs values. The router applies these quirks automatically before dispatching requests.

How does error normalization work when providers return different response formats?

The BaseProvider class implements providerHttpError and parseStatedRetryMs to extract HTTP status codes and retry timing from responses. Concrete providers can extend this logic to parse vendor-specific error structures (such as Google's google.rpc.RetryInfo), but all errors are ultimately transformed into a uniform ProviderHttpError with standardized fields like retryAfterMs, allowing the router and quota manager to handle failures consistently regardless of the underlying provider.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →