# How OmniRoute Free Tier Budget Calculation Works: Provider Limits vs. Token Efforts

> Understand OmniRoute free tier budget calculation. Learn how provider limits and token efforts determine your cap and prevent exceeding it.

- Repository: [Diego Rodrigues de Sa e Souza/OmniRoute](https://github.com/diegosouzapw/OmniRoute)
- Tags: how-to-guide
- Published: 2026-08-07

---

**OmniRoute calculates free-tier budgets by taking the minimum of provider-declared rate limits and token-derived reasoning efforts, rejecting any request that exceeds the resulting cap.**

OmniRoute (diegosouzapw/OmniRoute) treats every free-tier model as a budget-constrained endpoint. The system derives available capacity using two independent sources: static provider limits published in registry entries and dynamic token budgets converted to effort levels. This dual-source approach ensures that free-tier usage stays within both technical rate limits and cost-management boundaries defined in the source code.

## Provider-Declared Rate Limits in Registry Entries

Each provider offering a free tier publishes rate-limit descriptions in the **`freeNote`** field of their registry entry. These human-readable strings—such as “~40 RPM”, “2 req·s⁻¹”, or “100 req·hr⁻¹”—are parsed by [`open-sse/services/providerDefaultRateLimit.ts`](https://github.com/diegosouzapw/OmniRoute/blob/main/open-sse/services/providerDefaultRateLimit.ts) to compute a per-second request budget for that specific model. The service translates declarative strings into concrete numeric caps that the budget manager enforces at runtime.

## Token Budget to Effort Conversion

When an API request contains a **`thinking`** or **`budget_tokens`** field, OmniRoute invokes the XAI translator's `budgetToEffort` routine. This function maps the requested token count to a discrete reasoning-effort bucket that influences the final capacity calculation.

### The XAI Translator Logic

The conversion logic resides in the XAI translator layer and is validated by the unit tests in [`tests/unit/xai-translators.test.ts`](https://github.com/diegosouzapw/OmniRoute/blob/main/tests/unit/xai-translators.test.ts). The routine categorizes incoming requests into predefined effort levels based strictly on the numeric token value provided in the request body.

### Effort Bucket Mapping

The `budgetToEffort` function implements the following token-to-effort mapping:

- **0, negative, or NaN values**: Returns `undefined` (no extra budget allocated)
- **1–3,999 tokens**: Maps to the `low` effort bucket
- **4,000–15,999 tokens**: Maps to the `medium` effort bucket
- **16,000+ tokens**: Maps to the `high` effort bucket

## Free-Tier Restrictions and Default Behavior

Free-tier models are **restricted to the `low` bucket** unless the provider explicitly advertises a higher allowance, such as Gemini Flash’s “~40 RPM” rating. If the client omits a budget entirely, the translator falls back to the default effort for the model—typically **`low`** for pure free-tier models and **`medium`** for mixed free/paid offerings. This ensures that unspecified requests do not accidentally exhaust limited free quotas.

## Budget Resolution and Enforcement

The final per-request budget equals the **minimum** of the provider-level request budget (derived from `freeNote`) and the token-budget-derived effort cap. If the calculated budget exceeds the provider’s free quota, the request is rejected with the standard error message: "free tier of the model has been exhausted". The [`open-sse/services/accountFallback.ts`](https://github.com/diegosouzapw/OmniRoute/blob/main/open-sse/services/accountFallback.ts) service classifies this condition as a quota-exhausted state and handles the rejection accordingly.

## Code Implementation Examples

The following TypeScript example demonstrates how a client request flows through the budget calculation:

```typescript
// Example: a client asks a free-tier Gemini model for 10,000 tokens.
const request = {
  model: "gemini-2.5-flash-free",
  thinking: { type: "enabled", budget_tokens: 10_000 },
};

// XAI translator runs budgetToEffort → "medium"
const effort = budgetToEffort(request.thinking.budget_tokens); // "medium"

// Provider's freeNote says "~40 RPM" → translates to ≈ 0.66 req/s.
// The free-tier manager caps the request to the lower of the two budgets.
const finalBudget = Math.min(effortToTokens(effort), providerFreeTokenCap);

```

The `budgetToEffort` function follows this implementation pattern, as validated in the unit test suite:

```typescript
// Pseudo-implementation of budgetToEffort (see tests in
// tests/unit/xai-translators.test.ts)
function budgetToEffort(tokens: number): "low" | "medium" | "high" | undefined {
  if (!Number.isFinite(tokens) || tokens <= 0) return undefined;
  if (tokens <= 3_999) return "low";
  if (tokens <= 15_999) return "medium";
  return "high";
}

```

## Summary

- OmniRoute calculates free-tier budgets using the **minimum** of provider rate limits and token-derived effort levels.
- Provider limits are parsed from the **`freeNote`** field by [`open-sse/services/providerDefaultRateLimit.ts`](https://github.com/diegosouzapw/OmniRoute/blob/main/open-sse/services/providerDefaultRateLimit.ts).
- Token budgets convert to efforts (`low`, `medium`, `high`) via **`budgetToEffort`** in the XAI translator, with test coverage in [`tests/unit/xai-translators.test.ts`](https://github.com/diegosouzapw/OmniRoute/blob/main/tests/unit/xai-translators.test.ts).
- Free-tier models default to the **`low`** effort bucket unless explicitly allowed higher limits by the provider.
- Excess budgets trigger rejection via [`open-sse/services/accountFallback.ts`](https://github.com/diegosouzapw/OmniRoute/blob/main/open-sse/services/accountFallback.ts) with a quota-exhausted classification and the error "free tier of the model has been exhausted".

## Frequently Asked Questions

### What happens if I request 10,000 tokens on a standard free-tier model?

The `budgetToEffort` function maps 10,000 tokens to the **`medium`** effort bucket. However, free-tier models are typically restricted to the **`low`** bucket unless the provider explicitly supports higher limits. The system calculates the minimum of the provider's rate limit and the effort-derived budget, meaning the request will be capped at the `low` level or rejected if it exceeds the free quota.

### Where does OmniRoute store provider rate limits for free tiers?

Rate limits are stored in the **`freeNote`** field of each provider's registry entry. The [`open-sse/services/providerDefaultRateLimit.ts`](https://github.com/diegosouzapw/OmniRoute/blob/main/open-sse/services/providerDefaultRateLimit.ts) service parses these declarations—such as “~40 RPM” or “100 req·hr⁻¹”—into per-second request budgets that feed into the final calculation.

### What error occurs when the free-tier budget is exceeded?

When a request exceeds the calculated budget, OmniRoute returns the error message **"free tier of the model has been exhausted"**. The [`open-sse/services/accountFallback.ts`](https://github.com/diegosouzapw/OmniRoute/blob/main/open-sse/services/accountFallback.ts) service classifies this as a quota-exhausted condition and handles the rejection before the request reaches the provider.

### Can I use high reasoning effort on free-tier models?

Generally no. Free-tier models are restricted to the **`low`** effort bucket unless the provider explicitly advertises a higher allowance, such as Gemini Flash’s “~40 RPM” rating. Requests for **`medium`** or **`high`** effort (4,000+ tokens) on restricted free-tier models will be constrained to the `low` bucket minimum to prevent quota exhaustion.