# How Token Counting Works Consistently Across Different AI Providers in 5ire

> Learn how 5ire unifies token counting across OpenAI, Gemini, Moonshot, and Llama models. Get consistent AI results with a single utility module and uniform API.

- Repository: [Ironben/5ire](https://github.com/nanbingxyz/5ire)
- Tags: deep-dive
- Published: 2026-03-07

---

**The 5ire client unifies disparate tokenization schemes from OpenAI, Google Gemini, Moonshot, and local Llama models behind a single utility module that implements provider-specific strategies while exposing a uniform API.**

Accurate token counting is essential for managing API costs and staying within context window limits when building LLM applications. The open-source 5ire repository solves this challenge by abstracting provider-specific tokenization logic into [`src/utils/token.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/utils/token.ts), enabling consistent token counting across different AI providers without coupling the rest of the codebase to implementation details.

## The Challenge of Provider-Specific Tokenization

Each AI provider implements tokenization differently. OpenAI uses Byte Pair Encoding (BPE) via the **tiktoken** library with specific encoding schemes like `cl100k_base`, while Google Gemini and Moonshot use proprietary tokenization algorithms exposed through REST endpoints. Local Llama models require JavaScript-based tokenizers that mirror the original Python implementations. Without abstraction, client code would need to handle these differences explicitly, leading to fragmentation and maintenance overhead.

## The Unified Token Utility Architecture

The repository centralizes all token counting logic in [`src/utils/token.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/utils/token.ts). This module exports provider-specific functions that implement distinct strategies while maintaining consistent function signatures. The rest of the 5ire codebase interacts with these utilities through a uniform interface, requesting token counts without knowledge of the underlying provider implementation.

## Provider-Specific Implementation Strategies

### OpenAI and Tiktoken Encoding

For OpenAI models, the `countGPTTokens` function leverages the **tiktoken** library to reproduce OpenAI's exact tokenizer behavior. The implementation maps model names to concrete encodings—for example, `gpt-3.5-turbo-0613` and `gpt-4-0613` resolve to `cl100k_base`. The function also accounts for OpenAI's chat format overhead, adding **3 tokens per message** and **1 token per name** to match the API's actual consumption.

```typescript
// src/utils/token.ts (lines 26-55)
export function countGPTTokens(messages: IChatRequestMessage[], model: string): number {
  // Maps model to encoding (cl100k_base for GPT-3.5/4)
  // Adds tokensPerMessage = 3 and tokensPerName = 1
  // Returns exact count matching OpenAI API
}

```

If a model name is unrecognized, the function falls back to `cl100k_base` encoding in the catch block (lines 35-39), ensuring the application always returns a number rather than crashing.

### Google Gemini REST Endpoint

For Google Gemini models, 5ire delegates token counting to the provider's own infrastructure via the `countTokens` REST endpoint. The `countTokensOfGemini` function sends the message list to `/v1beta/models/{model}:countTokens` and returns the exact token count the model would consume. This approach eliminates guesswork about Google's internal tokenization rules.

```typescript
// src/utils/token.ts (lines 67-85)
export async function countTokensOfGemini(
  messages: IChatRequestMessage[],
  apiBase: string,
  apiKey: string,
  model: string
): Promise<number> {
  // POST to /v1beta/models/{model}:countTokens
  // Returns total_tokens from response
}

```

### Moonshot Token Estimation API

Moonshot follows a similar pattern to Gemini, exposing a dedicated token estimation endpoint. The `countTokensOfMoonshot` function queries `/tokenizers/estimate-token-count` with the message payload and extracts the `total_tokens` value. The implementation includes error handling that logs failures and returns `0`, allowing the caller to implement fallback logic when the service is unavailable.

```typescript
// src/utils/token.ts (lines 87-102)
export async function countTokensOfMoonshot(
  messages: IChatRequestMessage[],
  apiBase: string,
  apiKey: string,
  model: string
): Promise<number> {
  // POST to /tokenizers/estimate-token-count
  // Returns total_tokens or 0 on error
}

```

### Local Llama Models with JavaScript Tokenizers

For local Llama models, 5ire loads JavaScript tokenizers at runtime that mirror the original Python implementations. The `countTokenOfLlama` function selects between `llama3-tokenizer-js` and `llama-tokenizer-js` based on the model version. It applies the same chat format overhead used for OpenAI (**3 tokens per message**, **1 token per name**) to maintain consistency with the Llama chat template.

```typescript
// src/utils/token.ts (lines 121-147)
export async function countTokenOfLlama(
  messages: IChatRequestMessage[],
  model: string
): Promise<number> {
  // Dynamically imports llama3-tokenizer-js or llama-tokenizer-js
  // Adds tokensPerMessage = 3 and tokensPerName = 1
  // Returns deterministic count matching local model behavior
}

```

## Ensuring Consistency Across Providers

The 5ire repository achieves consistency through four key architectural decisions:

- **Provider-Specific APIs** – For Gemini and Moonshot, the code queries the provider's own endpoint, removing any guesswork about internal tokenization rules.
- **Standardised Overhead** – All chat-style models (OpenAI and Llama) apply a fixed token overhead per message (**3 tokens**) and per name (**1 token**), matching the chat format specifications.
- **Model Mapping** – OpenAI helpers normalize aliases (e.g., `gpt-3.5-turbo-0613`, `gpt-4-0613`) so the correct `tiktoken` encoding is always selected.
- **Graceful Degradation** – When a model isn't recognized, the fallback to `cl100k_base` encoding yields a deterministic count, preventing crashes.

These strategies let the rest of 5ire treat token counting as a black-box operation while still respecting each provider's exact rules.

## Practical Usage Examples

The following example demonstrates how the unified API handles different providers without exposing implementation details:

```typescript
import {
  countGPTTokens,
  countTokensOfGemini,
  countTokensOfMoonshot,
  countTokenOfLlama,
} from '@/utils/token';
import type { IChatRequestMessage } from 'intellichat/types';

// Example chat messages
const msgs: IChatRequestMessage[] = [
  { role: 'user', content: 'Explain token counting.' },
];

// OpenAI (GPT‑4)
const gptTokens = countGPTTokens(msgs, 'gpt-4');
console.log('GPT‑4 token count:', gptTokens);

// Google Gemini
const geminiTokens = await countTokensOfGemini(
  msgs,
  'https://generativelanguage.googleapis.com',
  process.env.GEMINI_API_KEY!,
  'gemini-1.5-pro'
);
console.log('Gemini token count:', geminiTokens);

// Moonshot
const moonshotTokens = await countTokensOfMoonshot(
  msgs,
  'https://api.moonshot.cn',
  process.env.MOONSHOT_API_KEY!,
  'moonshot-v1'
);
console.log('Moonshot token count:', moonshotTokens);

// Local Llama 3 model
const llamaTokens = await countTokenOfLlama(msgs, 'llama3-8b');
console.log('Llama token count:', llamaTokens);

```

## Summary

- The 5ire repository centralizes token counting logic in [`src/utils/token.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/utils/token.ts) to abstract provider-specific implementations.
- **OpenAI** models use the `tiktoken` library with `cl100k_base` encoding and fixed chat format overhead (3 tokens per message, 1 per name).
- **Google Gemini** and **Moonshot** delegate counting to provider REST endpoints (`countTokens` and `estimate-token-count` respectively) for exact alignment with internal tokenizers.
- **Local Llama** models use runtime JavaScript tokenizers (`llama3-tokenizer-js`, `llama-tokenizer-js`) with the same chat overhead calculations.
- Graceful fallback to `cl100k_base` encoding ensures the application never crashes on unrecognized models.

## Frequently Asked Questions

### Why does token counting vary between AI providers?

Different providers train their models with distinct tokenization algorithms—OpenAI uses Byte Pair Encoding (BPE) via tiktoken, while Google and Moonshot use proprietary schemes. These algorithms split text into subword units differently, causing the same prompt to consume varying token counts across providers. The 5ire repository handles these differences by implementing provider-specific counting strategies rather than using a one-size-fits-all estimator.

### How does 5ire handle token counting for unknown OpenAI models?

When the `countGPTTokens` function encounters an unrecognized model name, it catches the error and falls back to the `cl100k_base` encoding. This fallback represents the most common encoding scheme used across modern GPT models, providing a best-effort token count that prevents application crashes while maintaining reasonable accuracy for cost estimation purposes.

### Can I use the token counting utilities independently of the 5ire chat interface?

Yes, the token counting functions exported from [`src/utils/token.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/utils/token.ts) are designed as pure utilities that accept message arrays and model identifiers without dependencies on the UI layer. You can import `countGPTTokens`, `countTokensOfGemini`, `countTokensOfMoonshot`, or `countTokenOfLlama` into any TypeScript or JavaScript project to perform accurate token estimation for API cost management or context window validation.

### Why do local Llama models use the same overhead calculation as OpenAI?

Local Llama models in 5ire apply the same chat format overhead (3 tokens per message, 1 token per name) because they follow the same conversational message structure as OpenAI's chat completions format. The `countTokenOfLlama` function adds this overhead after tokenizing the content with provider-specific JavaScript tokenizers, ensuring the total count reflects both the raw content tokens and the structural metadata required by the chat template.