How Token Counting Works Consistently Across Different AI Providers in 5ire
The 5ire client unifies disparate tokenization schemes from OpenAI, Google Gemini, Moonshot, and local Llama models behind a single utility module that implements provider-specific strategies while exposing a uniform API.
Accurate token counting is essential for managing API costs and staying within context window limits when building LLM applications. The open-source 5ire repository solves this challenge by abstracting provider-specific tokenization logic into src/utils/token.ts, enabling consistent token counting across different AI providers without coupling the rest of the codebase to implementation details.
The Challenge of Provider-Specific Tokenization
Each AI provider implements tokenization differently. OpenAI uses Byte Pair Encoding (BPE) via the tiktoken library with specific encoding schemes like cl100k_base, while Google Gemini and Moonshot use proprietary tokenization algorithms exposed through REST endpoints. Local Llama models require JavaScript-based tokenizers that mirror the original Python implementations. Without abstraction, client code would need to handle these differences explicitly, leading to fragmentation and maintenance overhead.
The Unified Token Utility Architecture
The repository centralizes all token counting logic in src/utils/token.ts. This module exports provider-specific functions that implement distinct strategies while maintaining consistent function signatures. The rest of the 5ire codebase interacts with these utilities through a uniform interface, requesting token counts without knowledge of the underlying provider implementation.
Provider-Specific Implementation Strategies
OpenAI and Tiktoken Encoding
For OpenAI models, the countGPTTokens function leverages the tiktoken library to reproduce OpenAI's exact tokenizer behavior. The implementation maps model names to concrete encodings—for example, gpt-3.5-turbo-0613 and gpt-4-0613 resolve to cl100k_base. The function also accounts for OpenAI's chat format overhead, adding 3 tokens per message and 1 token per name to match the API's actual consumption.
// src/utils/token.ts (lines 26-55)
export function countGPTTokens(messages: IChatRequestMessage[], model: string): number {
// Maps model to encoding (cl100k_base for GPT-3.5/4)
// Adds tokensPerMessage = 3 and tokensPerName = 1
// Returns exact count matching OpenAI API
}
If a model name is unrecognized, the function falls back to cl100k_base encoding in the catch block (lines 35-39), ensuring the application always returns a number rather than crashing.
Google Gemini REST Endpoint
For Google Gemini models, 5ire delegates token counting to the provider's own infrastructure via the countTokens REST endpoint. The countTokensOfGemini function sends the message list to /v1beta/models/{model}:countTokens and returns the exact token count the model would consume. This approach eliminates guesswork about Google's internal tokenization rules.
// src/utils/token.ts (lines 67-85)
export async function countTokensOfGemini(
messages: IChatRequestMessage[],
apiBase: string,
apiKey: string,
model: string
): Promise<number> {
// POST to /v1beta/models/{model}:countTokens
// Returns total_tokens from response
}
Moonshot Token Estimation API
Moonshot follows a similar pattern to Gemini, exposing a dedicated token estimation endpoint. The countTokensOfMoonshot function queries /tokenizers/estimate-token-count with the message payload and extracts the total_tokens value. The implementation includes error handling that logs failures and returns 0, allowing the caller to implement fallback logic when the service is unavailable.
// src/utils/token.ts (lines 87-102)
export async function countTokensOfMoonshot(
messages: IChatRequestMessage[],
apiBase: string,
apiKey: string,
model: string
): Promise<number> {
// POST to /tokenizers/estimate-token-count
// Returns total_tokens or 0 on error
}
Local Llama Models with JavaScript Tokenizers
For local Llama models, 5ire loads JavaScript tokenizers at runtime that mirror the original Python implementations. The countTokenOfLlama function selects between llama3-tokenizer-js and llama-tokenizer-js based on the model version. It applies the same chat format overhead used for OpenAI (3 tokens per message, 1 token per name) to maintain consistency with the Llama chat template.
// src/utils/token.ts (lines 121-147)
export async function countTokenOfLlama(
messages: IChatRequestMessage[],
model: string
): Promise<number> {
// Dynamically imports llama3-tokenizer-js or llama-tokenizer-js
// Adds tokensPerMessage = 3 and tokensPerName = 1
// Returns deterministic count matching local model behavior
}
Ensuring Consistency Across Providers
The 5ire repository achieves consistency through four key architectural decisions:
- Provider-Specific APIs – For Gemini and Moonshot, the code queries the provider's own endpoint, removing any guesswork about internal tokenization rules.
- Standardised Overhead – All chat-style models (OpenAI and Llama) apply a fixed token overhead per message (3 tokens) and per name (1 token), matching the chat format specifications.
- Model Mapping – OpenAI helpers normalize aliases (e.g.,
gpt-3.5-turbo-0613,gpt-4-0613) so the correcttiktokenencoding is always selected. - Graceful Degradation – When a model isn't recognized, the fallback to
cl100k_baseencoding yields a deterministic count, preventing crashes.
These strategies let the rest of 5ire treat token counting as a black-box operation while still respecting each provider's exact rules.
Practical Usage Examples
The following example demonstrates how the unified API handles different providers without exposing implementation details:
import {
countGPTTokens,
countTokensOfGemini,
countTokensOfMoonshot,
countTokenOfLlama,
} from '@/utils/token';
import type { IChatRequestMessage } from 'intellichat/types';
// Example chat messages
const msgs: IChatRequestMessage[] = [
{ role: 'user', content: 'Explain token counting.' },
];
// OpenAI (GPT‑4)
const gptTokens = countGPTTokens(msgs, 'gpt-4');
console.log('GPT‑4 token count:', gptTokens);
// Google Gemini
const geminiTokens = await countTokensOfGemini(
msgs,
'https://generativelanguage.googleapis.com',
process.env.GEMINI_API_KEY!,
'gemini-1.5-pro'
);
console.log('Gemini token count:', geminiTokens);
// Moonshot
const moonshotTokens = await countTokensOfMoonshot(
msgs,
'https://api.moonshot.cn',
process.env.MOONSHOT_API_KEY!,
'moonshot-v1'
);
console.log('Moonshot token count:', moonshotTokens);
// Local Llama 3 model
const llamaTokens = await countTokenOfLlama(msgs, 'llama3-8b');
console.log('Llama token count:', llamaTokens);
Summary
- The 5ire repository centralizes token counting logic in
src/utils/token.tsto abstract provider-specific implementations. - OpenAI models use the
tiktokenlibrary withcl100k_baseencoding and fixed chat format overhead (3 tokens per message, 1 per name). - Google Gemini and Moonshot delegate counting to provider REST endpoints (
countTokensandestimate-token-countrespectively) for exact alignment with internal tokenizers. - Local Llama models use runtime JavaScript tokenizers (
llama3-tokenizer-js,llama-tokenizer-js) with the same chat overhead calculations. - Graceful fallback to
cl100k_baseencoding ensures the application never crashes on unrecognized models.
Frequently Asked Questions
Why does token counting vary between AI providers?
Different providers train their models with distinct tokenization algorithms—OpenAI uses Byte Pair Encoding (BPE) via tiktoken, while Google and Moonshot use proprietary schemes. These algorithms split text into subword units differently, causing the same prompt to consume varying token counts across providers. The 5ire repository handles these differences by implementing provider-specific counting strategies rather than using a one-size-fits-all estimator.
How does 5ire handle token counting for unknown OpenAI models?
When the countGPTTokens function encounters an unrecognized model name, it catches the error and falls back to the cl100k_base encoding. This fallback represents the most common encoding scheme used across modern GPT models, providing a best-effort token count that prevents application crashes while maintaining reasonable accuracy for cost estimation purposes.
Can I use the token counting utilities independently of the 5ire chat interface?
Yes, the token counting functions exported from src/utils/token.ts are designed as pure utilities that accept message arrays and model identifiers without dependencies on the UI layer. You can import countGPTTokens, countTokensOfGemini, countTokensOfMoonshot, or countTokenOfLlama into any TypeScript or JavaScript project to perform accurate token estimation for API cost management or context window validation.
Why do local Llama models use the same overhead calculation as OpenAI?
Local Llama models in 5ire apply the same chat format overhead (3 tokens per message, 1 token per name) because they follow the same conversational message structure as OpenAI's chat completions format. The countTokenOfLlama function adds this overhead after tokenizing the content with provider-specific JavaScript tokenizers, ensuring the total count reflects both the raw content tokens and the structural metadata required by the chat template.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →