How to Read Token Counts, Context Window Utilization, and Billing Info in the Copilot SDK
The Copilot SDK exposes token counts and billing data through ModelBilling metadata on available models and real-time session.usage events that report per-request consumption and context limits.
The GitHub Copilot SDK provides granular visibility into token consumption and cost management through type-safe TypeScript interfaces. Developers can query model-specific pricing tiers and context window limits before initiating sessions, then monitor actual usage through event-driven callbacks. This guide demonstrates how to access billing information, track context utilization, and calculate costs using the actual source structures found in github/copilot-sdk.
Understanding the Data Sources
The SDK surfaces usage and billing information through two distinct mechanisms that work together to provide a complete cost picture.
Model Metadata (ModelBilling): Available via client.getAvailableModels(), each model description contains a billing object that defines per-token-batch pricing and context window limits. According to the source code in nodejs/src/generated/rpc.ts (lines 9125-9140), the ModelBilling interface includes tokenPrices with inputPrice, outputPrice, and maxPromptTokens fields, plus an optional longContext tier for extended windows.
Session Usage Events: While a session runs, the SDK emits session.usage events containing real-time token consumption. These events implement the AssistantUsageEvent interface defined in nodejs/src/generated/session-events.ts (lines 2184-2190), providing tokenCount, contextMax, and detailed tokenDetails breakdowns. For cumulative tracking, the session.usage_checkpoint event (lines 2220-2240) emits UsageCheckpointEvent data containing aggregated totals for the entire session.
Both sources respect the account's billing mode, indicated by the token_based_billing boolean found in quota snapshots within rpc.ts (lines 2831-2833).
Retrieving Model-Level Billing and Context Limits
Before starting a session, inspect the ModelBilling data to understand pricing tiers and maximum context windows. This data lives in the billing field of each model returned by getAvailableModels().
import { CopilotClient } from '@github/copilot-sdk';
const client = new CopilotClient({ token: process.env.COPILOT_TOKEN });
async function listModelBilling() {
const models = await client.getAvailableModels();
for (const model of models) {
const billing = model.billing;
if (!billing?.tokenPrices) continue;
console.log(`Model: ${model.id}`);
console.log(` Default context max: ${billing.tokenPrices.maxPromptTokens}`);
console.log(` Long-context max: ${billing.tokenPrices.longContext?.maxPromptTokens}`);
console.log(` Input price/batch: ${billing.tokenPrices.inputPrice}`);
console.log(` Output price/batch: ${billing.tokenPrices.outputPrice}`);
}
}
listModelBilling();
The ModelBillingTokenPrices interface defines the batchSize property (typically representing the number of tokens per billing unit), which you will need later when calculating actual costs from raw token counts.
Monitoring Per-Request Token Usage and Context Utilization
During an active session, listen for session.usage events to track individual request consumption and verify context window utilization. These events fire immediately after each Copilot completion request.
async function monitorUsage() {
const session = await client.startSession({ modelId: 'gpt-4' });
session.on('session.usage', (event) => {
// event implements AssistantUsageEvent
console.log('Request token count:', event.tokenCount);
console.log('Context limit for this tier:', event.contextMax);
console.log('Token breakdown:', event.tokenDetails);
// Calculate utilization percentage
const utilization = (event.tokenCount / event.contextMax) * 100;
console.log(`Context window utilization: ${utilization.toFixed(1)}%`);
});
// Continue with normal session operations...
// await session.prompt('Explain this code...');
}
monitorUsage();
The contextMax field reflects the specific tier's limit (default or long-context) based on the model configuration, while tokenDetails provides optional categorization of prompt versus completion tokens.
Tracking Cumulative Usage with Checkpoints
For long-running sessions or accurate billing aggregation, consume the session.usage_checkpoint event. This emits a UsageCheckpointEvent containing cumulative totals since session start, ideal for resumable sessions or periodic cost reporting.
async function trackCheckpoint() {
const session = await client.startSession({ modelId: 'gpt-4' });
session.on('session.usage_checkpoint', (checkpoint) => {
// checkpoint.data contains aggregated totals
console.log('Cumulative input tokens:', checkpoint.data.totalInputTokens);
console.log('Cumulative output tokens:', checkpoint.data.totalOutputTokens);
// Store these values for external billing aggregation
const sessionTotals = {
input: checkpoint.data.totalInputTokens,
output: checkpoint.data.totalOutputTokens,
timestamp: checkpoint.timestamp
};
});
}
trackCheckpoint();
Unlike per-request events, checkpoints provide a definitive running total that accounts for all tokens consumed in the session up to that point, eliminating the need to sum individual requests manually.
Calculating Real-Time Costs
Combine model metadata with checkpoint data to compute actual AI credit costs. The calculation requires dividing token counts by the billing batch size, then multiplying by the respective input and output prices.
async function calculateAccumulatedCost() {
// Cache billing info for cost calculation
const models = await client.getAvailableModels();
const billingMap = new Map(
models.map(m => [m.id, m.billing?.tokenPrices])
);
const session = await client.startSession({ modelId: 'gpt-4' });
const prices = billingMap.get('gpt-4');
if (!prices) throw new Error('Billing info unavailable');
session.on('session.usage_checkpoint', (ckpt) => {
const { inputPrice = 0, outputPrice = 0, batchSize = 1 } = prices;
const inputBatches = Math.ceil(ckpt.data.totalInputTokens / batchSize);
const outputBatches = Math.ceil(ckpt.data.totalOutputTokens / batchSize);
const totalCost = (inputBatches * inputPrice) + (outputBatches * outputPrice);
console.log(`Accumulated AI credits: ${totalCost}`);
console.log(`Billing method: ${prices.tokenBasedBilling ? 'Per-token' : 'Request-based'}`);
});
}
calculateAccumulatedCost();
This approach accounts for the batch-based pricing model where costs are calculated per fixed-size token batch rather than per individual token.
Summary
- Query
ModelBillingfirst: Callclient.getAvailableModels()to retrieve pricing tiers (inputPrice,outputPrice) and context limits (maxPromptTokens,longContext.maxPromptTokens) fromnodejs/src/generated/rpc.ts. - Monitor
session.usageevents: Listen during active sessions to capture per-requesttokenCount,contextMax, andtokenDetailsvia theAssistantUsageEventinterface innodejs/src/generated/session-events.ts. - Aggregate with checkpoints: Use
session.usage_checkpointevents to get cumulativetotalInputTokensandtotalOutputTokenswithout manual summation. - Calculate costs by batch: Divide token totals by
batchSizefromModelBillingTokenPrices, multiply by respective prices, and sum for total AI credits consumed. - Check billing mode: Verify
token_based_billingon quota snapshots to determine if the account uses token-based or request-based billing.
Frequently Asked Questions
How do I check if my account uses token-based billing versus request-based billing?
Inspect the quota snapshot returned by the SDK for the token_based_billing boolean field. According to nodejs/src/generated/rpc.ts (lines 2831-2833), this flag indicates whether the account consumes AI credits per token or uses a fixed premium-request quota. This determines whether you need to calculate costs using the batch pricing method or simply count requests.
What is the difference between contextMax and maxPromptTokens?
maxPromptTokens (found in ModelBilling) defines the absolute maximum tokens the model accepts for a given tier, while contextMax (found in AssistantUsageEvent) reflects the actual limit applied to the current request. The latter may vary based on dynamic constraints or specific model configurations during the session, whereas the former is a static property of the billing tier.
When should I use session.usage_checkpoint instead of summing session.usage events?
Use session.usage_checkpoint when you need authoritative cumulative totals for resumable sessions, billing aggregation, or when network reliability might cause you to miss individual usage events. The checkpoint event provides a ground-truth snapshot of totalInputTokens and totalOutputTokens that accounts for all activity up to that point, eliminating drift from missed events.
Where are the TypeScript interfaces for billing and usage events defined?
The ModelBilling and ModelBillingTokenPrices interfaces reside in nodejs/src/generated/rpc.ts (around lines 9125-9140), while usage event types like AssistantUsageEvent and UsageCheckpointEvent are defined in nodejs/src/generated/session-events.ts (lines 2184-2190 and 2220-2240 respectively). These generated files contain the authoritative schema for all billing and token count fields.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →