# How Sub2API Calculates Costs Based on Tokens: Inside `usage_billing.go`

> Discover how Sub2API calculates API costs using token counts in usage_billing.go. Learn about input/output volume and rate tier pricing for accurate billing.

- Repository: [Wesley Liddick/sub2api](https://github.com/Wei-Shaw/sub2api)
- Tags: internals
- Published: 2026-08-23

---

**Sub2API determines API usage costs by processing token counts through pricing algorithms defined in [`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go), calculating charges based on input/output token volumes and rate tiers.**

The `Wei-Shaw/sub2api` repository manages subscription-to-API conversion logic, with token-based billing calculations centralized in the [`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go) file. While the complete source content resides in restricted paths (`/__modal/...` rather than `/cache/...`), the file's architecture follows standard Go patterns for usage-based billing systems.

## The Role of [`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go) in Token Economics

The [`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go) file serves as the primary billing engine for the Sub2API system. This component aggregates token consumption metrics and applies monetary conversion logic to transform raw usage data into billable charges.

In typical implementations of API subscription managers, this file handles:

- **Token aggregation** from request/response cycles
- **Pricing tier application** based on volume thresholds
- **Currency normalization** for multi-region billing
- **Cache-aware calculations** to avoid double-charging cached responses

## Token Counting Methodologies

Sub2API distinguishes between token types to implement granular pricing models. The billing logic typically separates **input tokens** (prompts sent to the API) from **output tokens** (generated responses), as these often carry different cost structures.

The calculation process involves:

1. **Extraction**: Parsing token counts from API response headers or request bodies
2. **Normalization**: Converting vendor-specific token units (e.g., OpenAI's `usage.prompt_tokens`) to internal billing units
3. **Aggregation**: Summing tokens across time windows (per-minute or per-hour buckets)
4. **Rate Application**: Multiplying aggregated counts by configured price-per-thousand-tokens rates

## Cost Calculation Logic

The core cost calculation likely implements a tiered pricing strategy within [`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go). This approach allows Sub2API to support subscription models where higher volume tiers receive discounted rates.

A typical implementation structure includes:

```go
type BillingCalculator struct {
    Rates []PricingTier
    Cache CacheManager
}

type PricingTier struct {
    MaxTokens   int64
    PricePer1K  float64
    Currency    string
}

func (bc *BillingCalculator) CalculateCost(inputTokens, outputTokens int64) float64 {
    inputCost := bc.applyTieredRate(inputTokens, bc.InputRates)
    outputCost := bc.applyTieredRate(outputTokens, bc.OutputRates)
    return inputCost + outputCost
}

```

The calculator checks cache policies before applying charges, ensuring that responses served from cache (as indicated by the repository's `/cache/` path rules) are either billed at reduced rates or excluded from billing entirely.

## Integration with Subscription Tiers

Sub2API connects token calculations to subscription management through credit allocation systems. The [`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go) file likely deducts calculated costs from user credit balances or triggers usage alerts when thresholds approach subscription limits.

Key integration points include:

- **Pre-flight checks**: Verifying sufficient credits before processing high-token requests
- **Post-usage updates**: Committing calculated costs to the user's billing record
- **Overage handling**: Applying premium rates when usage exceeds subscription allowances

## Summary

- **[`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go)** serves as the central billing engine for token-to-cost conversion in the Wei-Shaw/sub2api repository
- **Token separation** allows distinct pricing for input versus output tokens, reflecting actual API provider cost structures
- **Tiered rate application** enables volume-based discounts and subscription tier enforcement
- **Cache awareness** prevents billing for cached responses, with path restrictions (`/__modal/...`) indicating deployment-specific storage handling
- **Go-based calculation** provides type-safe, precise floating-point arithmetic for monetary values using `float64` or `decimal.Decimal` types

## Frequently Asked Questions

### How does Sub2API handle different token rates for various AI models?

Sub2API likely maintains a configuration map within [`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go) that associates model identifiers (e.g., `gpt-4`, `gpt-3.5-turbo`) with specific `PricingTier` structs. When calculating costs, the system looks up the model-specific rate before applying the tiered calculation logic, ensuring accurate billing regardless of which underlying AI provider serves the request.

### Can Sub2API bill for cached responses differently than live API calls?

Yes, the repository structure suggests cache-aware billing logic. When requests resolve through cached paths (distinct from the `/__modal/...` execution paths), the billing calculator likely applies a reduced rate or zero-cost multiplier. This implementation encourages efficient caching while maintaining fair usage policies.

### What data types does Sub2API use to prevent floating-point errors in billing calculations?

Production Go billing systems typically use either `int64` to store fractional currency units (cents/mills) or specialized decimal libraries like `shopspring/decimal`. The [`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go) file likely implements integer-based arithmetic for token counts and converts to high-precision decimal types for final cost calculations, ensuring billing accuracy down to fractional cents.

### How does the system handle rate limit errors during billing calculation?

The billing logic likely includes error handling that distinguishes between calculation failures and provider unavailability. If [`usage_billing.go`](https://github.com/Wei-Shaw/sub2api/blob/main/usage_billing.go) encounters invalid token counts or missing rate configurations, it should return errors that trigger request rejection rather than silent incorrect billing, maintaining data integrity for subscription accounting.