# How the Bella OpenAPI Endpoint Logging System Handles Metrics and Cost Reporting

> Discover how the Bella OpenAPI logging system efficiently collects metrics and reports costs using a high-performance Disruptor pipeline with specialized handlers.

- Repository: [Ke Technologies/bella-openapi](https://github.com/lianjiatech/bella-openapi)
- Tags: internals
- Published: 2026-03-06

---

**The Bella OpenAPI endpoint logging system uses a high-performance Disruptor pipeline with specialized handlers—`MetricsLogHandler` for real-time usage counters and `CostLogHandler` for dynamic price evaluation—to decouple metrics collection from monetary cost reporting.**

The `lianjiatech/bella-openapi` repository implements a sophisticated endpoint logging system designed to process high-throughput AI API requests without blocking the critical path. This system captures granular usage metrics for monitoring dashboards while simultaneously calculating and persisting per-request costs for billing and quota enforcement.

## Architecture of the Endpoint Logging Pipeline

The endpoint logging system is built around the **LMAX Disruptor** ring buffer pattern, ensuring lock-free, high-performance event processing. Each incoming request generates a `LogEvent` that flows through a chain of specialized handlers.

### Core Components

| Component | Role | Key Source File |
|-----------|------|-----------------|
| **EndpointLogHandler** | Marker interface defining the contract `void onEvent(LogEvent event, long sequence, boolean endOfBatch)` that all concrete handlers implement. | [`api/server/src/main/java/com/ke/bella/openapi/protocol/log/EndpointLogHandler.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/server/src/main/java/com/ke/bella/openapi/protocol/log/EndpointLogHandler.java) |
| **MetricsLogHandler** | Aggregates usage counters by extracting `RequestMetrics` from `LogEvent` and forwarding to `MetricsManager` for storage in Redis/Caffeine. | [`api/server/src/main/java/com/ke/bella/openapi/protocol/log/MetricsLogHandler.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/server/src/main/java/com/ke/bella/openapi/protocol/log/MetricsLogHandler.java) |
| **CostLogHandler** | Evaluates monetary cost using a pluggable `CostScripFetcher` to obtain price expressions, calculates via `CostCalculator`, and persists via `CostCounter`. | [`api/server/src/main/java/com/ke/bella/openapi/protocol/log/CostLogHandler.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/server/src/main/java/com/ke/bella/openapi/protocol/log/CostLogHandler.java) |

### Disruptor Ring Buffer Configuration

The pipeline guarantees processing order through explicit handler chaining configured in `BellaAutoConf`. The `CostLogHandler` executes first to ensure cost recording occurs before any throttling decisions:

```java
disruptor.handleEventsWith(
    new CostLogHandler(costCounter, costScripFetcher)
).then(
    new MetricsLogHandler(metricsManager),
    new LimiterLogHandler(limiterManager)
);

```

## Metrics Collection via MetricsLogHandler

The `MetricsLogHandler` processes lightweight usage data for real-time monitoring and rate-limiting dashboards.

When a `LogEvent` arrives, the handler extracts the embedded `RequestMetrics` object (defined in [`api/sdk/src/main/java/com/ke/bella/openapi/RequestMetrics.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/sdk/src/main/java/com/ke/bella/openapi/RequestMetrics.java)) containing:

- **`inputTokens`** and **`outputTokens`**
- **`latencyMs`** and **`responseSize`**

The handler delegates to `MetricsManager.record(requestMetrics)`, which aggregates counters using a **Redis Lua-script-based sliding-window** implementation (located in `api/server/src/main/resources/lua/`). This approach ensures atomic counter increments for high-precision rate limiting and usage dashboards without database round-trips.

## Cost Reporting via CostLogHandler

The `CostLogHandler` implements a dynamic pricing engine that calculates monetary costs based on configurable price expressions and actual resource consumption.

### Price Script Fetching

The handler uses `CostScripFetcher`, a functional interface defined in `BellaAutoConf`, to retrieve pricing expressions for the current endpoint. The fetcher queries the `EndpointService` to obtain JSON pricing configurations. A typical script for token-based pricing looks like:

```java
endpoint -> "price.input * usage.input_tokens + price.output * usage.output_tokens";

```

### Usage Object Evaluation

Each protocol adapter (such as `CompletionLogHandler` or `TtsLogHandler`) creates a specific **usage subclass** implementing the `Usage` marker interface. The `CostLogHandler` extracts this object from the `LogEvent` and passes it to the calculation engine.

### Cost Calculation and Persistence

The `CostCalculator` (located in [`api/server/src/main/java/com/ke/bella/openapi/protocol/cost/CostCalculator.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/server/src/main/java/com/ke/bella/openapi/protocol/cost/CostCalculator.java)) parses the price script using a custom expression evaluator supporting arithmetic operations (`+`, `*`) and property navigation. It multiplies token counts or tool invocation counts by the configured price per-thousand-tokens or per-tool fees.

The resulting `BigDecimal` is handed to `CostCounter.increment(tenantId, cost)`, which:

1. Persists the per-tenant cost to the **cost table** in MySQL
2. Updates an in-memory cache for immediate quota enforcement

```java
costCounter.increase(tenantId, cost);

```

## End-to-End Example: Chat Completion Flow

The following example from `CompletionLogHandler` demonstrates how a chat completion request flows through the endpoint logging system:

```java
// Inside CompletionLogHandler
RequestMetrics metrics = RequestMetrics.builder()
        .inputTokens(requestTokens)
        .outputTokens(responseTokens)
        .latencyMs(duration)
        .build();

LogEvent log = LogEvent.builder()
        .endpoint("/v1/chat/completions")
        .requestMetrics(metrics)
        .usage(new CompletionLogHandler.CompletionUsage(inputTokens, outputTokens))
        .build();

logRingBuffer.publish(log);

```

Once published to the ring buffer:

1. `CostLogHandler` receives the event, fetches the pricing script `"price.input * usage.input_tokens + price.output * usage.output_tokens"`, evaluates it against the `CompletionUsage` object, and records the charge via `CostCounter`.
2. `MetricsLogHandler` extracts the `RequestMetrics` and updates the Redis-based sliding window counters for real-time dashboards.
3. `LimiterLogHandler` performs final rate-limit checks using the updated metrics.

## Summary

- The endpoint logging system in Bella OpenAPI uses a **LMAX Disruptor** ring buffer to process requests through a chain of specialized handlers without blocking the critical path.
- **`MetricsLogHandler`** extracts `RequestMetrics` to update Redis/Caffeine counters for real-time monitoring and rate-limiting dashboards.
- **`CostLogHandler`** dynamically evaluates pricing expressions fetched by `CostScripFetcher` using `CostCalculator`, then persists charges via `CostCounter` to MySQL and in-memory quota caches.
- The pipeline ordering guarantees cost recording occurs before rate-limiting decisions, ensuring accurate billing even for throttled requests.
- Protocol-specific adapters like `CompletionLogHandler` create usage subclasses that enable granular cost calculation for different AI model types.

## Frequently Asked Questions

### How does Bella OpenAPI ensure accurate cost reporting for throttled requests?

The Disruptor pipeline explicitly orders handlers so that `CostLogHandler` executes before `LimiterLogHandler`. This guarantees that `CostCounter.increment(tenantId, cost)` persists the charge to MySQL before any rate-limiting decision might abort the request, ensuring every billed request is recorded regardless of throttling outcomes.

### What data structure carries usage information through the logging pipeline?

The `RequestMetrics` class defined in [`api/sdk/src/main/java/com/ke/bella/openapi/RequestMetrics.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/sdk/src/main/java/com/ke/bella/openapi/RequestMetrics.java) carries lightweight usage data including `inputTokens`, `outputTokens`, `latencyMs`, and `responseSize`. Additionally, protocol-specific usage subclasses implementing the `Usage` marker interface carry detailed consumption data required for cost calculation.

### How does the system support different pricing models for various AI endpoints?

The `CostScripFetcher` functional interface retrieves endpoint-specific pricing expressions from the `EndpointService`. For example, chat completions might use `"price.input * usage.input_tokens + price.output * usage.output_tokens"` while TTS endpoints could use per-character or per-second pricing. The `CostCalculator` evaluates these expressions against the concrete `Usage` object provided by the protocol adapter.

### Where are the aggregated metrics stored for real-time dashboards?

`MetricsManager` stores aggregated counters in **Redis** using Lua-script-based sliding windows (located in `api/server/src/main/resources/lua/`) and **Caffeine** in-memory caches. This dual-layer approach provides high-throughput writes for the logging pipeline and low-latency reads for admin dashboards and rate-limiting decisions.