# How Read Frog's Batch Request System Optimizes AI API Calls and Reduces Costs

> Optimize AI API calls and cut costs with Read Frog's batch request system. Learn how it groups requests to reduce overhead and fees while ensuring reliability.

- Repository: [MengXi/read-frog](https://github.com/mengxi-ream/read-frog)
- Tags: internals
- Published: 2026-03-07

---

**Read Frog groups multiple small translation requests into single batched API calls, dramatically reducing token overhead and per-request fees while maintaining reliability through intelligent retries and fallback mechanisms.**

The **batch request system** in the [mengxi-ream/read-frog](https://github.com/mengxi-ream/read-frog) repository is the core infrastructure that enables cost-effective AI translation. By buffering individual translation tasks and consolidating them into shared API calls, the system minimizes redundant network overhead and maximizes throughput. This article examines the internal implementation across [`src/utils/request/batch-queue.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts) and [`src/entrypoints/background/translation-queues.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/background/translation-queues.ts) to explain exactly how Read Frog achieves these optimizations.

## Core Architecture of the Batch Request System

Read Frog implements batching through two coordinated modules. The **BatchQueue** class in [`src/utils/request/batch-queue.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts) handles the buffering lifecycle, while [`src/entrypoints/background/translation-queues.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/background/translation-queues.ts) provides the domain-specific glue that instantiates queues for each language-provider combination.

The system uses a **batch key** derived from source language, target language, and provider ID to segregate compatible requests. This ensures that all tasks in a single batch share identical translation parameters, preventing context contamination while allowing parallel processing across different language pairs.

## How the BatchQueue Class Buffers Translation Requests

### Task Enqueueing and Batch Key Generation

When a UI component requests translation, the background script creates a `BatchTask` and enqueues it via `batchQueue.enqueue()`. The queue groups tasks using the `getBatchKey` function, which generates a SHA-256 hash of the language configuration and provider ID:

```typescript
const batchQueue = new BatchQueue<TranslateBatchData, string>({
  maxCharactersPerBatch,
  maxItemsPerBatch,
  batchDelay: 100,
  maxRetries: 3,
  enableFallbackToIndividual: true,
  getBatchKey: data => Sha256Hex(
    `${data.langConfig.sourceCode}-${data.langConfig.targetCode}-${data.providerConfig.id}`
  ),
  getCharacters: data => data.text.length,
  executeBatch,
  executeIndividual,
  onError,
});

```

The `getCharacters` callback enables precise enforcement of provider-specific token limits by measuring the character count of each incoming text. This prevents individual batches from exceeding API payload constraints.

### Flush Triggers and Timing Controls

The `BatchQueue` evaluates three conditions to determine when to dispatch a batch, implemented in `schedule()` and `shouldFlushBatch()`:

1. **Item threshold** – `batch.tasks.length >= maxItemsPerBatch`
2. **Character threshold** – `batch.totalCharacters >= maxCharactersPerBatch`
3. **Delay timeout** – `now >= batch.createdAt + batchDelay`

If any condition triggers, `flushPendingBatchByKey()` executes the batch immediately. The default `batchDelay` of 100 milliseconds balances latency against batch density, ensuring rapid response for single requests while capturing burst traffic for optimization.

## Executing Batched AI Translation Requests

### Constructing the Combined Prompt

The `executeBatch` function in [`translation-queues.ts`](https://github.com/mengxi-ream/read-frog/blob/main/translation-queues.ts) concatenates multiple texts using the `BATCH_SEPARATOR` constant (`\n\n---\n\n`) defined in [`src/utils/constants/prompt.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/constants/prompt.ts):

```typescript
const batchText = texts.join(`\n\n${BATCH_SEPARATOR}\n\n`);
const batchThunk = async (): Promise<string[]> => {
  await putBatchRequestRecord({ 
    originalRequestCount: dataList.length, 
    providerConfig 
  });
  const result = await executeTranslate(
    batchText, langConfig, providerConfig, promptResolver,
    { isBatch: true, content }
  );
  return parseBatchResult(result);
};

```

This separator serves a dual purpose: it delimits inputs for the LLM and provides a deterministic split point for parsing the response back into individual results via `parseBatchResult()`.

### Rate-Limited Request Handling

Before reaching the network layer, batched requests pass through a **RequestQueue** ([`src/utils/request/request-queue.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/request-queue.ts)). This secondary queue implements rate limiting and additional retry logic for the actual HTTP transport, preventing API quota exhaustion and handling transient network failures independently of batching logic.

## Error Handling, Retries, and Fallback Strategies

### Handling Batch Count Mismatches

When an LLM returns an incorrect number of translated segments, the system throws a `BatchCountMismatchError`. The `executeBatchWithRetry` method captures this specific failure mode and initiates recovery:

```typescript
if (retryCount < this.maxRetries && err instanceof BatchCountMismatchError) {
  const delay = this.calculateBackoffDelay(retryCount);
  await this.sleep(delay);
  return this.executeBatchWithRetry(tasks, batchKey, retryCount + 1);
}

```

The `calculateBackoffDelay` function implements exponential backoff, increasing wait intervals between retry attempts to avoid overwhelming the API during service degradation.

### Exponential Backoff and Individual Fallback

If the batch exhausts its `maxRetries` (default: 3) or encounters a non-count-related error, the system evaluates `enableFallbackToIndividual`. When enabled, the queue decomposes the failed batch into separate `executeIndividual` calls, ensuring that a single problematic text does not block valid translations. The `onError` callback in [`translation-queues.ts`](https://github.com/mengxi-ream/read-frog/blob/main/translation-queues.ts) logs the batch key, retry count, and failure mode for diagnostic visibility.

## Runtime Configuration for Cost Optimization

Read Frog exposes dynamic tuning through `setBatchConfig`, allowing users to optimize the latency-cost trade-off without restarting the extension. The configuration interface accepts:

- **maxCharactersPerBatch** – Caps total characters to respect provider token limits
- **maxItemsPerBatch** – Limits the number of discrete translations per API call
- **batchDelay** – Adjusts the buffering window for batch accumulation

These parameters propagate via message handlers like `setTranslateBatchQueueConfig` and `setSubtitlesBatchQueueConfig`, enabling real-time adjustments based on network conditions or cost constraints.

## Summary

- **BatchQueue** in [`src/utils/request/batch-queue.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts) manages the complete lifecycle of request aggregation, from buffering to flush execution.
- **Batch keys** group compatible requests by language pair and provider, while `getCharacters` enforces payload size limits.
- **Three triggers** (item count, character count, timeout) determine when batches flush, balancing speed against efficiency.
- **BATCH_SEPARATOR** (`\n\n---\n\n`) enables reliable concatenation and parsing of multiple translations in a single LLM response.
- **Intelligent retries** with exponential backoff handle `BatchCountMismatchError`, with optional fallback to individual requests ensuring reliability.
- **Runtime configurability** via `setBatchConfig` allows users to tune `maxItemsPerBatch` and `batchDelay` for specific cost or latency requirements.

## Frequently Asked Questions

### What is the primary benefit of Read Frog's batch request system?

The system reduces AI API costs by consolidating multiple small translation requests into single batched calls. This approach minimizes per-request overhead charges and reduces redundant token usage from repeated system prompts, while the `BATCH_SEPARATOR` ensures accurate result demultiplexing.

### How does Read Frog handle failed batch requests?

The `BatchQueue` implements exponential backoff for transient failures like `BatchCountMismatchError`. After exhausting `maxRetries`, the system can automatically fall back to individual requests when `enableFallbackToIndividual` is true, ensuring partial success rather than total batch failure.

### Can users adjust batch size limits in real-time?

Yes. Through message handlers such as `setTranslateBatchQueueConfig`, users can dynamically update `maxCharactersPerBatch`, `maxItemsPerBatch`, and `batchDelay` via the `setBatchConfig` method, allowing optimization for different network conditions or provider pricing models without extension restarts.

### What separator does Read Frog use to split batched translations?

The system uses the `BATCH_SEPARATOR` constant defined in [`src/utils/constants/prompt.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/constants/prompt.ts), which contains the string `\n\n---\n\n`. This delimiter joins input texts for the LLM and provides the split pattern for `parseBatchResult()` to separate the combined response back into individual translation results.