How Read Frog's Batch Request System Optimizes AI API Calls and Reduces Costs
Read Frog groups multiple small translation requests into single batched API calls, dramatically reducing token overhead and per-request fees while maintaining reliability through intelligent retries and fallback mechanisms.
The batch request system in the mengxi-ream/read-frog repository is the core infrastructure that enables cost-effective AI translation. By buffering individual translation tasks and consolidating them into shared API calls, the system minimizes redundant network overhead and maximizes throughput. This article examines the internal implementation across src/utils/request/batch-queue.ts and src/entrypoints/background/translation-queues.ts to explain exactly how Read Frog achieves these optimizations.
Core Architecture of the Batch Request System
Read Frog implements batching through two coordinated modules. The BatchQueue class in src/utils/request/batch-queue.ts handles the buffering lifecycle, while src/entrypoints/background/translation-queues.ts provides the domain-specific glue that instantiates queues for each language-provider combination.
The system uses a batch key derived from source language, target language, and provider ID to segregate compatible requests. This ensures that all tasks in a single batch share identical translation parameters, preventing context contamination while allowing parallel processing across different language pairs.
How the BatchQueue Class Buffers Translation Requests
Task Enqueueing and Batch Key Generation
When a UI component requests translation, the background script creates a BatchTask and enqueues it via batchQueue.enqueue(). The queue groups tasks using the getBatchKey function, which generates a SHA-256 hash of the language configuration and provider ID:
const batchQueue = new BatchQueue<TranslateBatchData, string>({
maxCharactersPerBatch,
maxItemsPerBatch,
batchDelay: 100,
maxRetries: 3,
enableFallbackToIndividual: true,
getBatchKey: data => Sha256Hex(
`${data.langConfig.sourceCode}-${data.langConfig.targetCode}-${data.providerConfig.id}`
),
getCharacters: data => data.text.length,
executeBatch,
executeIndividual,
onError,
});
The getCharacters callback enables precise enforcement of provider-specific token limits by measuring the character count of each incoming text. This prevents individual batches from exceeding API payload constraints.
Flush Triggers and Timing Controls
The BatchQueue evaluates three conditions to determine when to dispatch a batch, implemented in schedule() and shouldFlushBatch():
- Item threshold –
batch.tasks.length >= maxItemsPerBatch - Character threshold –
batch.totalCharacters >= maxCharactersPerBatch - Delay timeout –
now >= batch.createdAt + batchDelay
If any condition triggers, flushPendingBatchByKey() executes the batch immediately. The default batchDelay of 100 milliseconds balances latency against batch density, ensuring rapid response for single requests while capturing burst traffic for optimization.
Executing Batched AI Translation Requests
Constructing the Combined Prompt
The executeBatch function in translation-queues.ts concatenates multiple texts using the BATCH_SEPARATOR constant (\n\n---\n\n) defined in src/utils/constants/prompt.ts:
const batchText = texts.join(`\n\n${BATCH_SEPARATOR}\n\n`);
const batchThunk = async (): Promise<string[]> => {
await putBatchRequestRecord({
originalRequestCount: dataList.length,
providerConfig
});
const result = await executeTranslate(
batchText, langConfig, providerConfig, promptResolver,
{ isBatch: true, content }
);
return parseBatchResult(result);
};
This separator serves a dual purpose: it delimits inputs for the LLM and provides a deterministic split point for parsing the response back into individual results via parseBatchResult().
Rate-Limited Request Handling
Before reaching the network layer, batched requests pass through a RequestQueue (src/utils/request/request-queue.ts). This secondary queue implements rate limiting and additional retry logic for the actual HTTP transport, preventing API quota exhaustion and handling transient network failures independently of batching logic.
Error Handling, Retries, and Fallback Strategies
Handling Batch Count Mismatches
When an LLM returns an incorrect number of translated segments, the system throws a BatchCountMismatchError. The executeBatchWithRetry method captures this specific failure mode and initiates recovery:
if (retryCount < this.maxRetries && err instanceof BatchCountMismatchError) {
const delay = this.calculateBackoffDelay(retryCount);
await this.sleep(delay);
return this.executeBatchWithRetry(tasks, batchKey, retryCount + 1);
}
The calculateBackoffDelay function implements exponential backoff, increasing wait intervals between retry attempts to avoid overwhelming the API during service degradation.
Exponential Backoff and Individual Fallback
If the batch exhausts its maxRetries (default: 3) or encounters a non-count-related error, the system evaluates enableFallbackToIndividual. When enabled, the queue decomposes the failed batch into separate executeIndividual calls, ensuring that a single problematic text does not block valid translations. The onError callback in translation-queues.ts logs the batch key, retry count, and failure mode for diagnostic visibility.
Runtime Configuration for Cost Optimization
Read Frog exposes dynamic tuning through setBatchConfig, allowing users to optimize the latency-cost trade-off without restarting the extension. The configuration interface accepts:
- maxCharactersPerBatch – Caps total characters to respect provider token limits
- maxItemsPerBatch – Limits the number of discrete translations per API call
- batchDelay – Adjusts the buffering window for batch accumulation
These parameters propagate via message handlers like setTranslateBatchQueueConfig and setSubtitlesBatchQueueConfig, enabling real-time adjustments based on network conditions or cost constraints.
Summary
- BatchQueue in
src/utils/request/batch-queue.tsmanages the complete lifecycle of request aggregation, from buffering to flush execution. - Batch keys group compatible requests by language pair and provider, while
getCharactersenforces payload size limits. - Three triggers (item count, character count, timeout) determine when batches flush, balancing speed against efficiency.
- BATCH_SEPARATOR (
\n\n---\n\n) enables reliable concatenation and parsing of multiple translations in a single LLM response. - Intelligent retries with exponential backoff handle
BatchCountMismatchError, with optional fallback to individual requests ensuring reliability. - Runtime configurability via
setBatchConfigallows users to tunemaxItemsPerBatchandbatchDelayfor specific cost or latency requirements.
Frequently Asked Questions
What is the primary benefit of Read Frog's batch request system?
The system reduces AI API costs by consolidating multiple small translation requests into single batched calls. This approach minimizes per-request overhead charges and reduces redundant token usage from repeated system prompts, while the BATCH_SEPARATOR ensures accurate result demultiplexing.
How does Read Frog handle failed batch requests?
The BatchQueue implements exponential backoff for transient failures like BatchCountMismatchError. After exhausting maxRetries, the system can automatically fall back to individual requests when enableFallbackToIndividual is true, ensuring partial success rather than total batch failure.
Can users adjust batch size limits in real-time?
Yes. Through message handlers such as setTranslateBatchQueueConfig, users can dynamically update maxCharactersPerBatch, maxItemsPerBatch, and batchDelay via the setBatchConfig method, allowing optimization for different network conditions or provider pricing models without extension restarts.
What separator does Read Frog use to split batched translations?
The system uses the BATCH_SEPARATOR constant defined in src/utils/constants/prompt.ts, which contains the string \n\n---\n\n. This delimiter joins input texts for the LLM and provides the split pattern for parseBatchResult() to separate the combined response back into individual translation results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →