Read Frog Batch Request Processing Pipeline: Optimization Techniques and Performance Tuning

Tune batch size limits, flush delays, and token-bucket rates while leveraging exponential backoff with jitter and automatic request deduplication to maximize throughput and minimize latency in Read Frog's translation system.

Read Frog's translation engine batches individual text translation calls into efficient bulk operations through a sophisticated dual-queue architecture. Understanding how to configure the BatchQueue and RequestQueue components in src/utils/request/batch-queue.ts and src/utils/request/request-queue.ts allows you to optimize the batch request processing pipeline for high throughput, low latency, and resilient error handling.

Architecture Overview

The pipeline consists of two cooperating components that manage the lifecycle of translation requests:

  • BatchQueue: Groups tasks by a batch key, enforces per-batch character and item limits, and triggers flushes after a configurable delay or when limits are hit. It handles retry logic for count mismatches and optional fallback to individual requests.
  • RequestQueue: Enforces token-bucket rate limiting, timeout handling (timeoutMs), and exponential-backoff retries with jitter for each flushed batch or individual fallback request.

In src/utils/request/batch-queue.ts, the constructor parses configuration at lines 53-61, while the schedule method (lines 83-101) determines when to flush batches. The calculateBackoffDelay function (lines 17-19) manages retry timing. The RequestQueue in src/utils/request/request-queue.ts implements token-bucket fields at lines 30-33 (bucketTokens, lastRefill), with the scheduling loop at lines 87-112 and retry logic in executeTask (lines 63-80).

Performance-Critical Configuration Parameters

Tuning these knobs is the primary method for optimizing throughput versus latency:

Parameter Function Trade-off Consideration
maxCharactersPerBatch Upper bound on total characters per batch Larger batches reduce HTTP round-trips but risk provider payload limits
maxItemsPerBatch Max number of tasks per batch 独立于 payload size 的并发控制
batchDelay Milliseconds to wait before flushing partially-filled batches Small values reduce latency; larger values improve batch utilization
rate / capacity (RequestQueue) Token-bucket refill rate and burst size Higher values increase throughput but may violate external API rate limits
timeoutMs Maximum request duration before abort Short timeouts enable faster failure detection but risk false positives on slow networks
maxRetries / baseRetryDelayMs Retry attempts and initial back-off interval More retries improve reliability; exponential back-off prevents API hammering
enableFallbackToIndividual Toggle for per-item fallback after batch failure Guarantees progress for stubborn LLM responses at the cost of additional API calls

Core Optimization Techniques

Match Batch Sizes to Provider Limits

Most translation APIs impose maximum payload restrictions (e.g., 5,000 characters). Setting maxCharactersPerBatch slightly below this threshold maximizes payload efficiency while avoiding 413 Payload Too Large rejections.

// Example: 4,800 chars per batch provides a safe margin
const batchQueue = new BatchQueue<string, string>({
  maxCharactersPerBatch: 4800,
  maxItemsPerBatch: 10,
  batchDelay: 300,               // 0.3s latency budget
  getBatchKey: () => 'default',
  getCharacters: (txt) => txt.length,
  executeBatch: async (texts) => translateApi.bulk(texts),
  executeIndividual: async (txt) => translateApi.single(txt),
});

Source: Constructor implementation at src/utils/request/batch-queue.ts#L53-L61.

Balance Flush Delay and Latency

The batchDelay parameter directly impacts the latency-throughput trade-off. A short delay (e.g., 100ms) yields near real-time responses for sparse traffic but creates many small batches. A longer delay (e.g., 500ms) dramatically improves batch fill rates during high-traffic periods.

Rule of thumb: Configure batchDelay to approximately ½ × the average inter-arrival time of translation requests.

Source: Scheduling logic at src/utils/request/batch-queue.ts#L83-L101.

Implement Token-Bucket Rate Limiting

RequestQueue decouples batch flushing from API rate limits using a token-bucket algorithm. Adjusting rate (tokens per second) and capacity (burst allowance) smooths traffic bursts without exceeding provider quotas.

const requestQueue = new RequestQueue({
  rate: 5,               // 5 tokens per second sustained
  capacity: 10,          // burst up to 10 concurrent requests
  timeoutMs: 30_000,
  maxRetries: 3,
  baseRetryDelayMs: 500,
});

Source: Token bucket implementation at src/utils/request/request-queue.ts#L30-L33.

Configure Exponential Back-off with Jitter

Both queues implement exponential back-off to handle transient failures. RequestQueue adds a random jitter (0–10%) to prevent thundering-herd effects when many requests retry simultaneously:

// Back-off formula with jitter (see executeTask implementation)
const backoffDelayMs = this.options.baseRetryDelayMs * (2 ** (task.retryCount - 1));
const jitter = Math.random() * 0.1 * backoffDelayMs;

BatchQueue retries only on BatchCountMismatchError, utilizing fixed bounds (BASE_BACKOFF_DELAY_MS = 1000, MAX_BACKOFF_DELAY_MS = 8000) to prevent infinite loops while allowing LLM recovery time.

Source: Back-off calculation at src/utils/request/batch-queue.ts#L17-L19 and retry logic at src/utils/request/request-queue.ts#L63-L71.

Enable Fallback for Malformed Batch Responses

When a batch consistently returns mismatched result counts (e.g., the LLM returns fewer translations than requested), enabling enableFallbackToIndividual ensures request completion by splitting the batch into individual API calls. This sacrifices efficiency for reliability when provider batch endpoints are flaky.

const batchQueue = new BatchQueue(..., {
  enableFallbackToIndividual: true,
  executeIndividual: async (txt) => translateApi.single(txt),
});

Source: Fallback implementation at src/utils/request/batch-queue.ts#L98-L115.

Leverage Automatic Request Deduplication

RequestQueue.enqueue automatically detects duplicate payloads via hashing. When a task with an identical hash is already pending, the system returns the existing promise instead of creating redundant work:

const promise = requestQueue.enqueue(thunk, Date.now(), payloadHash);

Source: Duplicate detection logic at src/utils/request/request-queue.ts#L2-L8.

Apply Dynamic Runtime Configuration

Both queues expose mutator methods—BatchQueue.setBatchConfig (lines 25-33) and RequestQueue.setQueueOptions (lines 75-81)—that re-validate parameters using Zod schemas. This enables runtime tuning (e.g., increasing rate during off-peak hours) without requiring application restarts.

Practical Implementation Example

This complete example demonstrates wiring the two queues together for a production translation pipeline:

import { RequestQueue } from "./src/utils/request/request-queue";
import { BatchQueue } from "./src/utils/request/batch-queue";

// 1️⃣ Create the low-level request queue (rate-limit, retries)
const requestQueue = new RequestQueue({
  rate: 10,               // 10 calls/second
  capacity: 20,           // allow short bursts
  timeoutMs: 20_000,
  maxRetries: 4,
  baseRetryDelayMs: 300,
});

// 2️⃣ Helper that enqueues a thunk into the RequestQueue
function enqueueRequest<T>(thunk: () => Promise<T>, hash: string) {
  // Schedule immediately; RequestQueue will handle rate-limits & retries
  return requestQueue.enqueue(thunk, Date.now(), hash);
}

// 3️⃣ Build a BatchQueue that uses the enqueue helper
const batchQueue = new BatchQueue<string, string>({
  maxCharactersPerBatch: 4000,
  maxItemsPerBatch: 8,
  batchDelay: 250,
  getBatchKey: () => "translate-default",
  getCharacters: (txt) => txt.length,
  // Flush a batch by sending a single request through RequestQueue
  executeBatch: async (texts) => {
    const hash = "batch-" + texts.join("|");
    return enqueueRequest(() => translateApi.bulk(texts), hash);
  },
  // Optional per-item fallback
  executeIndividual: async (txt) => {
    const hash = "single-" + txt;
    return enqueueRequest(() => translateApi.single(txt), hash);
  },
});

Key integration points:

  • Rate limiting is centralized in RequestQueue; all batch and individual calls share the same token bucket.
  • Retries operate at both layers: RequestQueue retries any failing HTTP call, while BatchQueue retries specifically when the LLM returns incorrect result counts.

Summary

  • Size batches appropriately: Set maxCharactersPerBatch just below provider limits (e.g., 4,800 for a 5,000-character cap) to maximize payload efficiency.
  • Tune flush timing: Configure batchDelay to approximately half your average request inter-arrival time to balance latency and batch utilization.
  • Respect rate limits: Use the token-bucket parameters (rate, capacity) in RequestQueue to smooth bursts without exceeding API quotas.
  • Handle failures gracefully: Implement exponential back-off with jitter via baseRetryDelayMs and enable enableFallbackToIndividual for stubborn batch failures.
  • Eliminate redundant work: Leverage built-in request deduplication via payload hashing in RequestQueue.enqueue.
  • Adjust without restarts: Use setBatchConfig and setQueueOptions for dynamic pipeline tuning in production environments.

Frequently Asked Questions

What triggers a BatchCountMismatchError and how does the pipeline handle it?

A BatchCountMismatchError occurs when the translation API returns a different number of results than the number of tasks sent in the batch (defined at src/utils/request/batch-queue.ts#L1-L7). When this error is detected, BatchQueue automatically retries the batch using exponential back-off with delays capped between 1,000ms and 8,000ms. If retries exhaust and enableFallbackToIndividual is true, the system falls back to processing each text segment individually through executeIndividual.

How does the token-bucket algorithm prevent API quota violations?

The RequestQueue implements a token-bucket rate limiter using bucketTokens and lastRefill fields (src/utils/request/request-queue.ts#L30-L33). Tokens are consumed for each request and refilled at the configured rate (tokens per second) up to capacity. This smooths traffic bursts by queueing requests when tokens are depleted, ensuring the average request rate never exceeds the provider's allowed quota while still permitting short bursts up to the capacity limit.

When should I enable fallback to individual requests?

Enable enableFallbackToIndividual when your translation provider's batch endpoint exhibits instability or returns inconsistent result counts. This is particularly valuable for large language models that occasionally return fewer translations than requested due to content filtering or context window limitations. While this increases API call volume, it guarantees request completion rather than failing the entire batch permanently.

How do I calculate the optimal batchDelay for my workload?

Calculate batchDelay as approximately 50% of your average request inter-arrival time. For example, if your application receives translation requests every 200ms on average, set batchDelay to 100ms. This heuristic minimizes latency for individual requests while maximizing the probability of accumulating multiple requests into efficient batches. Monitor your batch fill rates in production and increase the delay if batches are consistently under-filled, or decrease it if latency becomes unacceptable.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →