# Read Frog Batch Request Processing Pipeline: Optimization Techniques and Performance Tuning

> Optimize Read Frog's batch request pipeline by tuning batch size, flush delays, and rates. Implement exponential backoff jitter and auto deduplication for peak performance and low latency.

- Repository: [MengXi/read-frog](https://github.com/mengxi-ream/read-frog)
- Tags: performance
- Published: 2026-03-07

---

**Tune batch size limits, flush delays, and token-bucket rates while leveraging exponential backoff with jitter and automatic request deduplication to maximize throughput and minimize latency in Read Frog's translation system.**

Read Frog's translation engine batches individual text translation calls into efficient bulk operations through a sophisticated dual-queue architecture. Understanding how to configure the **BatchQueue** and **RequestQueue** components in [`src/utils/request/batch-queue.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts) and [`src/utils/request/request-queue.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/request-queue.ts) allows you to optimize the batch request processing pipeline for high throughput, low latency, and resilient error handling.

## Architecture Overview

The pipeline consists of two cooperating components that manage the lifecycle of translation requests:

- **BatchQueue**: Groups tasks by a *batch key*, enforces per-batch character and item limits, and triggers flushes after a configurable delay or when limits are hit. It handles retry logic for count mismatches and optional fallback to individual requests.
- **RequestQueue**: Enforces token-bucket rate limiting, timeout handling (`timeoutMs`), and exponential-backoff retries with jitter for each flushed batch or individual fallback request.

In [`src/utils/request/batch-queue.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts), the constructor parses configuration at lines 53-61, while the `schedule` method (lines 83-101) determines when to flush batches. The `calculateBackoffDelay` function (lines 17-19) manages retry timing. The `RequestQueue` in [`src/utils/request/request-queue.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/request-queue.ts) implements token-bucket fields at lines 30-33 (`bucketTokens`, `lastRefill`), with the scheduling loop at lines 87-112 and retry logic in `executeTask` (lines 63-80).

## Performance-Critical Configuration Parameters

Tuning these knobs is the primary method for optimizing throughput versus latency:

| Parameter | Function | Trade-off Consideration |
|-----------|----------|-------------------------|
| **maxCharactersPerBatch** | Upper bound on total characters per batch | Larger batches reduce HTTP round-trips but risk provider payload limits |
| **maxItemsPerBatch** | Max number of tasks per batch |独立于 payload size 的并发控制 |
| **batchDelay** | Milliseconds to wait before flushing partially-filled batches | Small values reduce latency; larger values improve batch utilization |
| **rate** / **capacity** (RequestQueue) | Token-bucket refill rate and burst size | Higher values increase throughput but may violate external API rate limits |
| **timeoutMs** | Maximum request duration before abort | Short timeouts enable faster failure detection but risk false positives on slow networks |
| **maxRetries** / **baseRetryDelayMs** | Retry attempts and initial back-off interval | More retries improve reliability; exponential back-off prevents API hammering |
| **enableFallbackToIndividual** | Toggle for per-item fallback after batch failure | Guarantees progress for stubborn LLM responses at the cost of additional API calls |

## Core Optimization Techniques

### Match Batch Sizes to Provider Limits

Most translation APIs impose maximum payload restrictions (e.g., 5,000 characters). Setting `maxCharactersPerBatch` slightly below this threshold maximizes payload efficiency while avoiding `413 Payload Too Large` rejections.

```typescript
// Example: 4,800 chars per batch provides a safe margin
const batchQueue = new BatchQueue<string, string>({
  maxCharactersPerBatch: 4800,
  maxItemsPerBatch: 10,
  batchDelay: 300,               // 0.3s latency budget
  getBatchKey: () => 'default',
  getCharacters: (txt) => txt.length,
  executeBatch: async (texts) => translateApi.bulk(texts),
  executeIndividual: async (txt) => translateApi.single(txt),
});

```

*Source*: Constructor implementation at [`src/utils/request/batch-queue.ts#L53-L61`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts#L53-L61).

### Balance Flush Delay and Latency

The `batchDelay` parameter directly impacts the latency-throughput trade-off. A short delay (e.g., 100ms) yields near real-time responses for sparse traffic but creates many small batches. A longer delay (e.g., 500ms) dramatically improves batch fill rates during high-traffic periods.

**Rule of thumb**: Configure `batchDelay` to approximately ½ × the average inter-arrival time of translation requests.

*Source*: Scheduling logic at [`src/utils/request/batch-queue.ts#L83-L101`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts#L83-L101).

### Implement Token-Bucket Rate Limiting

`RequestQueue` decouples batch flushing from API rate limits using a token-bucket algorithm. Adjusting `rate` (tokens per second) and `capacity` (burst allowance) smooths traffic bursts without exceeding provider quotas.

```typescript
const requestQueue = new RequestQueue({
  rate: 5,               // 5 tokens per second sustained
  capacity: 10,          // burst up to 10 concurrent requests
  timeoutMs: 30_000,
  maxRetries: 3,
  baseRetryDelayMs: 500,
});

```

*Source*: Token bucket implementation at [`src/utils/request/request-queue.ts#L30-L33`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/request-queue.ts#L30-L33).

### Configure Exponential Back-off with Jitter

Both queues implement exponential back-off to handle transient failures. `RequestQueue` adds a random jitter (0–10%) to prevent thundering-herd effects when many requests retry simultaneously:

```typescript
// Back-off formula with jitter (see executeTask implementation)
const backoffDelayMs = this.options.baseRetryDelayMs * (2 ** (task.retryCount - 1));
const jitter = Math.random() * 0.1 * backoffDelayMs;

```

`BatchQueue` retries only on `BatchCountMismatchError`, utilizing fixed bounds (`BASE_BACKOFF_DELAY_MS = 1000`, `MAX_BACKOFF_DELAY_MS = 8000`) to prevent infinite loops while allowing LLM recovery time.

*Source*: Back-off calculation at [`src/utils/request/batch-queue.ts#L17-L19`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts#L17-L19) and retry logic at [`src/utils/request/request-queue.ts#L63-L71`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/request-queue.ts#L63-L71).

### Enable Fallback for Malformed Batch Responses

When a batch consistently returns mismatched result counts (e.g., the LLM returns fewer translations than requested), enabling `enableFallbackToIndividual` ensures request completion by splitting the batch into individual API calls. This sacrifices efficiency for reliability when provider batch endpoints are flaky.

```typescript
const batchQueue = new BatchQueue(..., {
  enableFallbackToIndividual: true,
  executeIndividual: async (txt) => translateApi.single(txt),
});

```

*Source*: Fallback implementation at [`src/utils/request/batch-queue.ts#L98-L115`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts#L98-L115).

### Leverage Automatic Request Deduplication

`RequestQueue.enqueue` automatically detects duplicate payloads via hashing. When a task with an identical hash is already pending, the system returns the existing promise instead of creating redundant work:

```typescript
const promise = requestQueue.enqueue(thunk, Date.now(), payloadHash);

```

*Source*: Duplicate detection logic at [`src/utils/request/request-queue.ts#L2-L8`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/request-queue.ts#L2-L8).

### Apply Dynamic Runtime Configuration

Both queues expose mutator methods—`BatchQueue.setBatchConfig` (lines 25-33) and `RequestQueue.setQueueOptions` (lines 75-81)—that re-validate parameters using Zod schemas. This enables runtime tuning (e.g., increasing `rate` during off-peak hours) without requiring application restarts.

## Practical Implementation Example

This complete example demonstrates wiring the two queues together for a production translation pipeline:

```typescript
import { RequestQueue } from "./src/utils/request/request-queue";
import { BatchQueue } from "./src/utils/request/batch-queue";

// 1️⃣ Create the low-level request queue (rate-limit, retries)
const requestQueue = new RequestQueue({
  rate: 10,               // 10 calls/second
  capacity: 20,           // allow short bursts
  timeoutMs: 20_000,
  maxRetries: 4,
  baseRetryDelayMs: 300,
});

// 2️⃣ Helper that enqueues a thunk into the RequestQueue
function enqueueRequest<T>(thunk: () => Promise<T>, hash: string) {
  // Schedule immediately; RequestQueue will handle rate-limits & retries
  return requestQueue.enqueue(thunk, Date.now(), hash);
}

// 3️⃣ Build a BatchQueue that uses the enqueue helper
const batchQueue = new BatchQueue<string, string>({
  maxCharactersPerBatch: 4000,
  maxItemsPerBatch: 8,
  batchDelay: 250,
  getBatchKey: () => "translate-default",
  getCharacters: (txt) => txt.length,
  // Flush a batch by sending a single request through RequestQueue
  executeBatch: async (texts) => {
    const hash = "batch-" + texts.join("|");
    return enqueueRequest(() => translateApi.bulk(texts), hash);
  },
  // Optional per-item fallback
  executeIndividual: async (txt) => {
    const hash = "single-" + txt;
    return enqueueRequest(() => translateApi.single(txt), hash);
  },
});

```

**Key integration points**:

- **Rate limiting** is centralized in `RequestQueue`; all batch and individual calls share the same token bucket.
- **Retries** operate at both layers: `RequestQueue` retries any failing HTTP call, while `BatchQueue` retries specifically when the LLM returns incorrect result counts.

## Summary

- **Size batches appropriately**: Set `maxCharactersPerBatch` just below provider limits (e.g., 4,800 for a 5,000-character cap) to maximize payload efficiency.
- **Tune flush timing**: Configure `batchDelay` to approximately half your average request inter-arrival time to balance latency and batch utilization.
- **Respect rate limits**: Use the token-bucket parameters (`rate`, `capacity`) in `RequestQueue` to smooth bursts without exceeding API quotas.
- **Handle failures gracefully**: Implement exponential back-off with jitter via `baseRetryDelayMs` and enable `enableFallbackToIndividual` for stubborn batch failures.
- **Eliminate redundant work**: Leverage built-in request deduplication via payload hashing in `RequestQueue.enqueue`.
- **Adjust without restarts**: Use `setBatchConfig` and `setQueueOptions` for dynamic pipeline tuning in production environments.

## Frequently Asked Questions

### What triggers a BatchCountMismatchError and how does the pipeline handle it?

A `BatchCountMismatchError` occurs when the translation API returns a different number of results than the number of tasks sent in the batch (defined at [`src/utils/request/batch-queue.ts#L1-L7`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/batch-queue.ts#L1-L7)). When this error is detected, `BatchQueue` automatically retries the batch using exponential back-off with delays capped between 1,000ms and 8,000ms. If retries exhaust and `enableFallbackToIndividual` is true, the system falls back to processing each text segment individually through `executeIndividual`.

### How does the token-bucket algorithm prevent API quota violations?

The `RequestQueue` implements a token-bucket rate limiter using `bucketTokens` and `lastRefill` fields ([`src/utils/request/request-queue.ts#L30-L33`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/request/request-queue.ts#L30-L33)). Tokens are consumed for each request and refilled at the configured `rate` (tokens per second) up to `capacity`. This smooths traffic bursts by queueing requests when tokens are depleted, ensuring the average request rate never exceeds the provider's allowed quota while still permitting short bursts up to the capacity limit.

### When should I enable fallback to individual requests?

Enable `enableFallbackToIndividual` when your translation provider's batch endpoint exhibits instability or returns inconsistent result counts. This is particularly valuable for large language models that occasionally return fewer translations than requested due to content filtering or context window limitations. While this increases API call volume, it guarantees request completion rather than failing the entire batch permanently.

### How do I calculate the optimal batchDelay for my workload?

Calculate `batchDelay` as approximately 50% of your average request inter-arrival time. For example, if your application receives translation requests every 200ms on average, set `batchDelay` to 100ms. This heuristic minimizes latency for individual requests while maximizing the probability of accumulating multiple requests into efficient batches. Monitor your batch fill rates in production and increase the delay if batches are consistently under-filled, or decrease it if latency becomes unacceptable.