How to Handle Memory Management for Processing Large PDF Files (100GB+) in Stirling-PDF

Stirling-PDF processes 100GB+ PDFs without exhausting RAM by isolating PDF.js workers, enforcing LRU-capped thumbnail caches, and automatically revoking blob URLs through centralized lifecycle management.

Stirling-Tools/Stirling-PDF is engineered to handle PDFs scaling to hundreds of gigabytes without browser crashes or server memory exhaustion. The application implements a coordinated memory management strategy that combines worker isolation, incremental parsing, and aggressive resource cleanup. This architecture ensures that memory management for processing large PDF files remains bounded even when underlying document streams exceed available system memory.

1. Worker-Based PDF.js Isolation

The PdfWorkerManager service in frontend/src/core/services/pdfWorkerManager.ts creates every PDF.js PDFDocumentProxy inside a dedicated WebWorker. This isolates the PDF parsing engine from the main thread and allows explicit termination when processing completes.

Key methods include:

  • createDocument(file: File) – Instantiates a new worker and loads the PDF document proxy.
  • destroyDocument(pdf: PDFDocumentProxy) – Sends a termination signal to the worker and removes the entry from the internal active documents map.
  • destroyAll() – Emergency cleanup that terminates all active workers (e.g., on page unload).

Typical usage follows a strict create-process-destroy pattern:

const pdf = await pdfWorkerManager.createDocument(file);
// ... read pages or extract metadata ...
pdfWorkerManager.destroyDocument(pdf);

This pattern appears in thumbnail generation and enhanced PDF processing, ensuring workers do not linger after their task completes.

2. Incremental Processing for Massive Files

When a file exceeds the configurable threshold of approximately 200MB, EnhancedPDFProcessingService (located in frontend/src/core/services/enhancedPDFProcessingService.ts) switches to metadata-only mode. The processFile method checks a largeFile flag and loads only the document catalog, outline, and page count while skipping full-page rasterization.

Benefits of this approach include:

  • Minimal heap allocation – No pixel buffers are created for unseen pages.
  • Structural operation support – Split, merge, and OCR front-end operations can proceed using only metadata.
  • Lazy loading – If the user later requests page-level operations, the service loads specific pages on demand rather than the entire document.

3. LRU-Capped Thumbnail Cache

The thumbnailGenerationService.ts implements a bounded cache with a default limit of 500MB. When generating previews, each entry stores its PDF document reference. Upon reaching the cache limit, the service evicts the oldest entry and explicitly destroys its associated worker.

The eviction logic ensures bounded memory usage:

if (this.cacheSize > this.maxCacheSize) {
  const [oldestKey, oldestEntry] = this.cache.entries().next().value;
  pdfWorkerManager.destroyDocument(oldestEntry.pdf);
  this.cache.delete(oldestKey);
}

This LRU (Least Recently Used) strategy guarantees that only a fixed number of PDFs remain resident regardless of how many large files the user opens.

4. Centralized Lifecycle Management

All UI-visible files live inside FileContext, defined in frontend/src/core/contexts/file/lifecycle.ts. Every file entry registers a cleanupFile callback that executes three critical actions:

  1. Revokes blob URLs via URL.revokeObjectURL.
  2. Destroys PDF workers by calling pdfWorkerManager.destroyDocument for any lingering document proxy.
  3. Removes entries from the LRU thumbnail cache.

When a user removes a file, switches tools, or closes the application, cleanupFile triggers automatically. This centralized registry prevents stray memory references from surviving navigation events.

5. Blob URL Revocation and Resource Tracking

The resourceManager.ts utility centralizes the creation of temporary object URLs. The createObjectUrl function returns both the URL and a bound cleanup function:

export function createObjectUrl(blob: Blob): { url: string; cleanup: () => void } {
  const url = URL.createObjectURL(blob);
  const cleanup = () => URL.revokeObjectURL(url);
  return { url, cleanup };
}

Hooks like useToolResources (in frontend/src/core/hooks/tools/shared/useToolResources.ts) aggregate these cleanup functions across a tool operation. When the operation finishes or the component unmounts, all pending blob URLs are revoked, releasing the underlying memory.

6. End-to-End Example: Streaming a 50GB Split Operation

The following flow demonstrates how memory-management layers cooperate when splitting an extremely large PDF into 10MB chunks:

// 1️⃣ Initialize the split tool with resource tracking
const { runOperation, cleanupBlobUrls } = useSplitOperation();

// 2️⃣ Execute the split workflow
const handleSplit = async (file: File) => {
  // Register file with FileContext for automatic cleanup
  addFiles([file]);

  // Run operation with streaming processing:
  // • pdfWorkerManager creates a temporary worker
  // • EnhancedPDFProcessingService detects large size → metadata-only mode
  // • Pages stream one at a time; each chunk written to a new Blob
  // • Worker destroyed immediately after each chunk to keep RAM low
  await runOperation({ file, chunkSizeMB: 10 });

  // 3️⃣ Revoke all temporary blob URLs created during processing
  cleanupBlobUrls();
};

// 4️⃣ Final cleanup when user removes the file
removeFile(file.id); // Triggers cleanupFile → revokes URLs + destroys workers

This streaming approach never holds the full 50GB document in memory. It processes one page at a time, releases the worker, and moves to the next segment.

Summary

  • PdfWorkerManager isolates PDF.js in dedicated WebWorkers that are explicitly destroyed after use.
  • EnhancedPDFProcessingService switches to metadata-only mode for files exceeding ~200MB.
  • ThumbnailGenerationService enforces a 500MB LRU cache that evicts old documents and terminates their workers.
  • FileContext registers per-file cleanup callbacks that revoke blob URLs and destroy document proxies on navigation or removal.
  • ResourceManager provides centralized blob URL creation with mandatory cleanup functions consumed by React hooks.

Frequently Asked Questions

What happens when a PDF exceeds the 200MB threshold in Stirling-PDF?

When EnhancedPDFProcessingService detects a file larger than the threshold (default ~200MB), it activates metadata-only mode. The service loads only the document outline, catalog, and page count while skipping rasterization. This allows structural operations like split or merge to proceed with minimal memory footprint, and pages are loaded lazily only when explicitly requested.

How does Stirling-PDF prevent memory leaks when switching between tools?

The FileContext lifecycle system registers a cleanupFile callback for every file added to the UI. When a user switches tools, removes a file, or closes the tab, these callbacks automatically revoke blob URLs, call pdfWorkerManager.destroyDocument on any active PDF proxies, and purge entries from the thumbnail cache. This guarantees no orphaned references survive navigation events.

What is the default thumbnail cache size and how does eviction work?

The default thumbnail cache limit is 500MB, defined in thumbnailGenerationService.ts. When the cache exceeds this limit, the service evicts the oldest entry using an LRU (Least Recently Used) algorithm. During eviction, the service calls pdfWorkerManager.destroyDocument on the associated PDF document, immediately releasing the underlying worker and buffers.

How are PDF.js workers terminated after processing completes?

Workers are terminated through PdfWorkerManager.destroyDocument(pdf), which sends a termination signal to the specific WebWorker handling that document and removes the reference from the internal active documents map. For emergency scenarios like page unload, destroyAll() terminates every active worker simultaneously.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →