How the DeepWiki-MCP Crawler Handles Partial Failures and Aggregates Error Responses
The DeepWiki-MCP crawler processes each URL in isolation using a promise queue, implements exponential backoff retries up to three times, and aggregates all failures into a structured errors array returned alongside successful results.
The regenrek/deepwiki-mcp repository implements a fault-tolerant web crawler designed to continue operation even when individual pages fail to load. Understanding how this system handles partial failures is essential for building resilient documentation pipelines that process large URL sets without terminating on isolated network errors.
Isolated Per-URL Execution with PQueue
The crawler leverages PQueue (promise queue) to ensure that failures remain isolated to individual URLs. Each fetch operation executes as an independent task within the queue, meaning a network timeout or HTTP error in one task does not propagate to others or halt the overall crawl process.
This architecture is implemented in src/lib/httpCrawler.ts within the main crawl function, where the queue processes URLs concurrently while maintaining strict error boundaries between tasks.
Retry Logic with Exponential Backoff
When a fetch fails, the crawler does not immediately record it as a permanent failure. Instead, it enters a retry loop with exponential backoff to handle transient network issues.
Configurable Retry Limits
The system uses two constants defined in the crawler:
RETRY_LIMIT: Defaults to 3 attempts per URLBACKOFF_BASE_MS: Base milliseconds for calculating delay intervals
Backoff Strategy Implementation
The retry logic resides inside the while (true) loop within each queue task. After a failure, the crawler calculates the delay using exponential backoff:
// src/lib/httpCrawler.ts#L78-L84
if (retries < RETRY_LIMIT) {
retries++;
await setTimeout(BACKOFF_BASE_MS * 2 ** (retries - 1));
continue;
}
This approach waits progressively longer between attempts (e.g., 1s, 2s, 4s) to avoid overwhelming struggling servers while maximizing the chance of recovery from temporary outages.
Error Capture and Aggregation
Once the retry limit is exhausted, the crawler transitions from recovery to recording mode, ensuring no error data is lost while maintaining workflow continuity.
Local Error Collection
Rather than throwing exceptions that would disrupt the queue, the crawler pushes failure details into a local errors array:
// src/lib/httpCrawler.ts#L84-L86
errors.push({ path: key, reason: String(err) })
Each entry contains the path (URL identifier) and the stringified error reason, creating a structured log of all failures that occurred during the crawl session.
CrawlResult Structure
After the queue drains completely, the crawl function returns a comprehensive CrawlResult object that unifies successes and failures:
// src/lib/httpCrawler.ts#L94-L100
return { html, errors, bytes: totalBytes, elapsedMs }
This structure includes:
html: Object mapping URLs to fetched contenterrors: Array of all failed URL recordsbytes: Total data transferredelapsedMs: Total execution time
Real-Time Progress Monitoring
The crawler exposes failure states through ProgressEvent emissions, allowing calling code to monitor partial failures in real time without interrupting the crawl.
Each request emits events containing the current retries count, enabling live dashboards or logs to display which URLs are struggling and how many retry attempts they have consumed. This mechanism is defined in src/types.ts and utilized throughout the crawler implementation.
Practical Implementation Example
The following example demonstrates initiating a crawl, monitoring progress events, and handling the aggregated error list upon completion:
import { crawl } from '@/lib/httpCrawler';
import { URL } from 'node:url';
// Start a crawl with a depth limit of 2
const result = await crawl({
root: new URL('https://example.com/'),
maxDepth: 2,
// Simple logger that prints progress
emit: (e) => console.log(`[${e.type}] ${e.url} – retries:${e.retries}`),
verbose: true,
});
console.log('Fetched pages:', Object.keys(result.html).length);
console.log('Failed pages:', result.errors.length);
result.errors.forEach(err => {
console.error(`❌ ${err.path} → ${err.reason}`);
});
This implementation showcases the fault-tolerant design: even if specific pages fail after three retry attempts, the crawl continues, and all failures are accessible via result.errors for post-processing or reporting.
Summary
The DeepWiki-MCP crawler handles partial failures through a multi-layered resilience strategy:
- Isolated execution via PQueue prevents single failures from aborting the entire crawl
- Exponential backoff retries (up to 3 attempts) recover transient network issues
- Structured error aggregation collects all failures into a returnable
errorsarray - Real-time progress events expose retry states without interrupting execution
- Comprehensive result objects unify successful fetches and error logs in
CrawlResult
Frequently Asked Questions
What happens when a URL fails permanently after all retries?
When a URL exhausts the RETRY_LIMIT (default 3), the crawler captures the error in the local errors array with the path and error message, then continues processing remaining URLs. The failure is included in the final CrawlResult returned by the crawl function in src/lib/httpCrawler.ts.
How does the crawler prevent one failure from stopping the entire job?
The crawler uses PQueue to execute each URL fetch as an independent task. Errors within individual tasks are caught and handled locally without propagating to the queue controller, ensuring that network timeouts or HTTP errors remain isolated to specific URLs while the broader crawl continues uninterrupted.
Can I adjust the retry behavior?
Yes, the retry logic uses configurable constants RETRY_LIMIT and BACKOFF_BASE_MS defined in src/lib/httpCrawler.ts. While the current implementation uses default values (3 retries with exponential backoff), these constants can be modified in the source to adjust the maximum attempts or the base delay interval between retries.
How do I access the list of failed URLs after crawling?
After the crawl function completes, access the errors property on the returned CrawlResult object. This array contains objects with path (the failed URL) and reason (the error message) for every URL that failed after exhausting all retry attempts, allowing you to log, report, or reprocess failed pages as needed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →