Understanding the keepAliveTimeout Configuration in deepwiki-mcp's HTTP Agent
The keepAliveTimeout setting in deepwiki-mcp's Undici HTTP Agent controls how long idle TCP connections remain open for reuse, set to 5 seconds by default to balance connection pooling efficiency with resource cleanup.
The regenrek/deepwiki-mcp repository implements a micro-crawling pipeline that relies on efficient HTTP connection management to fetch documentation pages rapidly. At the heart of this system lies a carefully tuned keepAliveTimeout configuration that determines how long the crawler maintains idle sockets before releasing them.
What Is keepAliveTimeout in deepwiki-mcp?
In src/lib/httpCrawler.ts at line 35, deepwiki-mcp instantiates a shared Undici Agent with a specific timeout value:
new Agent({ keepAliveTimeout: 5_000 })
The keepAliveTimeout parameter defines the duration in milliseconds that the HTTP agent keeps an idle TCP connection open after completing a request. When a socket finishes transmitting data, the agent starts a timer. If no new request targets that same host before the timer expires, the agent forcibly closes the connection to free the underlying file descriptor.
Why deepwiki-mcp Uses a 5-Second keepAliveTimeout
The 5-second default represents a deliberate engineering trade-off designed specifically for web crawling workloads.
Connection Reuse and Performance
By maintaining idle sockets for brief periods, deepwiki-mcp enables HTTP persistent connections (HTTP/1.1 keep-alive). Subsequent fetch calls to the same host can reuse existing TCP connections, eliminating the latency overhead of:
- TCP three-way handshake (SYN, SYN-ACK, ACK)
- TLS/SSL negotiation for HTTPS endpoints
This optimization proves critical when crawling documentation sites that require fetching dozens of interlinked pages from the same domain.
Resource Management and Cleanup
Unlike browsers that might hold connections open for minutes, deepwiki-mcp aggressively closes idle sockets after 5 seconds. This prevents the Node.js process from accumulating unused file descriptors or consuming memory for connection state tracking. In high-concurrency crawling scenarios, this constraint protects against "too many open files" errors and ensures predictable memory usage.
Balancing Act for Web Crawling
The 5-second window accommodates the crawler's typical access pattern: rapid sequential requests to the same host as it follows internal links. If the crawler encounters a delay between requests (e.g., processing complex pages), the connection closes automatically, forcing a fresh handshake for the next request—a reasonable penalty for ensuring resource hygiene.
Configuring keepAliveTimeout in Your Implementation
When adapting deepwiki-mcp for different workloads, you may adjust this parameter based on your specific latency and resource requirements.
Default Implementation
import { Agent } from 'undici';
// Standard deepwiki-mcp configuration
const agent = new Agent({ keepAliveTimeout: 5_000 });
Extended Timeout for Slow Crawling
If your crawler processes large documents with significant delays between requests, increase the timeout to maintain persistent connections longer:
const slowAgent = new Agent({ keepAliveTimeout: 30_000 }); // 30 seconds
Aggressive Cleanup for High-Concurrency
For burst-crawling scenarios where you rapidly spawn many concurrent requests to diverse hosts, reduce the timeout to free resources faster:
const aggressiveAgent = new Agent({ keepAliveTimeout: 1_000 }); // 1 second
Summary
keepAliveTimeoutinsrc/lib/httpCrawler.tscontrols how long idle TCP sockets remain open after completing HTTP requests.- The 5-second default balances connection reuse efficiency against resource cleanup requirements for web crawling workloads.
- Connection reuse eliminates TCP handshake and TLS negotiation overhead when fetching multiple pages from the same host.
- Resource management prevents file descriptor exhaustion by aggressively closing unused connections.
- The parameter is configurable based on specific crawling patterns, with trade-offs between latency and memory usage.
Frequently Asked Questions
What happens if keepAliveTimeout is set to 0 in deepwiki-mcp?
Setting keepAliveTimeout: 0 forces the Undici agent to immediately close every TCP connection after completing the HTTP response. While this eliminates any risk of file descriptor leaks, it forces the crawler to perform a full TCP handshake and TLS negotiation for every single request, significantly degrading performance when crawling multiple pages from the same domain.
How does keepAliveTimeout differ from headersTimeout in Undici?
keepAliveTimeout governs the lifecycle of the underlying TCP socket between requests, determining how long an idle connection stays open. In contrast, headersTimeout (also configurable in the Agent options) controls the maximum duration the client waits to receive HTTP response headers after sending a request. The former manages connection state, while the latter enforces response latency limits.
Can increasing keepAliveTimeout improve crawling performance?
Yes, increasing keepAliveTimeout can improve performance when crawling large sites with rapid sequential access patterns. By maintaining connections open for 30 seconds or longer, the crawler avoids repeated TCP handshakes and TLS negotiations when fetching dozens of pages from the same host. However, longer timeouts consume more file descriptors and memory, potentially causing resource exhaustion under high concurrency or when crawling thousands of different domains.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →