How the MCP Server Implements Progress Event Emission During Crawling
The DeepWiki MCP server generates progress events through the httpCrawler library but does not forward them to clients because the MCP SDK lacks a sendEvent method, leaving the emitProgress handler as a no-op.
The regenrek/deepwiki-mcp repository demonstrates how to structure progress tracking during web crawling operations within an MCP (Model Context Protocol) server. While the underlying crawler implements robust event emission capabilities, the MCP server integration reveals a critical gap in the current SDK implementation regarding real-time progress reporting.
Event Generation in the HTTP Crawler
The crawling logic resides in src/lib/httpCrawler.ts, where the crawl function orchestrates page fetching and emits progress updates through a callback mechanism. For every successful page retrieval, the crawler constructs a ProgressEvent payload containing granular telemetry about the operation.
// src/lib/httpCrawler.ts
emit({
type: 'progress',
url: url.href,
bytes,
elapsedMs,
fetched: Object.keys(html).length,
queued: queue.size + queue.pending,
retries,
} as any)
This emission occurs synchronously within the fetch loop, providing real-time visibility into URL processing, byte counts, timing metrics, and queue depth.
Progress Event Schema Definition
The structure of progress events is formally defined in src/schemas/deepwiki.ts using Zod for runtime validation. The ProgressEvent schema enforces strict typing on all telemetry fields:
// src/schemas/deepwiki.ts
export const ProgressEvent = z.object({
type: z.literal('progress'),
url: z.string(),
bytes: z.number().int().nonnegative(),
elapsedMs: z.number().int().nonnegative(),
fetched: z.number().int().nonnegative(),
queued: z.number().int().nonnegative(),
retries: z.number().int().nonnegative(),
})
This schema ensures that every progress emission contains valid, non-negative integers for metrics and properly identifies the event type and target URL.
MCP Server Integration and the Emission Gap
Despite the crawler's capabilities, the DeepWiki tool in src/tools/deepwiki.ts implements a stubbed emitProgress function that prevents event propagation to MCP clients:
// src/tools/deepwiki.ts
function emitProgress(e: any) {
// Progress reporting is not supported in this context because McpServer does not have a sendEvent method.
}
...
await crawl({
root,
maxDepth: req.maxDepth,
emit: emitProgress,
verbose: req.verbose,
})
The comment explicitly states that McpServer does not expose a sendEvent method, which is required to forward progress events to connected clients. Consequently, while the crawler generates rich telemetry data, the MCP server acts as a dead end for these events in the current implementation.
Summary
- The
httpCrawlerlibrary insrc/lib/httpCrawler.tsgenerates detailed ProgressEvent objects for every fetched page, including URL, bytes, timing, and queue statistics. - The event schema in
src/schemas/deepwiki.tsdefines strict Zod validation for progress telemetry using non-negative integers and literal type constraints. - The DeepWiki MCP tool in
src/tools/deepwiki.tsreceives these events through theemitProgresscallback but cannot forward them to clients due to the lack of asendEventmethod in the MCP SDK. - Real-time progress reporting remains unimplemented in this repository, requiring SDK updates to enable event streaming to MCP clients.
Frequently Asked Questions
What fields are included in a ProgressEvent?
A ProgressEvent contains seven fields: type (always "progress"), url (the fetched URL), bytes (response size), elapsedMs (request duration), fetched (count of processed pages), queued (remaining queue size plus pending items), and retries (retry attempt count). All numeric fields are validated as non-negative integers.
Why doesn't the DeepWiki MCP server forward progress events to clients?
The MCP server cannot forward events because the McpServer class from the @modelcontextprotocol/sdk does not implement a sendEvent method required for server-to-client streaming. The emitProgress function in src/tools/deepwiki.ts explicitly documents this limitation as the reason for its no-op implementation.
How does the httpCrawler generate progress events during crawling?
The crawler invokes the emit callback function supplied in its configuration options after every successful page fetch. It constructs a progress payload containing the current URL, response byte count, elapsed milliseconds, total fetched pages, current queue depth, and retry statistics, then passes this object to the emitter.
What would be required to enable real-time progress reporting?
Enabling real-time progress would require the MCP SDK to expose a sendEvent or similar server-to-client streaming method. Once available, the emitProgress function in src/tools/deepwiki.ts would need to invoke this method with the ProgressEvent payload instead of returning silently.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →