How the DeepWiki-MCP Crawler Uses Content-Type Headers to Process Only HTML
The DeepWiki-MCP crawler inspects the Content-Type response header immediately after fetching a URL and skips any resource that does not include the substring text/html, ensuring only HTML documents are parsed and stored.
The deepwiki-mcp crawler is designed to extract and convert web documentation into structured Markdown, but it must avoid wasting resources on images, stylesheets, and binary files. By implementing a strict Content-Type header validation mechanism in src/lib/httpCrawler.ts, the crawler guarantees that only HTML responses enter the parsing pipeline.
Two-Layer Filtering Strategy in deepwiki-mcp
The crawler employs a defense-in-depth approach to ensure only HTML content is processed. This strategy combines URL-level pre-filtering with response-level header inspection.
Extension-Based Pre-Filtering
Before any HTTP request is dispatched, the crawler examines the URL's file extension. A whitelist of known non-HTML extensions—such as .css, .js, .png, .pdf, and .jpg—is maintained in the enqueue logic. URLs matching these patterns are rejected immediately, preventing unnecessary network overhead.
Content-Type Header Verification
Even if a URL passes the extension check, the crawler validates the actual response content. Immediately after receiving the HTTP response, the code inspects the Content-Type header to confirm the resource is HTML before proceeding with body parsing or link extraction.
Implementation Details in src/lib/httpCrawler.ts
The core validation logic resides in src/lib/httpCrawler.ts, specifically within the fetch handling routine (around lines 40-44). The implementation uses Undici, a Node.js HTTP client, to perform requests and access response headers.
const res = await fetch(url, { dispatcher: agent })
const contentType = res.headers.get('content-type') || ''
if (!contentType.includes('text/html')) {
return // non-HTML resources are ignored
}
This code performs three critical operations:
- Fetches the resource using Undici's
fetchwith a customdispatcher(agent) for connection pooling. - Extracts the header using
res.headers.get('content-type'), defaulting to an empty string if missing. - Validates the content by checking if the header string includes
'text/html'. If not, the function returns early, skipping all subsequent parsing, storage, and link extraction.
Practical Code Examples
Running the Crawler with Automatic Filtering
When using the deepwiki-mcp crawler, the Content-Type validation happens automatically. You do not need to manually filter URLs—the crawl function handles both extension checks and header validation internally.
import { crawl } from '@/lib/httpCrawler'
import { URL } from 'node:url'
async function runCrawler() {
const root = new URL('https://example.com/')
const result = await crawl({
root,
maxDepth: 2,
emit: e => console.log('progress:', e),
verbose: true,
})
console.log('Fetched HTML pages:', Object.keys(result.html).length)
// Non-HTML resources such as images, PDFs, etc. are automatically omitted.
}
runCrawler()
Debugging Content-Type Detection
To understand why certain URLs are being skipped, you can replicate the header inspection logic manually using Undici:
import { fetch } from 'undici'
async function checkContentType(url: string) {
const res = await fetch(url)
const ct = res.headers.get('content-type') ?? ''
console.log(`${url} → Content-Type: ${ct}`)
console.log('Will be processed?', ct.includes('text/html'))
}
// Example checks
checkContentType('https://example.com/style.css') // Likely false
checkContentType('https://example.com/index.html') // Likely true
Key Files and Architecture
The Content-Type filtering mechanism spans several files in the deepwiki-mcp repository:
| File | Role | Direct Link |
|---|---|---|
src/lib/httpCrawler.ts |
Implements the breadth-first crawler, including the Content-Type check that restricts processing to HTML. |
View on GitHub |
src/server.ts |
Exposes the crawler via the HTTP API used by the CLI and external clients. | View on GitHub |
src/types.ts |
Defines ProgressEvent and other types used by the crawler's progress callbacks. |
View on GitHub |
These files together show how DeepWiki-MCP ensures that only HTML content is fetched, stored, and later transformed into Markdown.
Summary
- The deepwiki-mcp crawler uses a two-layer filtering system to ensure only HTML is processed: extension-based pre-filtering and Content-Type header validation.
- The critical check occurs in
src/lib/httpCrawler.ts(lines 40-44), where the code inspects theContent-Typeheader for the substringtext/html. - If the header check fails, the crawler returns early, skipping body parsing, link extraction, and storage, which conserves bandwidth and processing resources.
- The implementation uses Undici for HTTP requests and operates automatically without requiring manual URL filtering from the user.
Frequently Asked Questions
What happens if a server returns the wrong Content-Type?
If a server misconfigures its headers—for example, serving HTML with Content-Type: text/plain—the deepwiki-mcp crawler will skip the page because the header string does not include text/html. The crawler strictly adheres to the header value rather than attempting to sniff content, ensuring predictable behavior and avoiding the processing of binary data disguised as text.
Does deepwiki-mcp follow redirects before checking Content-Type?
Yes, the crawler uses Undici's fetch implementation, which automatically follows HTTP redirects (3xx responses) by default. The Content-Type check is performed on the final response after all redirects resolve. This ensures that if a URL redirects to a non-HTML resource (such as a PDF), the crawler correctly identifies and skips the final destination based on its Content-Type header.
Can I configure the crawler to accept non-HTML content?
No, the Content-Type filter is hardcoded in src/lib/httpCrawler.ts to specifically check for text/html. There is no configuration option to modify the accepted MIME types. If you need to process other content types (such as XML or JSON), you would need to fork the repository and modify the conditional check in the crawler source code to include additional MIME type strings.
Which HTTP library does deepwiki-mcp use for fetching?
The deepwiki-mcp crawler uses Undici, a modern HTTP/1.1 client for Node.js. This is evident in the import statements and the usage of fetch with a custom dispatcher (agent) for connection pooling. Undici provides the res.headers.get() API used to retrieve the Content-Type header for validation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →