# How the DeepWiki-MCP Crawler Uses Content-Type Headers to Process Only HTML

> Learn how the deepwiki-mcp crawler uses Content-Type headers to identify and process only HTML, skipping other resource types for efficient data collection.

- Repository: [Kevin Kern/deepwiki-mcp](https://github.com/regenrek/deepwiki-mcp)
- Tags: internals
- Published: 2026-02-16

---

**The DeepWiki-MCP crawler inspects the `Content-Type` response header immediately after fetching a URL and skips any resource that does not include the substring `text/html`, ensuring only HTML documents are parsed and stored.**

The deepwiki-mcp crawler is designed to extract and convert web documentation into structured Markdown, but it must avoid wasting resources on images, stylesheets, and binary files. By implementing a strict **Content-Type header** validation mechanism in [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts), the crawler guarantees that only HTML responses enter the parsing pipeline.

## Two-Layer Filtering Strategy in deepwiki-mcp

The crawler employs a defense-in-depth approach to ensure only HTML content is processed. This strategy combines URL-level pre-filtering with response-level header inspection.

### Extension-Based Pre-Filtering

Before any HTTP request is dispatched, the crawler examines the URL's file extension. A whitelist of known non-HTML extensions—such as `.css`, `.js`, `.png`, `.pdf`, and `.jpg`—is maintained in the enqueue logic. URLs matching these patterns are rejected immediately, preventing unnecessary network overhead.

### Content-Type Header Verification

Even if a URL passes the extension check, the crawler validates the actual response content. Immediately after receiving the HTTP response, the code inspects the `Content-Type` header to confirm the resource is HTML before proceeding with body parsing or link extraction.

## Implementation Details in [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts)

The core validation logic resides in **[`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts)**, specifically within the fetch handling routine (around lines 40-44). The implementation uses **Undici**, a Node.js HTTP client, to perform requests and access response headers.

```typescript
const res = await fetch(url, { dispatcher: agent })
const contentType = res.headers.get('content-type') || ''
if (!contentType.includes('text/html')) {
  return               // non-HTML resources are ignored
}

```

This code performs three critical operations:

1. **Fetches the resource** using Undici's `fetch` with a custom `dispatcher` (agent) for connection pooling.
2. **Extracts the header** using `res.headers.get('content-type')`, defaulting to an empty string if missing.
3. **Validates the content** by checking if the header string includes `'text/html'`. If not, the function returns early, skipping all subsequent parsing, storage, and link extraction.

## Practical Code Examples

### Running the Crawler with Automatic Filtering

When using the deepwiki-mcp crawler, the Content-Type validation happens automatically. You do not need to manually filter URLs—the `crawl` function handles both extension checks and header validation internally.

```typescript
import { crawl } from '@/lib/httpCrawler'
import { URL } from 'node:url'

async function runCrawler() {
  const root = new URL('https://example.com/')
  const result = await crawl({
    root,
    maxDepth: 2,
    emit: e => console.log('progress:', e),
    verbose: true,
  })

  console.log('Fetched HTML pages:', Object.keys(result.html).length)
  // Non-HTML resources such as images, PDFs, etc. are automatically omitted.
}
runCrawler()

```

### Debugging Content-Type Detection

To understand why certain URLs are being skipped, you can replicate the header inspection logic manually using Undici:

```typescript
import { fetch } from 'undici'

async function checkContentType(url: string) {
  const res = await fetch(url)
  const ct = res.headers.get('content-type') ?? ''
  console.log(`${url} → Content-Type: ${ct}`)
  console.log('Will be processed?', ct.includes('text/html'))
}

// Example checks
checkContentType('https://example.com/style.css')   // Likely false
checkContentType('https://example.com/index.html') // Likely true

```

## Key Files and Architecture

The Content-Type filtering mechanism spans several files in the deepwiki-mcp repository:

| File | Role | Direct Link |
|------|------|-------------|
| [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts) | Implements the breadth-first crawler, including the `Content-Type` check that restricts processing to HTML. | [View on GitHub](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts) |
| [`src/server.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/server.ts) | Exposes the crawler via the HTTP API used by the CLI and external clients. | [View on GitHub](https://github.com/regenrek/deepwiki-mcp/blob/main/src/server.ts) |
| [`src/types.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/types.ts) | Defines `ProgressEvent` and other types used by the crawler's progress callbacks. | [View on GitHub](https://github.com/regenrek/deepwiki-mcp/blob/main/src/types.ts) |

These files together show how DeepWiki-MCP ensures that only HTML content is fetched, stored, and later transformed into Markdown.

## Summary

- The deepwiki-mcp crawler uses a **two-layer filtering system** to ensure only HTML is processed: extension-based pre-filtering and Content-Type header validation.
- The critical check occurs in **[`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts)** (lines 40-44), where the code inspects the `Content-Type` header for the substring `text/html`.
- If the header check fails, the crawler returns early, skipping body parsing, link extraction, and storage, which conserves bandwidth and processing resources.
- The implementation uses **Undici** for HTTP requests and operates automatically without requiring manual URL filtering from the user.

## Frequently Asked Questions

### What happens if a server returns the wrong Content-Type?

If a server misconfigures its headers—for example, serving HTML with `Content-Type: text/plain`—the deepwiki-mcp crawler will skip the page because the header string does not include `text/html`. The crawler strictly adheres to the header value rather than attempting to sniff content, ensuring predictable behavior and avoiding the processing of binary data disguised as text.

### Does deepwiki-mcp follow redirects before checking Content-Type?

Yes, the crawler uses Undici's `fetch` implementation, which automatically follows HTTP redirects (3xx responses) by default. The Content-Type check is performed on the final response after all redirects resolve. This ensures that if a URL redirects to a non-HTML resource (such as a PDF), the crawler correctly identifies and skips the final destination based on its Content-Type header.

### Can I configure the crawler to accept non-HTML content?

No, the Content-Type filter is hardcoded in [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts) to specifically check for `text/html`. There is no configuration option to modify the accepted MIME types. If you need to process other content types (such as XML or JSON), you would need to fork the repository and modify the conditional check in the crawler source code to include additional MIME type strings.

### Which HTTP library does deepwiki-mcp use for fetching?

The deepwiki-mcp crawler uses **Undici**, a modern HTTP/1.1 client for Node.js. This is evident in the import statements and the usage of `fetch` with a custom `dispatcher` (agent) for connection pooling. Undici provides the `res.headers.get()` API used to retrieve the Content-Type header for validation.