# How the Deepwiki-MCP Crawler Handles Non-HTML Files and Filters Binary Assets

> Discover how the deepwiki-mcp crawler efficiently filters non-HTML files and binary assets using a two-stage system, ensuring only HTML documents are processed.

- Repository: [Kevin Kern/deepwiki-mcp](https://github.com/regenrek/deepwiki-mcp)
- Tags: internals
- Published: 2026-02-16

---

**The deepwiki-mcp crawler uses a two-stage filtering system—pre-enqueue extension blocking and runtime Content-Type verification—to ensure only HTML documents are processed while binary assets like images, PDFs, and executables are discarded.**

The deepwiki-mcp crawler is engineered to traverse websites and extract only meaningful HTML content, deliberately excluding non-textual resources that would waste bandwidth and processing power. By implementing strict filtering mechanisms directly in the crawl pipeline, the system guarantees that its output contains clean, parseable HTML without binary noise. This article examines the specific implementation details found in the source code, including the exact file paths and logic used to distinguish HTML from non-HTML resources.

## Two-Stage Filtering Architecture in Deepwiki-MCP

The crawler employs a defense-in-depth strategy with two independent validation layers. The first layer intercepts URLs before any network request occurs, while the second layer verifies the actual HTTP response headers. Together, these stages ensure that only documents with both acceptable extensions and correct MIME types enter the processing queue.

### Stage 1: Pre-Enqueue Extension Filtering

Before a URL is added to the crawl queue, the `enqueue` helper function in [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts) performs an aggressive pathname analysis. The code maintains a comprehensive hard-coded array of **non-HTML extensions** spanning images, stylesheets, scripts, archives, media files, documents, and binaries.

```typescript
// src/lib/httpCrawler.ts – lines 56-102
const nonHtmlExt = [
  '.css', '.js', '.mjs', '.json', '.png', '.jpg', '.jpeg', '.gif', 
  '.svg', '.ico', '.webp', '.mp4', '.webm', '.mp3', '.wav', '.ogg',
  '.pdf', '.zip', '.tar', '.gz', '.rar', '.7z', '.exe', '.dll', 
  '.bin', '.dat', '.db', '.sqlite', '.sql', '.bak', '.xml', '.txt', 
  '.csv', '.md', '.yml', '.yaml', '.log', '.docx', '.pptx', '.xlsx'
]

const lowerPath = url.pathname.toLowerCase()
if (nonHtmlExt.some(ext => lowerPath.endsWith(ext))) {
  return  // ← URL discarded before fetch
}

```

This whitelist approach ensures that resources like [`/assets/style.css`](https://github.com/regenrek/deepwiki-mcp/blob/main//assets/style.css) or `/images/logo.png` are rejected immediately, preventing unnecessary HTTP requests and conserving crawler resources.

### Stage 2: Runtime Content-Type Verification

Even URLs passing the extension filter undergo a secondary validation after the HTTP response arrives. The crawler inspects the `Content-Type` header to confirm the response actually contains HTML, catching edge cases where extension-less URLs serve binary data (such as dynamically generated images or API endpoints returning PDFs).

```typescript
// src/lib/httpCrawler.ts – lines 39-44
const res = await fetch(url, { dispatcher: agent })
const contentType = res.headers.get('content-type') || ''

if (!contentType.includes('text/html')) {
  return  // ← Binary or non-HTML response discarded
}

```

This runtime check acts as a safety net, ensuring that only responses with `text/html` in their MIME type proceed to HTML parsing and storage in the `CrawlResult` object.

## Implementation Details in httpCrawler.ts

The filtering logic resides primarily in **[`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts)**, which implements the breadth-first crawl engine. The `nonHtmlExt` array (lines 56-102) represents a deliberately exhaustive blacklist covering common web asset categories:

- **Static assets**: CSS, JavaScript, images (PNG, JPG, SVG, WebP)
- **Media**: Video and audio files (MP4, WebM, MP3)
- **Documents**: PDFs, Office formats (DOCX, PPTX, XLSX), OpenDocument
- **Data files**: JSON, XML, CSV, YAML, SQLite databases
- **Archives and executables**: ZIP, TAR, EXE, DLL, BIN

The crawler processes these checks within the `enqueue` function before URLs enter the `p-queue` instance, ensuring the crawl frontier remains strictly limited to HTML documents.

## Practical Usage Example

When invoking the crawler, the filtering happens transparently. The following example demonstrates how binary assets are automatically excluded while HTML pages are retrieved:

```typescript
import { crawl } from './src/lib/httpCrawler.js'

const root = new URL('https://en.wikipedia.org/')
await crawl({
  root,
  maxDepth: 2,
  emit: e => console.log(`[${e.type}] ${e.url} – ${e.bytes} B`),
  verbose: true,
})

```

**Observed behavior:**
- Requests to [`/wiki/Main_Page.css`](https://github.com/regenrek/deepwiki-mcp/blob/main//wiki/Main_Page.css) or `/wiki/image.jpg` are **skipped** during the enqueue phase—no network traffic occurs.
- If a URL like `/api/download` returns `Content-Type: application/pdf`, the response is **discarded** after the fetch completes, despite having no `.pdf` extension.

## Summary

- The deepwiki-mcp crawler implements a **two-stage filtering system** in [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts) to eliminate non-HTML resources.
- **Extension filtering** occurs pre-fetch using a comprehensive blacklist (`nonHtmlExt` array, lines 56-102) that blocks URLs ending with known binary or non-HTML extensions.
- **Content-Type verification** runs post-fetch to catch extension-less URLs serving binary data, ensuring only `text/html` responses enter the processing pipeline.
- This architecture minimizes bandwidth consumption, reduces parsing overhead, and guarantees that [`CrawlResult.html`](https://github.com/regenrek/deepwiki-mcp/blob/main/CrawlResult.html) contains exclusively HTML document content.

## Frequently Asked Questions

### What file extensions does the deepwiki-mcp crawler block?

The crawler maintains an extensive whitelist of non-HTML extensions in [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts) covering static assets (`.css`, `.js`, `.png`, `.jpg`, `.svg`), media files (`.mp4`, `.mp3`, `.pdf`), documents (`.docx`, `.pptx`, `.xlsx`), data formats (`.json`, `.xml`, `.csv`, `.yaml`), archives (`.zip`, `.tar`, `.gz`), and executables (`.exe`, `.dll`, `.bin`). Any URL pathname ending with these extensions is rejected before fetching.

### How does the crawler handle URLs without file extensions?

URLs lacking extensions undergo **runtime Content-Type verification** after the HTTP response arrives. The crawler inspects the `Content-Type` header and discards any response that does not include `text/html`. This catches dynamically generated binary content served by extension-less endpoints, such as API routes returning PDFs or images.

### Can I customize the list of blocked extensions?

Currently, the `nonHtmlExt` array in [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts) (lines 56-102) is hard-coded. The crawler does not expose a configuration option to modify this list at runtime. To change filtering behavior, you would need to fork the repository and edit the source code directly, then rebuild the project.

### Where is the filtering logic located in the codebase?

The primary filtering logic resides in **[`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts)**. The extension-based filtering occurs within the `enqueue` helper function (around lines 56-102), while the Content-Type verification happens immediately after the `fetch` call (around lines 39-44). These two stages work sequentially to ensure only HTML documents proceed through the crawl pipeline.