# How URL Normalization in Deepwiki-MCP Handles Short-Form Repository Inputs

> Learn how Deepwiki-MCP normalizes short-form owner/repo inputs to absolute URLs using regex and prepending https://deepwiki.com/ Discover the mechanics behind this efficient handling.

- Repository: [Kevin Kern/deepwiki-mcp](https://github.com/regenrek/deepwiki-mcp)
- Tags: internals
- Published: 2026-02-16

---

**Deepwiki-MCP normalizes short-form `owner/repo` inputs by detecting the pattern with the regex `^[^/]+/[^/]+$`, preserving the string unchanged, and prepending `https://deepwiki.com/` to construct the final absolute URL.**

The `deepwiki-mcp` repository provides a Model Context Protocol (MCP) server that fetches documentation from Deepwiki. When developers invoke the `deepwiki_fetch` tool, the implementation performs robust URL normalization in [`src/tools/deepwiki.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/tools/deepwiki.ts) to handle various input formats, including convenient short-form repository references like `vercel/ai`.

## The Normalization Pipeline in [`src/tools/deepwiki.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/tools/deepwiki.ts)

The normalization logic executes a six-step pipeline that processes the `url` parameter before any network request occurs.

### Step 1: Input Validation and HTTP URL Detection

First, the code trims whitespace and verifies the input is a string. If the value already begins with `http://` or `https://`, the pipeline exits early and leaves the URL untouched.

### Step 2: Detecting the Short-Form owner/repo Pattern

The critical step for short-form inputs uses the regex `^[^/]+/[^/]+$` to detect strings matching the `owner/repo` pattern. When a match occurs, the code considers this a valid short-form reference and preserves it exactly as provided. No additional keyword extraction or repository resolution occurs at this stage.

### Step 3: Handling Single-Word and Free-Form Inputs

For inputs that do not match the short-form pattern, the implementation employs fallback mechanisms. Single-word terms trigger the `extractKeyword` utility from [`src/utils/extractKeyword.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/extractKeyword.ts), followed by `resolveRepo` in [`src/utils/resolveRepoFetch.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/resolveRepoFetch.ts) to query the GitHub Search API. If resolution fails, the system falls back to `defaultuser/<term>`. Free-form phrases containing slashes that do not fit the strict `owner/repo` pattern undergo identical processing.

### Step 4: Constructing the Absolute Deepwiki URL

Finally, the pipeline prefixes `https://deepwiki.com/` to the processed string, transforming `owner/repo` into `https://deepwiki.com/owner/repo`.

## Code Examples: Normalizing Short-Form Inputs

The following TypeScript examples demonstrate how the `deepwiki_fetch` tool processes different input formats:

```typescript
// Example 1 – Short-form repository reference
await mcp.tools.deepwiki_fetch({
  url: 'vercel/ai',      // Matches ^[^/]+/[^/]+$
  maxDepth: 1,
  verbose: false,
  mode: 'markdown',
});
// Normalized request URL → https://deepwiki.com/vercel/ai

```

```typescript
// Example 2 – Full HTTP URL (bypasses normalization)
await mcp.tools.deepwiki_fetch({
  url: 'https://deepwiki.com/vercel/ai',
  maxDepth: 1,
  verbose: false,
  mode: 'markdown',
});
// Already absolute; returned unchanged

```

## Supporting Utilities and Fallback Mechanisms

The normalization pipeline relies on several specialized utilities:

- **[`src/utils/resolveRepoFetch.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/resolveRepoFetch.ts)** – Queries the GitHub Search API to translate plain keywords into `owner/repo` strings when automatic resolution is required.

- **[`src/utils/extractKeyword.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/extractKeyword.ts)** – Implements NLP logic to extract likely library names from free-form text inputs (e.g., parsing "React router" to extract the relevant identifier).

- **[`src/schemas/deepwiki.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/schemas/deepwiki.ts)** – Defines the Zod `FetchRequest` schema that validates the `url` field structure before normalization begins.

- **[`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts)** – Consumes the fully normalized Deepwiki URL to perform the actual content crawling, relying on the guaranteed format produced by the pipeline.

## Summary

- URL normalization in `deepwiki-mcp` occurs in [`src/tools/deepwiki.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/tools/deepwiki.ts) before any network requests execute.

- Short-form `owner/repo` inputs match the regex `^[^/]+/[^/]+$` and pass through unchanged except for the final domain prefix.

- HTTP(S) URLs bypass normalization entirely to preserve existing absolute references.

- Non-conforming inputs trigger fallback mechanisms using `extractKeyword` and `resolveRepo` with GitHub Search API integration.

- All paths ultimately resolve to `https://deepwiki.com/{normalized_path}`.

## Frequently Asked Questions

### What regex pattern does deepwiki-mcp use to detect short-form repository URLs?

The implementation uses the pattern `^[^/]+/[^/]+$` to identify valid `owner/repo` strings. This regex ensures exactly one slash separates two non-empty segments, distinguishing short-form references from absolute URLs or multi-segment paths.

### Does deepwiki-mcp modify short-form inputs like 'owner/repo' during normalization?

No, the code preserves short-form inputs exactly as provided. When the regex matches, the string remains unchanged through steps 1-5 of the pipeline. Only in the final construction step does the system prepend `https://deepwiki.com/` to create the absolute URL.

### What happens if I provide a single-word repository name without a slash?

Single-word inputs trigger the fallback resolution pipeline. The system calls `extractKeyword` to parse the term, then `resolveRepo` to query the GitHub Search API. If no matching repository is found, the code defaults to `defaultuser/<term>` before final URL construction.

### Where is the URL normalization logic located in the repository?

The core normalization logic resides in [`src/tools/deepwiki.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/tools/deepwiki.ts) within the `deepwiki_fetch` tool implementation. Supporting utilities for keyword extraction and repository resolution are located in [`src/utils/extractKeyword.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/extractKeyword.ts) and [`src/utils/resolveRepoFetch.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/resolveRepoFetch.ts) respectively.