# How the extractKeyword Utility Enables NLP-Based Search in DeepWiki-MCP

> Discover how the extractKeyword utility uses NLP to parse text and extract tech keywords for DeepWiki-MCP, resolving user queries into GitHub repository references.

- Repository: [Kevin Kern/deepwiki-mcp](https://github.com/regenrek/deepwiki-mcp)
- Tags: deep-dive
- Published: 2026-02-16

---

**The extractKeyword utility uses lightweight NLP processing to parse free-form text and extract technology keywords, enabling the DeepWiki-MCP tool to resolve ambiguous user queries into concrete GitHub repository references.**

The `regenrek/deepwiki-mcp` repository implements a Model Context Protocol (MCP) server that fetches and converts DeepWiki documentation into markdown. At the core of its search functionality lies the **extractKeyword utility for NLP-based search**, which bridges the gap between natural language input and structured repository resolution.

## What Is the extractKeyword Utility?

The `extractKeyword` function is a specialized NLP helper located in [[`src/utils/extractKeyword.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/extractKeyword.ts)](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/extractKeyword.ts). It processes free-form text—such as user queries like "how to use react for UI" or URL fragments—and returns the most likely technology or library name.

The implementation leverages the lightweight **wink-NLP** library with the English lite model. This choice keeps the bundle size minimal while providing sufficient linguistic analysis for keyword extraction tasks.

## How extractKeyword Implements NLP-Based Search

The utility follows a deterministic pipeline to isolate meaningful technology terms from noisy input:

1. **Tokenization** – The wink-NLP engine segments input text into individual tokens with part-of-speech tagging.
2. **Part-of-speech filtering** – The function specifically targets **nouns** and **proper nouns**, discarding verbs, adjectives, and other non-essential word types.
3. **Stop-word removal** – A curated stop-list filters out generic terms like `how`, `what`, `new`, `use`, and `for` that do not contribute to technology identification.
4. **First-match return** – The function returns the first remaining token that survives filtering. If no candidates remain, it yields `undefined`.

This approach ensures that inputs like "awesome chart library" resolve to `chart` while "react component library" resolves to `react`.

## Using extractKeyword in the DeepWiki Tool

The primary consumer of this utility is the DeepWiki fetch tool implemented in [[`src/tools/deepwiki.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/tools/deepwiki.ts)](https://github.com/regenrek/deepwiki-mcp/blob/main/src/tools/deepwiki.ts). Here, the `extractKeyword` function serves as the critical bridge between natural language input and repository resolution.

When processing user input that lacks a URL protocol (indicating free-form text rather than a direct link), the tool:

1. Invokes `extractKeyword` to parse the input phrase
2. Passes the extracted keyword to `resolveRepo`, which attempts to map the term to a concrete `owner/repo` reference on GitHub
3. Falls back to a default placeholder (`defaultuser/${extracted}`) if resolution fails, maintaining backward compatibility while maximizing success rates through keyword extraction

This integration allows users to query the MCP server with conversational phrases rather than requiring exact repository names or URLs.

## Code Examples

### Direct Utility Usage

```typescript
import { extractKeyword } from '@/utils/extractKeyword';

// Extracts technology from natural language
console.log(extractKeyword('how to use react for UI'));
// → "react"

console.log(extractKeyword('awesome chart library'));
// → "chart"

```

### Integration in DeepWiki Tool

```typescript
// Inside src/tools/deepwiki.ts (simplified)
if (!/^https?:\/\//.test(url)) {
  // Input is free-form text like "vue state management"
  const extracted = extractKeyword(url);
  
  if (extracted) {
    try {
      // Map keyword to GitHub repository
      const repo = await resolveRepo(extracted); // e.g. "vuejs/vuex"
      url = repo;
    } catch {
      // Fallback when resolution fails
      url = `defaultuser/${extracted}`;
    }
  }
}

```

## Summary

- The **extractKeyword utility** in [`src/utils/extractKeyword.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/extractKeyword.ts) provides lightweight NLP-based keyword extraction using the wink-NLP library.
- It filters input text for nouns and proper nouns while removing common stop words to isolate technology terms.
- The function enables the DeepWiki tool to **normalize ambiguous user queries** into concrete repository identifiers through integration with the `resolveRepo` function.
- When extraction succeeds but repository resolution fails, the system maintains **backward compatibility** through fallback placeholder generation.

## Frequently Asked Questions

### How does extractKeyword handle multi-word technology names?

The `extractKeyword` function returns only the first valid noun or proper noun it encounters after filtering. For multi-word names like "next js" or "vue router," it extracts the first component (e.g., "next" or "vue"). The `resolveRepo` function in [`src/tools/deepwiki.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/tools/deepwiki.ts) then handles the mapping logic to match these partial terms to full repository names.

### What happens when extractKeyword cannot identify a keyword?

When the input contains no nouns or proper nouns that survive the stop-word filtering process, the function returns `undefined`. In [`src/tools/deepwiki.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/tools/deepwiki.ts), this triggers the fallback logic where the system attempts to use the raw input or a default placeholder, ensuring the tool degrades gracefully rather than failing completely.

### Why does the utility use wink-NLP instead of larger NLP libraries?

The implementation prioritizes **bundle size and performance** over comprehensive linguistic analysis. The wink-NLP library with its English lite model provides sufficient part-of-speech tagging for keyword extraction without the heavy dependencies of larger frameworks like spaCy or NLTK. This aligns with the MCP server's goal of remaining lightweight and fast to initialize.

### Can extractKeyword be used for languages other than English?

Currently, the utility is limited to English input due to its reliance on the wink-NLP English lite model and an English-specific stop-word list. Processing non-English queries would require extending the implementation to support additional language models and stop-word dictionaries, which is not currently implemented in [`src/utils/extractKeyword.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/utils/extractKeyword.ts).