# How the MCP Server Handles Stop Words for Recommendations

> Learn how the MCP server handles stop words. It filters common terms from queries to boost recommendation relevance and surface technical insights effectively.

- Repository: [Henry Ndubuaku/maths-cs-ai-compendium](https://github.com/HenryNdubuaku/maths-cs-ai-compendium)
- Tags: how-to-guide
- Published: 2026-07-16

---

**The MCP server filters out common English words and domain-specific verbs from user queries before matching them against section metadata, ensuring that only meaningful technical terms influence the recommendation scoring.**

The Model-Context-Protocol (MCP) server in the HenryNdubuaku/maths-cs-ai-compendium repository provides intelligent content recommendations through its `recommend` tool. To maintain precision when suggesting relevant sections from the compendium, the server implements a stop-word filtering mechanism that removes generic terms before scoring potential matches.

## The Stop-Word Set (STOP_WORDS)

In [`mcp/src/index.ts`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/mcp/src/index.ts), the server defines a constant **`STOP_WORDS`** (lines 86-95) as a `Set<string>` containing high-frequency English words such as "the", "and", and "for", alongside domain-specific verbs like "understand" and "explain". This collection acts as a filter to prevent generic instructional language from diluting the relevance of technical keyword matches.

The use of a `Set` data structure ensures O(1) lookup time when checking tokens during the extraction phase, making the filtering process efficient even for longer queries.

## The Recommendation Pipeline

When the `recommend` tool is invoked, the server executes a multi-stage pipeline:

1. **Metadata Loading**: Parses the [`llms.txt`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/llms.txt) file to build a catalogue of sections with titles and descriptions.
2. **Tokenization**: Splits the user query into tokens using the regex `/\W+/` to separate words from punctuation.
3. **Filtering**: Removes tokens shorter than three characters or appearing in the `STOP_WORDS` set.
4. **Scoring**: Assigns points to sections based on filtered keyword matches.
5. **Grouping**: Sorts results by chapter and section number for logical reading order.

## Keyword Extraction and Stop-Word Filtering

The actual filtering occurs within the `recommend` tool implementation (lines 200-212 in [`mcp/src/index.ts`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/mcp/src/index.ts)). The server normalizes the query to lowercase, then applies length and stop-word checks:

```typescript
const keywords = query
  .toLowerCase()
  .split(/\W+/)                     // split on non-word characters
  .filter((w) => w.length > 2 && !STOP_WORDS.has(w));

```

This logic ensures that only substantive terms—those longer than two characters and not on the stop-word list—proceed to the scoring phase. The `filter` callback combines two conditions: `w.length > 2` eliminates short artifacts like "a" or "an" (if not already in the set), while `!STOP_WORDS.has(w)` removes the predefined common vocabulary.

## Scoring Sections with Filtered Keywords

Once extracted, the filtered keywords drive a weighted scoring algorithm. For each section in the metadata, the server calculates relevance as follows:

```typescript
const scored = meta.map((entry) => {
  const descLower = entry.description.toLowerCase();
  const nameLower = `${entry.chapterName} ${entry.sectionName}`.toLowerCase();

  let score = 0;
  for (const kw of keywords) {
    if (descLower.includes(kw)) score += 2;   // description match
    if (nameLower.includes(kw)) score += 3;   // title match
  }
  return { ...entry, score };
});

```

Sections receive **2 points** for each keyword appearing in the description and **3 points** for matches in the title or chapter name. Because stop words were removed prior to this loop, the scores reflect genuine topical overlap rather than coincidental matches with common English words.

## Result Formatting and Grouping

After scoring, the server filters out sections with zero or negative scores, then organizes the results by chapter:

```typescript
const lines: string[] = ["Recommended sections (in suggested reading order):\n"];
for (const [chNum, sections] of [...byChapter.entries()].sort((a, b) => a[0] - b[0])) {
  const ch = sections[0];
  sections.sort((a, b) => a.section - b.section);
  lines.push(`## Chapter ${chNum}: ${ch.chapterName}`);

  for (const sec of sections) {
    lines.push(`  ${sec.section}. ${sec.sectionName} — ${sec.description}`);
  }
  lines.push("");
}

```

This grouping ensures that users receive a structured reading list where priority is given to sections with the highest concentration of relevant technical terms.

## Summary

- The MCP server maintains a predefined **`STOP_WORDS`** set in [`mcp/src/index.ts`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/mcp/src/index.ts) (lines 86-95) containing common English words and educational verbs like "understand".
- During the **`recommend`** tool execution (lines 200-212), the server splits queries on non-word characters and filters out tokens shorter than three characters or present in the stop-word set.
- Only the remaining keywords participate in scoring, where description matches add 2 points and title/chapter matches add 3 points.
- Stop-word filtering prevents generic language from skewing results, ensuring recommendations target sections with genuine technical relevance to the user's learning goal.

## Frequently Asked Questions

### What specific words are included in the STOP_WORDS set?

The `STOP_WORDS` set includes common English articles and conjunctions such as "the", "and", and "for", as well as domain-specific instructional verbs like "understand" and "explain". These terms are defined as a `Set<string>` near the top of [`mcp/src/index.ts`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/mcp/src/index.ts) (lines 86-95) to ensure frequent educational vocabulary does not interfere with technical keyword matching.

### Why does the server filter tokens shorter than three characters?

The server applies a minimum length filter (`w.length > 2`) in addition to the stop-word check to remove short artifacts like "a", "an", or "to" that might slip through regex splitting, ensuring that only meaningful, substantive terms contribute to the section scoring algorithm.

### How does stop-word removal improve recommendation accuracy?

By removing generic vocabulary before scoring, the recommendation engine focuses exclusively on technical terms that differentiate one compendium section from another. This prevents common words from inflating scores for irrelevant sections and ensures that matches reflect genuine topical overlap with the user's specific learning goal.

### Where is the recommendation logic implemented in the codebase?

The core recommendation logic resides in [`mcp/src/index.ts`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/mcp/src/index.ts) within the HenryNdubuaku/maths-cs-ai-compendium repository. Specifically, the `STOP_WORDS` definition appears at lines 86-95, while the keyword extraction, filtering, and scoring implementation for the `recommend` tool occupies lines 200-212.