How the MCP Recommends Relevant Sections from the Maths‑CS‑AI Compendium

The MCP server uses a keyword‑scoring algorithm that parses the llms.txt index, weights matches in section titles and descriptions, and returns the top 15 results grouped by chapter.

The HenryNdubuaku/maths‑cs‑ai‑compendium repository ships a Model‑Context‑Protocol (MCP) server that exposes a recommend tool. This tool enables AI assistants to suggest specific reading material from the compendium by matching natural‑language queries against a structured index of sections and chapters.

How the Recommend Tool Works

At its core, the recommendation engine transforms a user query into weighted keywords, scores every section in the compendium, and surfaces the strongest matches. The algorithm prioritizes semantic alignment over exact string matching by using a custom scoring heuristic implemented in the MCP server.

Step‑by‑Step Recommendation Algorithm

Metadata Extraction from llms.txt

When the recommend tool executes, it first loads and parses the llms.txt file located at the repository root. The parser extracts five fields for every section to build an array of SectionMeta objects:

  • Chapter number
  • Chapter name
  • Section number
  • Section name
  • Short description

This extraction logic resides in mcp/src/index.ts between lines 58 and 84. The resulting metadata serves as the searchable knowledge base for all subsequent scoring operations.

Query Tokenization and Stop‑Word Filtering

The user’s query undergoes aggressive normalization before scoring begins. The pipeline lower‑cases the input, splits it on non‑word characters, and discards any token shorter than three characters. It also filters out common stop words defined in the STOP_WORDS constant (lines 86‑95 of mcp/src/index.ts).

This preprocessing ensures that only meaningful keywords contribute to the final score, reducing noise from articles and prepositions.

Keyword Scoring Logic

For each SectionMeta entry, the algorithm calculates a cumulative score based on keyword presence:

  • +2 points for every keyword appearing in the section’s description
  • +3 points for every keyword appearing in the combined chapter and section name

This weighting scheme favors matches in titles over description matches, assuming that explicit keyword mentions in headings indicate higher relevance. The scoring loop is implemented on lines 21‑27 of mcp/src/index.ts.

Ranking and Filtering Results

After scoring, the system drops any section with a score of zero. The remaining entries are sorted by descending score, with ties broken by chapter number. The algorithm retains only the top 15 results to keep the response concise. This filtering and ranking logic appears on lines 29‑33 of mcp/src/index.ts.

Grouping by Chapter for Readable Output

The selected sections are reorganized for human consumption. The code groups results by chapter, sorts sections within each chapter by section number, and formats them into a bulleted list. The final output always begins with the header: Recommended sections (in suggested reading order):.

This presentation logic occupies lines 38‑55 of mcp/src/index.ts. The tool returns a JSON object containing a single content element of type "text" with the formatted string (lines 55‑56).

Implementation Details in the Source Code

The recommendation engine is self‑contained within mcp/src/index.ts. It relies on the llms.txt file as its sole data source, requiring no external vector database or embedding model. This design keeps the MCP server lightweight while still providing semantic relevance through careful keyword weighting.

Key constants like STOP_WORDS and the scoring multipliers (2 for description, 3 for titles) are hard‑coded, making the behavior deterministic and easy to audit.

Calling the Recommend Tool

You can invoke the recommendation engine from any MCP client using the standard SDK:

import { McpClient } from "@modelcontextprotocol/sdk/client";

const client = new McpClient({ transport: "stdio" });
await client.connect();

const response = await client.callTool("recommend", {
  query: "How do transformers work?",
});

console.log(response.content[0].text);

The tool returns a formatted text block similar to this:


Recommended sections (in suggested reading order):

## Chapter 10: multimodal learning

  02. vision language models — How to combine text and images in a single model
  03. image and video tokenisation — Tokenising visual data for transformers

## Chapter 12: graph neural networks

  04. graph attention networks — Extending attention mechanisms to graph structures

Summary

  • The MCP recommend tool parses llms.txt to build a metadata index of all compendium sections.
  • User queries are tokenized and filtered against a STOP_WORDS list to extract meaningful keywords.
  • Sections receive +2 points for description matches and +3 points for title matches.
  • Only the top 15 scoring sections are returned, sorted by chapter and section number.
  • The implementation lives entirely in mcp/src/index.ts and requires no external AI services.

Frequently Asked Questions

What file does the MCP use to know about compendium sections?

The MCP reads the llms.txt file at the repository root. This plain‑text index lists every chapter, section, and description, allowing the server to build an in‑memory array of SectionMeta objects without scanning the actual PDF or markdown content.

How does the MCP prevent common words from skewing results?

The algorithm references a STOP_WORDS constant defined around lines 86‑95 of mcp/src/index.ts. During tokenization, it removes any word in this set, ensuring that terms like “the,” “and,” or “how” do not contribute to the relevance score.

Why are title matches worth more than description matches?

Title matches receive +3 points while description matches receive +2 points. This reflects the assumption that if a keyword appears in a chapter or section heading, that section is likely a primary source on the topic, whereas description matches may only touch on the subject tangentially.

How many recommendations does the tool return by default?

The server limits output to the top 15 sections with positive scores. This cap is hard‑coded in the filtering logic on lines 29‑33 of mcp/src/index.ts to balance comprehensiveness with readability.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →