# What Regex Patterns Are Used by the MCP Server in the Maths-CS-AI Compendium?

> Discover the six regex patterns powering the MCP server in the Maths-CS-AI Compendium. Learn how they parse directories, files, and code blocks from the HenryNdubuaku repository.

- Repository: [Henry Ndubuaku/maths-cs-ai-compendium](https://github.com/HenryNdubuaku/maths-cs-ai-compendium)
- Tags: how-to-guide
- Published: 2026-07-16

---

**The MCP server in the HenryNdubuaku/maths-cs-ai-compendium repository uses six specific regular expressions defined in [`mcp/src/index.ts`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/mcp/src/index.ts) to parse chapter directories, section files, curated index entries, and code blocks.**

The Model Context Protocol (MCP) server powering the Maths-CS-AI Compendium relies on precise regex patterns to map the repository's file system into a structured knowledge base. These patterns enable the server to discover content hierarchies and extract metadata for LLM consumption without hardcoded file lists. Understanding these regex patterns reveals exactly how the server organizes topics and serves relevant educational content.

## Chapter and Section Discovery Patterns

The server uses two primary patterns to enumerate the repository's content structure during initialization.

### Chapter Directory Matching

The pattern `^chapter (\d{2}): (.+)$` at line 11 of [`mcp/src/index.ts`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/mcp/src/index.ts) matches folder names such as `chapter 01: Introduction`. This regex captures the two-digit chapter number in the first group and the chapter title in the second group, powering the `list_topics` tool by identifying valid chapter directories.

### Section File Matching

For markdown content files, the server applies `^(\d{2})\. (.+)\.md$` at line 12. This pattern identifies files like `01. What is AI?.md`, extracting the section number and title to filter the file system and surface only relevant educational content.

## Index File Parsing Patterns

The server parses [`llms.txt`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/llms.txt) to support the recommendation engine, using specialized regexes for structured headings and list entries.

### Chapter Heading Extraction

The pattern `^### Chapter (\d+): (.+)$` at line 65 detects chapter headers in the curated index file. Unlike the directory pattern, this one matches markdown headings (###) and uses a single digit capture group, enabling the server to group sections by chapter when building recommendation contexts.

### Section Entry Parsing

To extract section metadata, the server uses `^- \[(.+?)\]\(.+?\): (.+)$` at line 72. This pattern matches markdown list entries like `- [Attention](link): Description`, capturing the section title and its description. The non-greedy `(.+?)` ensures proper extraction of the link text for the `recommend` tool's keyword relevance ranking.

## Code Block Detection Patterns

To support the `get_examples` tool, the server extracts fenced code blocks from markdown content using `^```(\w*)$` at line 82. This pattern detects the start of a code block and optionally captures the language identifier (e.g., `python`, `cpp`, or `javascript`). The closing fence is detected through a simple line comparison in the extraction loop (lines 88-90), allowing the server to return complete code snippets with their associated language tags.

## Implementation Examples

Here are the actual TypeScript implementations from the source:

```typescript
// Chapter discovery (mcp/src/index.ts, line 11)
const match = entry.match(/^chapter (\d{2}): (.+)$/);
if (match) {
  const chapterNumber = match[1]; // e.g., "01"
  const chapterTitle = match[2];  // e.g., "Introduction"
}

```

```typescript
// Section file discovery (mcp/src/index.ts, line 12)
const match = entry.match(/^(\d{2})\. (.+)\.md$/);
if (match) {
  const sectionNumber = match[1];
  const sectionTitle = match[2];
}

```

```typescript
// llms.txt chapter parsing (mcp/src/index.ts, line 65)
if (line.match(/^### Chapter (\d+): (.+)$/)) {

  // Process chapter header from curated index
}

```

```typescript
// llms.txt section entry parsing (mcp/src/index.ts, line 72)
if (line.match(/^- \[(.+?)\]\(.+?\): (.+)$/)) {
  // Extract section metadata for recommendations
}

```

```typescript
// Code block detection (mcp/src/index.ts, line 82)
const openMatch = lines[i].match(/^```(\w*)$/);
const language = openMatch?.[1] || "text";

```

## Summary

- The MCP server uses **six distinct regex patterns** in [`mcp/src/index.ts`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/mcp/src/index.ts) to parse the repository structure.
- **Chapter and section discovery** relies on `^chapter (\d{2}): (.+)$` and `^(\d{2})\. (.+)\.md$` at lines 11-12.
- The **[`llms.txt`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/llms.txt) parser** uses `^### Chapter (\d+): (.+)$` (line 65) and `^- \[(.+?)\]\(.+?\): (.+)$` (line 72) to build the recommendation index.

- **Code extraction** depends on `^```(\w*)$` (line 82) to identify fenced blocks and their languages.
- These patterns enable the `list_topics`, `recommend`, and `get_examples` tools to function dynamically without static file lists.

## Frequently Asked Questions

### Where are the MCP server regex patterns defined?

All regex patterns are defined in `mcp/src/index.ts` within the HenryNdubuaku/maths-cs-ai-compendium repository. The chapter and section patterns appear at lines 11 and 12, while the `llms.txt` parsing patterns are located at lines 65 and 72. The code block detection regex is defined at line 82.

### How does the MCP server use regex to find content?

The server applies `^chapter (\d{2}): (.+)$` to identify valid chapter directories during file system enumeration. It then uses `^(\d{2})\. (.+)\.md$` to filter markdown files within those directories. This regex-driven discovery populates the `list_topics` tool and enables dynamic content loading without maintaining a static index.

### What regex pattern extracts code examples from the Compendium?

The pattern `^```(\w*)$` at line 82 of [`mcp/src/index.ts`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/mcp/src/index.ts) detects the opening of fenced code blocks. The optional capture group extracts the language identifier (such as `python` or `javascript`), while the closing fence is detected through line comparison in the extraction loop, enabling the `get_examples` tool to return complete code snippets.

### Why does the server parse llms.txt with regex instead of a YAML parser?

The server uses `^### Chapter (\d+): (.+)$` and `^- \[(.+?)\]\(.+?\): (.+)$` to parse the curated [`llms.txt`](https://github.com/HenryNdubuaku/maths-cs-ai-compendium/blob/main/llms.txt) index because the file follows a specific markdown structure designed for human readability and LLM context. Regex extraction provides a lightweight method to build the recommendation graph without importing heavy parsing dependencies, keeping the MCP server implementation lean.