# How the Search Functionality Works in the CS Self-Learning Documentation

> Discover how the CS Self-Learning documentation search works. It uses static, client-side MkDocs-Material and lunr.js for a fast, compressed bilingual index without servers.

- Repository: [Yinmin Zhong/cs-self-learning](https://github.com/PKUFlyingPig/cs-self-learning)
- Tags: internals
- Published: 2026-03-02

---

**The CS Self-Learning documentation uses a static, client-side search powered by MkDocs-Material and lunr.js, generating a compressed bilingual search index at build time that requires no server-side infrastructure.**

The PKUFlyingPig/cs-self-learning repository provides a comprehensive curriculum for computer science self-study, organizing hundreds of resources across multiple languages. Understanding how the search functionality operates within this documentation helps contributors optimize content discoverability and enables users to navigate the extensive material efficiently.

## Architecture Overview

The search system is built entirely into the static site generation process. Unlike dynamic websites that query a database or search engine backend, this implementation relies on **MkDocs-Material**'s built-in client-side search capabilities. All indexing occurs during the build phase (`mkdocs build`), producing a self-contained search experience that works on any static file host.

## Configuration in mkdocs.yml

The behavior of the search functionality is controlled through specific settings in the [`mkdocs.yml`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/mkdocs.yml) file at the repository root.

### Disabling the Dedicated Search Page

Rather than generating a separate `/search/` page, the configuration embeds the search interface directly into the navigation bar of every page. This is achieved by setting `include_search_page: false` on lines 19–20 of [`mkdocs.yml`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/mkdocs.yml):

```yaml
theme:
  include_search_page: false
  search_index_only: true

```

### Generating the Search Index

When `search_index_only: true` is set (line 21), MkDocs-Material creates a compressed JSON file called [`search_index.json`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/search_index.json) during the build process. This file contains tokenized extracts of all headings, titles, and plain-text content from the Markdown files under the `docs/` directory. The theme ships only this index file rather than additional search page templates, keeping the deployment footprint minimal.

### Bilingual Language Support

The documentation serves both Chinese and English readers, so the **lunr.js** search engine must handle both character sets correctly. Lines 121–124 of [`mkdocs.yml`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/mkdocs.yml) configure the `search` plugin with dual language support:

```yaml
plugins:
  - search:
      lang:
        - zh
        - en

```

This configuration ensures that **lunr.js** builds separate language-specific indices, enabling proper tokenization that respects Chinese character boundaries while applying English word stemming algorithms.

### UI Enhancement Features

The theme enables three interactive features to improve the search experience, configured on lines 26–28:

```yaml
theme:
  features:
    - search.highlight
    - search.share
    - search.suggest

```

These options activate:
- **Highlight** — automatic highlighting of matching terms in result snippets
- **Share** — a button that copies a URL containing the current search query parameter
- **Suggest** — autocomplete suggestions that appear as users type in the search bar

## How the Search Index Is Built

During the static site generation process, MkDocs-Material traverses all Markdown files in the `docs/` directory. It extracts structured content including page titles, heading hierarchies (H1–H6), and body text paragraphs. This content is then normalized, tokenized according to the configured languages (Chinese and English), and serialized into the [`search_index.json`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/search_index.json) file.

The index supports additional metadata fields such as `tags` or `categories` when defined in a page's front-matter, though the current repository relies primarily on the default behavior of indexing titles, headings, and body text.

## Client-Side Search Execution Flow

When a user interacts with the search bar, the following process occurs entirely within the browser:

1. **Index Loading** — The JavaScript loads the [`search_index.json`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/search_index.json) file asynchronously
2. **Query Parsing** — **lunr.js** processes the user's query using the appropriate language analyzer (Chinese or English) to handle tokenization and stemming
3. **Scoring and Ranking** — The engine calculates relevance scores for each document based on term frequency and field weights (titles and headings rank higher than body text)
4. **Result Rendering** — Top matches render in a dropdown overlay; pressing **Enter** displays a full results page
5. **Navigation and Highlighting** — Clicking a result navigates to the target page with matched terms automatically highlighted

## Customizing Search Behavior

Contributors can extend the search functionality through optional configuration changes or front-matter additions.

### Adding Custom Keywords via Front-Matter

To improve discoverability of specific pages, add a `tags` field to any Markdown file's front-matter:

```yaml
---
title: "Stanford CS229: Machine Learning"
tags: [machine-learning, Andrew-Ng, regression]
---

```

These tags are automatically indexed and support targeted queries using syntax like `tags:regression`.

### Enabling a Dedicated Search Page

If you fork the repository and prefer a standalone search page, modify [`mkdocs.yml`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/mkdocs.yml) as follows:

```yaml
theme:
  include_search_page: true   # changed from false

  search_index_only: false    # provide full search page template

```

After rebuilding the site with `mkdocs build`, a dedicated `/search/` page becomes available alongside the navigation bar search.

### Debugging the Search Index Programmatically

For development or troubleshooting, you can inspect the generated index directly using browser developer tools:

```html
<script>
  fetch('/search_index.json')
    .then(r => r.json())
    .then(idx => {
      const lunr = window.lunr;
      const search = lunr(function () {
        this.use(lunr.zh);   // Use Chinese language support
        this.ref('location');
        this.field('title');
        this.field('content');
        idx.forEach(doc => this.add(doc));
      });
      console.log(search.search('深度学习'));
    });
</script>

```

## Summary

- The search functionality in PKUFlyingPig/cs-self-learning is **client-side and static**, requiring no backend server or database.
- Configuration resides in **[`mkdocs.yml`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/mkdocs.yml)** (lines 19–21, 26–28, and 121–124), controlling index generation, UI features, and bilingual support.
- The build process generates **[`search_index.json`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/search_index.json)**, a compressed index of all documentation content in Chinese and English.
- **lunr.js** powers the search engine, handling language-specific tokenization for both character sets.
- Users can extend search metadata through **front-matter tags** or modify the theme settings to enable a dedicated search page.

## Frequently Asked Questions

### Does the search require a server to function?

No. The search operates entirely within the user's browser using JavaScript. Once the static site is built, the [`search_index.json`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/search_index.json) file and the lunr.js library handle all querying and ranking without server-side processing, making it compatible with GitHub Pages, CDNs, or any static file host.

### How does the search handle Chinese and English content simultaneously?

The MkDocs-Material `search` plugin is configured in [`mkdocs.yml`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/mkdocs.yml) with both `zh` (Chinese) and `en` (English) language options. During indexing, lunr.js creates separate tokenization pipelines for each language, ensuring that Chinese characters are segmented correctly while English words undergo stemming. The engine can match queries across both languages within the same index.

### Where is the search index stored and how large is it?

The search index is generated as **[`search_index.json`](https://github.com/PKUFlyingPig/cs-self-learning/blob/main/search_index.json)** in the site root during the `mkdocs build` process. Because MkDocs-Material compresses the index and the repository primarily contains text-based Markdown files, the resulting JSON file remains lightweight (typically under a few megabytes), ensuring fast initial page loads even on slower connections.

### Can I search for content within specific sections or categories?

While the default configuration indexes all titles, headings, and body text, you can enable filtered searching by adding **front-matter metadata** to your Markdown files. Define `tags` or custom fields in the YAML header, and lunr.js will index these values. Users can then perform targeted searches using field-specific syntax like `tags:distributed-systems` to narrow results.