How the Content-Structure Integration Parses Markdown and Generates Navigation in Astro Big Doc

The content-structure integration walks the content/ directory, extracts front-matter from Markdown files, and transforms the flat file list into a hierarchical JSON tree consumed by the site's navigation components.

The astro-big-doc repository uses a custom content-structure integration to automate navigation generation. This integration eliminates manual menu maintenance by parsing Markdown front-matter and file system structure to produce a dynamic navigation tree that drives the sidebar components.

Integration Architecture Overview

The content-structure integration operates as an Astro integration registered in astro.config.mjs. It hooks into the astro:config:setup lifecycle event to execute a three-phase pipeline before the site builds. The integration delegates Markdown parsing to the external content-structure package while handling menu transformation through internal utilities in integrations/create_menu.js and integrations/process_menu.js.

Step 1: Collecting and Parsing Markdown Files

The collect() Function

During astro:config:setup, the integration calls collect(collect_config) from the content-structure package. This function traverses the directory specified in collect_config.contentdir (default: content/), reads each .md file, and extracts structured metadata.

The collection process emits a JSON file named document_list.json to collect_config.outdir. This flat list serves as the single source of truth for all subsequent menu generation steps.

Extracted Front-Matter Fields

The parser extracts the following fields from each Markdown file's front-matter and file system properties:

Field Source Purpose
title Front-matter title: Display label in navigation
path File system relative path Internal reference
url Derived from file name URL slug for routing
url_type File system check "file" or "dir" classification
level Heading depth calculation Hierarchy depth (1 for home, 2 for sections)
order Front-matter order: Sorting priority (default 100)

Step 2: Building the Raw Menu Structure

Loading menu.yaml or Auto-Generating

The create_menu() function in integrations/create_menu.js begins by checking for a manual menu definition at content/menu.yaml. If present, it loads this file directly via load_yaml_abs(). If absent, it generates a default menu structure from the previously created document_list.json.

const menu_file = join(collect_config.contentdir, "menu.yaml");
if (await exists(menu_file)) {
  raw_menu = await load_yaml_abs(menu_file);
} else {
  const document_list = await load_json_abs(join(collect_config.outdir, "document_list.json"));
  raw_menu = await create_raw_menu(collect_config.contentdir, document_list);
}

Processing Auto-Generated Sections

For entries marked with autogenerate: { directory: "." }, the integration queries getDocuments({format:"markdown"}) to retrieve all Markdown documents. It maps each document to a navigation item object containing label, path, url, level, and order properties.

The link property is constructed by combining the base path, section name, and document URL using add_base() from src/libs/assets.js. Each item defaults to order: 100 if not specified in front-matter.

Step 3: Converting Flat Lists to Hierarchical Trees

The pages_list_to_tree() Algorithm

The pages_list_to_tree() function in integrations/process_menu.js transforms the flat list of navigation items into a nested tree structure suitable for recursive rendering. The algorithm executes five distinct phases:

  1. Parent Injection: Identifies missing parent folders via get_new_parents() and appends them to the entries array until all parents exist.
  2. Initialization: Prepares each entry with empty items arrays and flags (parent: true, expanded: true).
  3. Tree Construction: Iterates entries and attaches children to their parents using get_parent(). Entries with level > 2 become children; others populate the root array.
  4. Leaf Pruning: Removes empty items arrays and flags from leaf nodes to minimize payload size.
  5. Sorting: Sorts children and root items by the order field using sort((a,b) => a.order - b.order).

The resulting tree preserves the hierarchical relationships defined by file system paths and front-matter order values.

Step 4: Output and Client-Side Consumption

The final navigation object contains base_menu (top-level links) and sections (section-specific trees). The integration generates an MD5 hash of the JSON content for cache-busting and writes the result to collect_config.out_menu (default: public/menu.json).

Client-side components SideMenu.astro and ClientNavMenu.astro load this JSON at runtime. They render base_menu as the primary navigation and lookup the current section in menu.sections to display the expandable sidebar tree.

Configuration and Customization

Register the integration in your Astro configuration:

// astro.config.mjs
import { defineConfig } from 'astro/config';
import { config } from './config.js';
import { collect_content } from './integrations/integration-content-structure.js';
import yaml from '@rollup/plugin-yaml';

export default defineConfig({
  integrations: [collect_content(config.collect_content)],
  output: "static",
  outDir: config.outDir,
  base: config.base,
  trailingSlash: 'ignore',
  vite: { plugins: [yaml()] },
});

Optionally define a manual menu structure in content/menu.yaml:

- label: Home
  link: /
  autogenerate:
    directory: .
- label: About
  link: /about/
  items:
    - label: Team
      link: /about/team/
    - label: License
      link: /about/license/

Summary

  • The content-structure integration hooks into Astro's astro:config:setup lifecycle to preprocess Markdown before the build starts.
  • collect() parses front-matter from all Markdown files in content/ and emits document_list.json.
  • create_menu() either loads a manual menu.yaml or auto-generates navigation entries, mapping documents to navigation items with order, level, and link properties.
  • pages_list_to_tree() transforms the flat list into a hierarchical tree by injecting missing parents, attaching children, pruning leaves, and sorting by the order field.
  • The final menu.json is consumed by SideMenu.astro and ClientNavMenu.astro to render the site's navigation sidebar.

Frequently Asked Questions

How does the content-structure integration handle missing parent folders?

The integration detects missing parent directories during the tree conversion phase. The pages_list_to_tree() function in integrations/process_menu.js calls get_new_parents() to identify gaps in the hierarchy and injects placeholder parent entries until every child has a valid parent node. This ensures the navigation tree remains complete even when intermediate index files are absent.

Can I manually define the navigation instead of auto-generating it?

Yes. If you create a menu.yaml file in your content/ directory, the integration will load it directly via load_yaml_abs() instead of generating a menu from the document list. This allows you to specify exact labels, links, and nested items arrays. You can also mix manual entries with autogenerate properties to dynamically populate specific sections while keeping others manually curated.

What front-matter fields are required for the navigation to work?

Only the title field is strictly required, as it provides the display label for navigation items. The order field is optional but recommended for controlling sort order; entries without an explicit order value default to 100. The integration automatically derives path, url, level, and url_type from the file system, so you do not need to specify these manually unless overriding defaults.

How is the navigation menu cached or updated during development?

The integration generates an MD5 hash of the final menu JSON using createHash('md5') and appends the first eight characters to the output. During development, the integration runs inside astro:config:setup every time Astro restarts or when content changes trigger a rebuild. The collect() and create_menu() functions execute synchronously during the setup phase, ensuring menu.json is always current before the client-side components request it.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →