# How the Content-Structure Integration Parses Markdown and Generates Navigation in Astro Big Doc

> Discover how Astro Big Doc's content-structure integration parses markdown files and builds a hierarchical JSON tree for your site navigation. Learn more about this efficient process.

- Repository: [Micro Web Stacks/astro-big-doc](https://github.com/microwebstacks/astro-big-doc)
- Tags: deep-dive
- Published: 2026-03-07

---

**The content-structure integration walks the `content/` directory, extracts front-matter from Markdown files, and transforms the flat file list into a hierarchical JSON tree consumed by the site's navigation components.**

The astro-big-doc repository uses a custom **content-structure** integration to automate navigation generation. This integration eliminates manual menu maintenance by parsing Markdown front-matter and file system structure to produce a dynamic navigation tree that drives the sidebar components.

## Integration Architecture Overview

The content-structure integration operates as an Astro integration registered in `astro.config.mjs`. It hooks into the `astro:config:setup` lifecycle event to execute a three-phase pipeline before the site builds. The integration delegates Markdown parsing to the external `content-structure` package while handling menu transformation through internal utilities in [`integrations/create_menu.js`](https://github.com/microwebstacks/astro-big-doc/blob/main/integrations/create_menu.js) and [`integrations/process_menu.js`](https://github.com/microwebstacks/astro-big-doc/blob/main/integrations/process_menu.js).

## Step 1: Collecting and Parsing Markdown Files

### The collect() Function

During `astro:config:setup`, the integration calls `collect(collect_config)` from the `content-structure` package. This function traverses the directory specified in `collect_config.contentdir` (default: `content/`), reads each `.md` file, and extracts structured metadata.

The collection process emits a JSON file named [`document_list.json`](https://github.com/microwebstacks/astro-big-doc/blob/main/document_list.json) to `collect_config.outdir`. This flat list serves as the single source of truth for all subsequent menu generation steps.

### Extracted Front-Matter Fields

The parser extracts the following fields from each Markdown file's front-matter and file system properties:

| Field | Source | Purpose |
|-------|--------|---------|
| `title` | Front-matter `title:` | Display label in navigation |
| `path` | File system relative path | Internal reference |
| `url` | Derived from file name | URL slug for routing |
| `url_type` | File system check | `"file"` or `"dir"` classification |
| `level` | Heading depth calculation | Hierarchy depth (1 for home, 2 for sections) |
| `order` | Front-matter `order:` | Sorting priority (default 100) |

## Step 2: Building the Raw Menu Structure

### Loading menu.yaml or Auto-Generating

The `create_menu()` function in [`integrations/create_menu.js`](https://github.com/microwebstacks/astro-big-doc/blob/main/integrations/create_menu.js) begins by checking for a manual menu definition at [`content/menu.yaml`](https://github.com/microwebstacks/astro-big-doc/blob/main/content/menu.yaml). If present, it loads this file directly via `load_yaml_abs()`. If absent, it generates a default menu structure from the previously created [`document_list.json`](https://github.com/microwebstacks/astro-big-doc/blob/main/document_list.json).

```javascript
const menu_file = join(collect_config.contentdir, "menu.yaml");
if (await exists(menu_file)) {
  raw_menu = await load_yaml_abs(menu_file);
} else {
  const document_list = await load_json_abs(join(collect_config.outdir, "document_list.json"));
  raw_menu = await create_raw_menu(collect_config.contentdir, document_list);
}

```

### Processing Auto-Generated Sections

For entries marked with `autogenerate: { directory: "." }`, the integration queries `getDocuments({format:"markdown"})` to retrieve all Markdown documents. It maps each document to a navigation item object containing `label`, `path`, `url`, `level`, and `order` properties.

The `link` property is constructed by combining the base path, section name, and document URL using `add_base()` from [`src/libs/assets.js`](https://github.com/microwebstacks/astro-big-doc/blob/main/src/libs/assets.js). Each item defaults to `order: 100` if not specified in front-matter.

## Step 3: Converting Flat Lists to Hierarchical Trees

### The pages_list_to_tree() Algorithm

The `pages_list_to_tree()` function in [`integrations/process_menu.js`](https://github.com/microwebstacks/astro-big-doc/blob/main/integrations/process_menu.js) transforms the flat list of navigation items into a nested tree structure suitable for recursive rendering. The algorithm executes five distinct phases:

1. **Parent Injection**: Identifies missing parent folders via `get_new_parents()` and appends them to the entries array until all parents exist.
2. **Initialization**: Prepares each entry with empty `items` arrays and flags (`parent: true`, `expanded: true`).
3. **Tree Construction**: Iterates entries and attaches children to their parents using `get_parent()`. Entries with `level > 2` become children; others populate the root array.
4. **Leaf Pruning**: Removes empty `items` arrays and flags from leaf nodes to minimize payload size.
5. **Sorting**: Sorts children and root items by the `order` field using `sort((a,b) => a.order - b.order)`.

The resulting tree preserves the hierarchical relationships defined by file system paths and front-matter order values.

## Step 4: Output and Client-Side Consumption

The final navigation object contains `base_menu` (top-level links) and `sections` (section-specific trees). The integration generates an MD5 hash of the JSON content for cache-busting and writes the result to `collect_config.out_menu` (default: [`public/menu.json`](https://github.com/microwebstacks/astro-big-doc/blob/main/public/menu.json)).

Client-side components `SideMenu.astro` and `ClientNavMenu.astro` load this JSON at runtime. They render `base_menu` as the primary navigation and lookup the current section in `menu.sections` to display the expandable sidebar tree.

## Configuration and Customization

Register the integration in your Astro configuration:

```javascript
// astro.config.mjs
import { defineConfig } from 'astro/config';
import { config } from './config.js';
import { collect_content } from './integrations/integration-content-structure.js';
import yaml from '@rollup/plugin-yaml';

export default defineConfig({
  integrations: [collect_content(config.collect_content)],
  output: "static",
  outDir: config.outDir,
  base: config.base,
  trailingSlash: 'ignore',
  vite: { plugins: [yaml()] },
});

```

Optionally define a manual menu structure in [`content/menu.yaml`](https://github.com/microwebstacks/astro-big-doc/blob/main/content/menu.yaml):

```yaml
- label: Home
  link: /
  autogenerate:
    directory: .
- label: About
  link: /about/
  items:
    - label: Team
      link: /about/team/
    - label: License
      link: /about/license/

```

## Summary

- The **content-structure integration** hooks into Astro's `astro:config:setup` lifecycle to preprocess Markdown before the build starts.
- **`collect()`** parses front-matter from all Markdown files in `content/` and emits [`document_list.json`](https://github.com/microwebstacks/astro-big-doc/blob/main/document_list.json).
- **`create_menu()`** either loads a manual [`menu.yaml`](https://github.com/microwebstacks/astro-big-doc/blob/main/menu.yaml) or auto-generates navigation entries, mapping documents to navigation items with `order`, `level`, and `link` properties.
- **`pages_list_to_tree()`** transforms the flat list into a hierarchical tree by injecting missing parents, attaching children, pruning leaves, and sorting by the `order` field.
- The final [`menu.json`](https://github.com/microwebstacks/astro-big-doc/blob/main/menu.json) is consumed by `SideMenu.astro` and `ClientNavMenu.astro` to render the site's navigation sidebar.

## Frequently Asked Questions

### How does the content-structure integration handle missing parent folders?

The integration detects missing parent directories during the tree conversion phase. The `pages_list_to_tree()` function in [`integrations/process_menu.js`](https://github.com/microwebstacks/astro-big-doc/blob/main/integrations/process_menu.js) calls `get_new_parents()` to identify gaps in the hierarchy and injects placeholder parent entries until every child has a valid parent node. This ensures the navigation tree remains complete even when intermediate index files are absent.

### Can I manually define the navigation instead of auto-generating it?

Yes. If you create a [`menu.yaml`](https://github.com/microwebstacks/astro-big-doc/blob/main/menu.yaml) file in your `content/` directory, the integration will load it directly via `load_yaml_abs()` instead of generating a menu from the document list. This allows you to specify exact labels, links, and nested `items` arrays. You can also mix manual entries with `autogenerate` properties to dynamically populate specific sections while keeping others manually curated.

### What front-matter fields are required for the navigation to work?

Only the `title` field is strictly required, as it provides the display label for navigation items. The `order` field is optional but recommended for controlling sort order; entries without an explicit `order` value default to `100`. The integration automatically derives `path`, `url`, `level`, and `url_type` from the file system, so you do not need to specify these manually unless overriding defaults.

### How is the navigation menu cached or updated during development?

The integration generates an MD5 hash of the final menu JSON using `createHash('md5')` and appends the first eight characters to the output. During development, the integration runs inside `astro:config:setup` every time Astro restarts or when content changes trigger a rebuild. The `collect()` and `create_menu()` functions execute synchronously during the setup phase, ensuring [`menu.json`](https://github.com/microwebstacks/astro-big-doc/blob/main/menu.json) is always current before the client-side components request it.