# How to Retrieve Blog Post Metadata Without an External CMS

> Retrieve blog post metadata from local MDX files using Node.js and gray-matter. Get static generation without an external CMS. Learn how to manage your content efficiently.

- Repository: [Ege Chelebi/blog](https://github.com/woosal1337/blog)
- Tags: how-to-guide
- Published: 2026-08-06

---

**You can retrieve blog post metadata directly from local MDX files using Node.js filesystem APIs and the `gray-matter` parser, eliminating the need for external content management systems while maintaining full static generation capabilities.**

The `woosal1337/blog` repository demonstrates a complete implementation to retrieve blog post metadata without an external CMS by leveraging the Next.js App Router alongside file-based MDX content. By storing posts as version-controlled MDX files with YAML frontmatter, the application extracts structured metadata at build time using lightweight Node.js utilities, creating a fast, SEO-optimized static site that requires no third-party CMS infrastructure.

## The Architecture: File-Based Metadata Storage

### Content Directory Structure

Posts reside as individual directories within `app/(website)/blog/(post)/`, where each folder contains a `page.mdx` file. This structure leverages Next.js route groups—denoted by parentheses—to organize content without affecting the URL path. Each MDX file includes a YAML frontmatter block at the top defining metadata fields such as `title`, `date`, `description`, and `tags`, followed by the article content.

### Metadata Schema and Type Safety

The frontmatter acts as the single source of truth for post metadata. When parsed by `gray-matter`, these key-value pairs become typed objects that populate the UI and SEO components. This approach ensures that all content remains in Git, enabling version control, pull request workflows, and offline development without API dependencies.

## Extracting Metadata with Node.js and gray-matter

### The Core Scanner Utility

The [`lib/blog-utils.ts`](https://github.com/woosal1337/blog/blob/main/lib/blog-utils.ts) file implements the primary extraction logic. It reads the filesystem synchronously during the build process, locates every MDX file, and parses the frontmatter using the `gray-matter` library. This function returns an array of post objects containing the metadata and raw content.

```typescript
// lib/blog-utils.ts
import fs from "node:fs";
import path from "node:path";
import matter from "gray-matter";

const POSTS_DIR = path.join(process.cwd(), "app/(website)/blog/(post)");

export function getAllPosts() {
  const dirs = fs.readdirSync(POSTS_DIR);
  return dirs.map((slug) => {
    const mdxPath = path.join(POSTS_DIR, slug, "page.mdx");
    const file = fs.readFileSync(mdxPath, "utf8");
    const { data, content } = matter(file);
    return {
      slug,
      ...data,      // spreads title, date, description, tags, etc.
      content,      // raw MDX string for rendering
    };
  });
}

```

### Data Abstraction Helpers

To keep page components clean, [`lib/blog.ts`](https://github.com/woosal1337/blog/blob/main/lib/blog.ts) provides a thin abstraction layer over the scanner. These helper functions filter and retrieve specific posts by slug or return lists for index pages.

```typescript
// lib/blog.ts
import { getAllPosts } from "./blog-utils";

export function getPostBySlug(slug: string) {
  const posts = getAllPosts();
  return posts.find((p) => p.slug === slug);
}

export function getAllSlugs() {
  const posts = getAllPosts();
  return posts.map((post) => post.slug);
}

```

## Consuming Metadata in the Next.js App Router

### Generate Static Routes

The dynamic route segments utilize `generateStaticParams` to build paths at compile time. By calling `getAllSlugs()`, Next.js creates a static page for every post directory found in the filesystem.

```typescript
// app/(website)/blog/page.tsx
import { getAllSlugs } from "@/lib/blog";

export async function generateStaticParams() {
  const slugs = getAllSlugs();
  return slugs.map((slug) => ({ slug }));
}

```

### Page-Level Metadata Exports

Individual MDX pages export a `metadata` object that Next.js automatically injects into the document head. While `gray-matter` parses the data for component consumption, this export handles standard meta tags and Open Graph properties.

```mdx
export const metadata = {
  title: "The Next Sessions",
  description: "Exploring the future of continuous learning.",
};

---

# The Next Sessions

Post content begins here after the frontmatter...

```

## SEO Enhancement with Structured Data

For rich search results, the [`components/seo/json-ld.tsx`](https://github.com/woosal1337/blog/blob/main/components/seo/json-ld.tsx) component transforms the parsed frontmatter into Schema.org JSON-LD structured data. This script injects article metadata—such as author, publication date, and description—directly into the page markup, improving visibility in search engine results without requiring external CMS plugins.

## Summary

- Store blog posts as MDX files with YAML frontmatter in the `app/(website)/blog/(post)/` directory structure to keep content version-controlled
- Parse metadata at build time using `gray-matter` inside [`lib/blog-utils.ts`](https://github.com/woosal1337/blog/blob/main/lib/blog-utils.ts) to avoid runtime database queries
- Access specific posts via helper functions like `getPostBySlug` defined in [`lib/blog.ts`](https://github.com/woosal1337/blog/blob/main/lib/blog.ts)
- Generate static routes dynamically with `generateStaticParams` consuming `getAllSlugs` for optimal static site generation
- Export Next.js metadata objects from MDX files and augment with JSON-LD in [`components/seo/json-ld.tsx`](https://github.com/woosal1337/blog/blob/main/components/seo/json-ld.tsx) for comprehensive SEO coverage

## Frequently Asked Questions

### What is YAML frontmatter and why does this approach use it?

YAML frontmatter is a block of metadata placed at the beginning of a markdown file between triple dashes that provides a human-readable way to define structured data like titles and dates alongside content. This approach uses it because `gray-matter` parses it efficiently into JavaScript objects at build time, enabling type-safe metadata handling without external databases.

### How does file-based metadata retrieval compare to using a headless CMS?

File-based retrieval offers complete version control through Git, zero API latency during builds, and offline development capabilities, whereas a headless CMS requires network requests, rate limiting considerations, and external service dependencies. The trade-off is that non-technical editors must commit files to Git rather than using a web-based GUI, though Git-based CMS interfaces can bridge this gap.

### Can this metadata retrieval method work with the Next.js Pages Router instead of the App Router?

Yes, the core utilities in [`lib/blog-utils.ts`](https://github.com/woosal1337/blog/blob/main/lib/blog-utils.ts) remain identical, but you would export `getStaticProps` from your page files instead of using `generateStaticParams`. The `getAllPosts` and `getPostBySlug` functions return serializable JSON that Next.js passes as props to your React components during static generation.

### Is the gray-matter library necessary, or can frontmatter be parsed manually?

While you could implement a custom regex parser to extract YAML frontmatter, `gray-matter` is the industry standard because it handles edge cases like different newline formats, complex YAML structures, and delimiter detection reliably. For production applications, using `gray-matter` reduces maintenance burden and ensures consistent parsing behavior across different operating systems.