How to Retrieve Blog Post Metadata Without an External CMS

You can retrieve blog post metadata directly from local MDX files using Node.js filesystem APIs and the gray-matter parser, eliminating the need for external content management systems while maintaining full static generation capabilities.

The woosal1337/blog repository demonstrates a complete implementation to retrieve blog post metadata without an external CMS by leveraging the Next.js App Router alongside file-based MDX content. By storing posts as version-controlled MDX files with YAML frontmatter, the application extracts structured metadata at build time using lightweight Node.js utilities, creating a fast, SEO-optimized static site that requires no third-party CMS infrastructure.

The Architecture: File-Based Metadata Storage

Content Directory Structure

Posts reside as individual directories within app/(website)/blog/(post)/, where each folder contains a page.mdx file. This structure leverages Next.js route groups—denoted by parentheses—to organize content without affecting the URL path. Each MDX file includes a YAML frontmatter block at the top defining metadata fields such as title, date, description, and tags, followed by the article content.

Metadata Schema and Type Safety

The frontmatter acts as the single source of truth for post metadata. When parsed by gray-matter, these key-value pairs become typed objects that populate the UI and SEO components. This approach ensures that all content remains in Git, enabling version control, pull request workflows, and offline development without API dependencies.

Extracting Metadata with Node.js and gray-matter

The Core Scanner Utility

The lib/blog-utils.ts file implements the primary extraction logic. It reads the filesystem synchronously during the build process, locates every MDX file, and parses the frontmatter using the gray-matter library. This function returns an array of post objects containing the metadata and raw content.

// lib/blog-utils.ts
import fs from "node:fs";
import path from "node:path";
import matter from "gray-matter";

const POSTS_DIR = path.join(process.cwd(), "app/(website)/blog/(post)");

export function getAllPosts() {
  const dirs = fs.readdirSync(POSTS_DIR);
  return dirs.map((slug) => {
    const mdxPath = path.join(POSTS_DIR, slug, "page.mdx");
    const file = fs.readFileSync(mdxPath, "utf8");
    const { data, content } = matter(file);
    return {
      slug,
      ...data,      // spreads title, date, description, tags, etc.
      content,      // raw MDX string for rendering
    };
  });
}

Data Abstraction Helpers

To keep page components clean, lib/blog.ts provides a thin abstraction layer over the scanner. These helper functions filter and retrieve specific posts by slug or return lists for index pages.

// lib/blog.ts
import { getAllPosts } from "./blog-utils";

export function getPostBySlug(slug: string) {
  const posts = getAllPosts();
  return posts.find((p) => p.slug === slug);
}

export function getAllSlugs() {
  const posts = getAllPosts();
  return posts.map((post) => post.slug);
}

Consuming Metadata in the Next.js App Router

Generate Static Routes

The dynamic route segments utilize generateStaticParams to build paths at compile time. By calling getAllSlugs(), Next.js creates a static page for every post directory found in the filesystem.

// app/(website)/blog/page.tsx
import { getAllSlugs } from "@/lib/blog";

export async function generateStaticParams() {
  const slugs = getAllSlugs();
  return slugs.map((slug) => ({ slug }));
}

Page-Level Metadata Exports

Individual MDX pages export a metadata object that Next.js automatically injects into the document head. While gray-matter parses the data for component consumption, this export handles standard meta tags and Open Graph properties.

export const metadata = {
  title: "The Next Sessions",
  description: "Exploring the future of continuous learning.",
};

---

# The Next Sessions

Post content begins here after the frontmatter...

SEO Enhancement with Structured Data

For rich search results, the components/seo/json-ld.tsx component transforms the parsed frontmatter into Schema.org JSON-LD structured data. This script injects article metadata—such as author, publication date, and description—directly into the page markup, improving visibility in search engine results without requiring external CMS plugins.

Summary

  • Store blog posts as MDX files with YAML frontmatter in the app/(website)/blog/(post)/ directory structure to keep content version-controlled
  • Parse metadata at build time using gray-matter inside lib/blog-utils.ts to avoid runtime database queries
  • Access specific posts via helper functions like getPostBySlug defined in lib/blog.ts
  • Generate static routes dynamically with generateStaticParams consuming getAllSlugs for optimal static site generation
  • Export Next.js metadata objects from MDX files and augment with JSON-LD in components/seo/json-ld.tsx for comprehensive SEO coverage

Frequently Asked Questions

What is YAML frontmatter and why does this approach use it?

YAML frontmatter is a block of metadata placed at the beginning of a markdown file between triple dashes that provides a human-readable way to define structured data like titles and dates alongside content. This approach uses it because gray-matter parses it efficiently into JavaScript objects at build time, enabling type-safe metadata handling without external databases.

How does file-based metadata retrieval compare to using a headless CMS?

File-based retrieval offers complete version control through Git, zero API latency during builds, and offline development capabilities, whereas a headless CMS requires network requests, rate limiting considerations, and external service dependencies. The trade-off is that non-technical editors must commit files to Git rather than using a web-based GUI, though Git-based CMS interfaces can bridge this gap.

Can this metadata retrieval method work with the Next.js Pages Router instead of the App Router?

Yes, the core utilities in lib/blog-utils.ts remain identical, but you would export getStaticProps from your page files instead of using generateStaticParams. The getAllPosts and getPostBySlug functions return serializable JSON that Next.js passes as props to your React components during static generation.

Is the gray-matter library necessary, or can frontmatter be parsed manually?

While you could implement a custom regex parser to extract YAML frontmatter, gray-matter is the industry standard because it handles edge cases like different newline formats, complex YAML structures, and delimiter detection reliably. For production applications, using gray-matter reduces maintenance burden and ensures consistent parsing behavior across different operating systems.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →