How validate-markdown.py Ensures Markdown Syntax Integrity in Claude-Skills

The validate-markdown.py script performs four targeted structural checks—detecting HTML comments in tables, missing separator rows, column mismatches, and unclosed code fences—to prevent rendering failures in Markdown files.

The claude-skills repository relies on pristine Markdown documentation to ensure AI assistants and developers can consume skill definitions without parsing errors. At scripts/validate-markdown.py, a dedicated validation utility implements a two-pass scanning algorithm that catches common syntactic pitfalls before they reach production pipelines.

The Four Core Validation Checks in validate-markdown.py

The script organizes its logic around four discrete IssueType categories, each handled by specific detector functions within the source.

Detecting HTML Comments Inside Tables

HTML comments (<!-- … -->) placed between table headers and separator rows, or within data rows, cause Markdown parsers to abort table rendering entirely. The script identifies these violations using is_html_comment() (lines 59–62), which scans for comment delimiters while processing table contexts. When detected, the validator records an HTML_IN_TABLE issue (lines 121–131) pointing to the exact line number where the comment interrupts table structure.

Validating Table Separator Rows

Every Markdown table must contain a separator line (|---|---|) immediately following the header row. The script enforces this requirement through coordinated checks: is_table_row() (lines 47–50) identifies potential headers, then is_separator_row() (lines 53–56) validates the subsequent line. If the separator is absent, validate-markdown.py raises a MISSING_SEPARATOR issue (lines 134–144) to prevent downstream rendering of malformed tables.

Ensuring Column Count Consistency

Data rows must maintain the same column count as their parent headers. The script implements count_columns() (lines 41–45) to tally pipe characters while ignoring escaped pipes (\|). During the table scanning phase, each data row’s column count is compared against the stored header_cols value. Discrepancies trigger a COLUMN_MISMATCH issue (lines 167–178), ensuring tables remain structurally coherent across all rows.

Trapping Unclosed Fenced Code Blocks

A code block initiated with triple backticks (```) must be properly closed; otherwise, the remainder of the file is consumed as code content. The script performs a dedicated first pass (lines 71–80) that toggles a boolean in_code_block flag whenever fence delimiters are encountered. If the flag remains True after processing all lines, validate-markdown.py generates an UNCLOSED_CODE_BLOCK issue pointing to the opening fence location (lines 81–89).

How validate-markdown.py Processes Files

The validation pipeline operates through a structured two-pass approach implemented in validate_file() (lines 68–70).

First Pass: Code Block Sanity The script iterates through all lines to detect dangling code fences before any other processing occurs. This prevents false positives in subsequent table checks, as content inside code blocks should not be parsed as Markdown syntax.

Second Pass: Table Structure Validation Using a manual index variable i to maintain context, the script walks through lines again (lines 95–182). It skips content inside fenced blocks using the in_code_block toggle (lines 99–103). When is_table_row() identifies a potential table header, the script enters a validation loop that checks for separators, HTML comments, and column consistency across subsequent rows.

Issue Aggregation All discovered MarkdownIssue objects are collected and returned to the caller (line 84). The validate_directory() function (lines 87–95) extends this capability recursively using Path.rglob("*.md") to process entire documentation trees.

Running validate-markdown.py from the Command Line

The script exposes a CLI through main() (lines 98–63) that supports flexible execution patterns for individual files or entire directories.


# Validate the default "skills" folder with human-readable output

python scripts/validate-markdown.py --check

# Validate a specific directory and output JSON for CI integration

python scripts/validate-markdown.py --path docs --format json

# Check a single file (useful for pre-commit hooks)

python scripts/validate-markdown.py --check --path README.md

The script exits with status 0 when no issues are detected, or 1 when violations are found (line 61), making it suitable for automated quality gates.

Typical text output groups violations by issue type:


HTML_IN_TABLE (2 issues):
  docs/guide.md:45: [html-in-table] HTML comment interrupts table structure
  docs/guide.md:78: [html-in-table] HTML comment interrupts table structure

UNCLOSED_CODE_BLOCK (1 issue):
  docs/tutorial.md:102: [unclosed-code-block] Code block opened at line 102 is never closed

Total: 3 issues found

Summary

  • validate-markdown.py implements four targeted checks to catch structural Markdown errors before they reach production.
  • The script detects HTML comments inside tables, missing separator rows, column count mismatches, and unclosed code fences through dedicated detector functions.
  • A two-pass processing architecture first validates code block integrity, then scans table structures while skipping fenced content.
  • The CLI supports directory recursion, JSON output, and configurable exit codes for integration with CI/CD pipelines and pre-commit hooks.
  • Located at scripts/validate-markdown.py, this utility serves as the primary quality gate for Markdown documentation in the claude-skills repository.

Frequently Asked Questions

What types of Markdown errors does validate-markdown.py detect?

The script specifically targets four categories of structural errors that cause rendering failures: HTML comments embedded within table structures, missing separator rows between table headers and data, column count inconsistencies across table rows, and unclosed fenced code blocks. Each check is implemented as a discrete function within scripts/validate-markdown.py to ensure precise error localization.

How does validate-markdown.py handle code blocks during validation?

The script employs a two-pass scanning strategy where the first pass exclusively tracks fenced code blocks using a boolean toggle. This ensures that content inside triple-backtick fences is excluded from table validation during the second pass. The is_table_row() and related detection functions skip any lines flagged as being inside code blocks, preventing false positives from code snippets that happen to contain pipe characters.

Can validate-markdown.py be integrated into CI/CD pipelines?

Yes, the script is designed for automation through its command-line interface and exit code behavior. When invoked with --format json, it outputs machine-parseable results suitable for CI ingestion. The script exits with status 0 when no issues are found and status 1 when violations are detected, allowing build systems to fail pipelines automatically when Markdown syntax errors are introduced. The --path argument enables targeting specific documentation directories within repository structures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →