What Markdown Syntax Issues Does validate-markdown.py Check For?

The validate-markdown.py script in the Jeffallan/claude-skills repository detects four specific structural problems: HTML comments inside tables, missing table separator rows, column-count mismatches in tables, and unclosed fenced code blocks.

Maintaining clean markdown syntax is critical for documentation that renders correctly across different parsers and platforms. The validate-markdown.py utility, located at scripts/validate-markdown.py, provides a automated quality gate that catches common markdown syntax issues before they reach production. This self-contained Python script implements a two-pass validation pipeline specifically designed to identify structural defects that standard linters often miss.

The Four Critical Markdown Syntax Checks

The validator implements four distinct detection routines, each targeting a specific class of rendering failure.

HTML Comments Inside Tables

HTML comments (<!-- … -->) placed within table structures cause immediate rendering failures in most markdown parsers. The script detects this via the is_html_comment() function (lines 59-62), which scans table content for comment delimiters. When found between a table header and its separator, or inside data rows, the script generates an HTML_IN_TABLE issue (lines 121-131). This check ensures that documentation tables render consistently without interruption from hidden comments.

Missing Table Separator Rows

Standard markdown tables require a separator line (|---|---|) immediately following the header row. The validator enforces this structural requirement using is_table_row() (lines 47-50) to identify potential headers, then immediately checks the subsequent line with is_separator_row() (lines 53-56). If the separator is absent, the script records a MISSING_SEPARATOR issue (lines 134-144). This prevents tables from being parsed as plain text or malformed headers.

Column-Count Mismatches

Table integrity requires consistent column counts across all rows. The script implements count_columns() (lines 41-45) to count pipe characters while intelligently ignoring escaped pipes (\|). During the table validation pass, each data row's column count is compared against the header_cols value established from the header row. Discrepancies trigger a COLUMN_MISMATCH issue (lines 167-178), ensuring tables align correctly when rendered.

Unclosed Fenced Code Blocks

Fenced code blocks initiated with triple backticks (```) must be properly closed to prevent the remainder of the document from being interpreted as code. The script performs a dedicated first pass through the file (lines 71-80), toggling an in_code_block boolean flag whenever it encounters fence markers. If the flag remains true after processing all lines, the script creates an UNCLOSED_CODE_BLOCK issue pointing to the opening fence location (lines 81-89).

How the Validation Pipeline Works

The validate-markdown.py script processes files through a structured two-pass pipeline implemented in validate_file() (lines 68-70).

First Pass: Code Block Sanity The script walks the entire file once to detect unclosed fenced code blocks (lines 71-89). This must occur before table validation because code blocks often contain pipe characters that should not be interpreted as table markup.

Second Pass: Table Structure Validation The script walks the lines again with a manual index to maintain context (lines 95-182). During this pass:

  • Content inside fenced code blocks is skipped (lines 99-103)
  • When is_table_row() identifies a potential table header, the script validates the separator row and subsequent data rows for column consistency
  • HTML comments are detected via is_html_comment() during table traversal

Directory Recursion The validate_directory() function (lines 87-95) uses Path.rglob("*.md") to apply these checks recursively across entire documentation trees.

Running the Validator

The script provides a command-line interface through main() (lines 98-63) that supports both single-file and directory validation with multiple output formats.

Validate the default "skills" folder with human-readable output:

python scripts/validate-markdown.py --check

Validate a specific documentation directory and output JSON for CI pipelines:

python scripts/validate-markdown.py --path docs --format json

Validate a single file for pre-commit hooks:

python scripts/validate-markdown.py --check --path README.md

The script exits with 0 on success or 1 when markdown syntax issues are detected (line 61), making it suitable for automated quality gates.

Summary

  • HTML comments in tables break rendering and are detected by is_html_comment() at lines 59-62
  • Missing separator rows after table headers trigger MISSING_SEPARATOR issues via is_separator_row() at lines 53-56
  • Column-count mismatches are caught by comparing count_columns() results against the header column count at lines 167-178
  • Unclosed code blocks are identified in the first validation pass by tracking fence markers at lines 71-89

Frequently Asked Questions

What happens if a markdown file contains nested code blocks with different fence lengths?

The validator tracks all fenced code blocks using the same in_code_block boolean toggle regardless of fence length (lines 71-80). As implemented in scripts/validate-markdown.py, the script treats any line containing triple backticks as a toggle, so nested blocks with different fence lengths may cause false positives in table validation but will still correctly detect unclosed outer blocks.

Can the validator detect malformed table separators that use incorrect characters?

The is_separator_row() function (lines 53-56) specifically checks for separator rows containing only pipes, hyphens, and colons (for alignment). If a separator contains other characters or incorrect formatting, the script treats it as a missing separator and generates a MISSING_SEPARATOR issue (lines 134-144) rather than a specific "malformed separator" error.

How does the script handle escaped pipe characters inside table cells?

The count_columns() function (lines 41-45) implements logic to ignore escaped pipes (\|) when counting column delimiters. This ensures that table cells containing literal pipe characters do not trigger false COLUMN_MISMATCH issues, maintaining accurate validation even when table content includes technical syntax like regular expressions or command-line examples containing pipes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →