Document Conversion Formats Supported by Pandoc in the p2r3/convert Project

The p2r3/convert project supports over 60 document conversion formats through its Pandoc handler, including Markdown, Word, EPUB, LaTeX Beamer, and JATS XML, while explicitly excluding PDF and RevealJS outputs.

The p2r3/convert repository implements a TypeScript-based conversion layer that wraps a WebAssembly-compiled Pandoc engine. Understanding which document conversion formats are available requires examining the runtime discovery mechanism in the Pandoc handler, which dynamically queries the underlying engine to populate its supported format registry.

Runtime Discovery of Supported Formats

According to the source code in src/handlers/pandoc.ts, the handler does not hardcode its format list. Instead, it discovers document conversion formats dynamically during initialization through a three-step process:

  1. Query Pandoc's native capabilities: The handler invokes query({ query: "input-formats" }) to retrieve readable formats and query({ query: "output-formats" }) to retrieve writable formats.
  2. Merge capabilities: The two returned arrays are combined into a single allFormats Set to ensure both reading and writing are supported for each identifier.
  3. Filter exclusions: Lines 74-78 explicitly filter out pdf and revealjs before constructing the final registry.

Complete Catalog of Document Conversion Formats

After filtering, the handler instantiates a FileFormat object for each supported identifier using the static formatNames and formatExtensions lookup tables defined at lines 8-87 of src/handlers/pandoc.ts. The following table lists the document conversion formats exposed by the project, mapped to their human-readable names and default extensions:

Pandoc Identifier Display Name Default Extension
ansi ANSI terminal —
asciidoc modern AsciiDoc adoc
asciidoc_legacy AsciiDoc for asciidoc-py adoc
asciidoctor AsciiDoctor (= modern AsciiDoc) adoc
bbcode BBCode —
beamer LaTeX Beamer slides tex
biblatex BibLaTeX bibliography bib
bibtex BibTeX bibliography bib
bits BITS XML, alias for jats —
chunkedhtml zip of linked HTML files zip
commonmark CommonMark Markdown md
commonmark_x CommonMark with extensions md
context ConTeXt tex
creole Creole 1.0 —
csljson CSL JSON bibliography json
csv CSV table —
djot Djot markup dj
docbook DocBook v4 xml
docbook5 DocBook v5 xml
docx Word docx
dokuwiki DokuWiki markup —
dzslides DZSlides HTML slides html
endnotexml EndNote XML bibliography —
epub EPUB v3 epub
epub2 EPUB v2 epub
epub3 EPUB v3 epub
fb2 FictionBook2 fb2
gfm GitHub-Flavored Markdown md
haddock Haddock markup —
html HTML html
html4 XHTML 1.0 Transitional html
html5 HTML html
icml InDesign ICML —
ipynb Jupyter notebook ipynb
jats JATS XML xml
zimwiki ZimWiki markup —

(The em-dash (—) indicates formats where formatExtensions does not define a default extension.)

Excluded Formats and Architecture Constraints

Despite being valid Pandoc outputs, pdf and revealjs are deliberately removed from the supported list at lines 74-78 of src/handlers/pandoc.ts. PDF generation requires external LaTeX engines or Chromium dependencies incompatible with the WebAssembly environment loaded from src/handlers/pandoc/pandoc.js. RevealJS output requires additional template assets and JavaScript dependencies that complicate the standalone conversion pipeline. All other formats returned by the runtime query() calls remain available.

Implementation Through FileFormat Objects

Each supported format is represented as a FileFormat object conforming to the interface in src/FormatHandler.ts. The handler constructs these objects using two static mappings:

  • formatNames: Maps technical identifiers like gfm or docbook5 to human-readable labels.
  • formatExtensions: Maps identifiers to their default file extensions, ensuring outputs receive correct suffixes.

This architecture decouples the presentation layer from Pandoc's internal naming conventions while maintaining type safety. The dynamic discovery mechanism ensures that any new formats added to the underlying Pandoc engine become available immediately without requiring updates to the static maps.

Programmatic Usage Example

To interact with these document conversion formats programmatically, initialize the handler and invoke the conversion method:

import { PandocHandler } from './src/handlers/pandoc';

const handler = new PandocHandler();
await handler.initialize(); // Queries input/output formats via query()

// Convert GitHub-Flavored Markdown to AsciiDoc
const result = await handler.convert({
  input: "README.md",
  inputFormat: "gfm",
  outputFormat: "asciidoc"
});

The convert method validates both the source and target formats against the internally cached lists derived from the runtime queries before executing the transformation through the WebAssembly interface.

Summary

  • The project exposes over 60 document conversion formats discovered dynamically from Pandoc's native capabilities via query({ query: "input-formats" }) and query({ query: "output-formats" }).
  • Format metadata is defined in static maps at lines 8-87 of src/handlers/pandoc.ts, providing human-readable names and default extensions.
  • PDF and RevealJS are intentionally filtered out at lines 74-78 to avoid dependency and asset complications.
  • Each format is encapsulated as a FileFormat object per the interface in src/FormatHandler.ts, enabling consistent handling across the application layer.
  • The WebAssembly-based Pandoc engine (src/handlers/pandoc/pandoc.js) executes conversions after format validation against the dynamically built registry.

Frequently Asked Questions

Why are PDF and RevealJS excluded from the supported formats?

The handler explicitly removes pdf and revealjs from the available document conversion formats because PDF generation requires external LaTeX distributions or browser engines unavailable in the WebAssembly runtime, while RevealJS depends on external template files and JavaScript assets that complicate the standalone conversion environment.

How does the project discover new formats added to Pandoc?

The handler calls query({ query: "input-formats" }) and query({ query: "output-formats" }) against the compiled Pandoc module during initialization, building the supported list dynamically rather than using hardcoded arrays. This ensures new formats added to Pandoc become available automatically without modifying the TypeScript source code.

What determines the file extension for a converted document?

The formatExtensions map in src/handlers/pandoc.ts assigns default extensions (e.g., docx for Word, adoc for AsciiDoc) when creating FileFormat objects. If a format lacks a mapping in this static table, the extension is omitted and must be specified manually by the caller during the conversion process.

Can I convert between any two supported formats bidirectionally?

While the allFormats Set merges both input and output capabilities, bidirectional conversion is not guaranteed for every pair. The handler validates that the source format exists in the input-formats list and the target exists in the output-formats list before executing the transformation, respecting Pandoc's actual read/write limitations for each specific format.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →