Document Conversion Formats Supported by Pandoc in the p2r3/convert Project
The p2r3/convert project supports over 60 document conversion formats through its Pandoc handler, including Markdown, Word, EPUB, LaTeX Beamer, and JATS XML, while explicitly excluding PDF and RevealJS outputs.
The p2r3/convert repository implements a TypeScript-based conversion layer that wraps a WebAssembly-compiled Pandoc engine. Understanding which document conversion formats are available requires examining the runtime discovery mechanism in the Pandoc handler, which dynamically queries the underlying engine to populate its supported format registry.
Runtime Discovery of Supported Formats
According to the source code in src/handlers/pandoc.ts, the handler does not hardcode its format list. Instead, it discovers document conversion formats dynamically during initialization through a three-step process:
- Query Pandoc's native capabilities: The handler invokes
query({ query: "input-formats" })to retrieve readable formats andquery({ query: "output-formats" })to retrieve writable formats. - Merge capabilities: The two returned arrays are combined into a single
allFormatsSet to ensure both reading and writing are supported for each identifier. - Filter exclusions: Lines 74-78 explicitly filter out
pdfandrevealjsbefore constructing the final registry.
Complete Catalog of Document Conversion Formats
After filtering, the handler instantiates a FileFormat object for each supported identifier using the static formatNames and formatExtensions lookup tables defined at lines 8-87 of src/handlers/pandoc.ts. The following table lists the document conversion formats exposed by the project, mapped to their human-readable names and default extensions:
| Pandoc Identifier | Display Name | Default Extension |
|---|---|---|
ansi |
ANSI terminal | — |
asciidoc |
modern AsciiDoc | adoc |
asciidoc_legacy |
AsciiDoc for asciidoc-py | adoc |
asciidoctor |
AsciiDoctor (= modern AsciiDoc) | adoc |
bbcode |
BBCode | — |
beamer |
LaTeX Beamer slides | tex |
biblatex |
BibLaTeX bibliography | bib |
bibtex |
BibTeX bibliography | bib |
bits |
BITS XML, alias for jats | — |
chunkedhtml |
zip of linked HTML files | zip |
commonmark |
CommonMark Markdown | md |
commonmark_x |
CommonMark with extensions | md |
context |
ConTeXt | tex |
creole |
Creole 1.0 | — |
csljson |
CSL JSON bibliography | json |
csv |
CSV table | — |
djot |
Djot markup | dj |
docbook |
DocBook v4 | xml |
docbook5 |
DocBook v5 | xml |
docx |
Word | docx |
dokuwiki |
DokuWiki markup | — |
dzslides |
DZSlides HTML slides | html |
endnotexml |
EndNote XML bibliography | — |
epub |
EPUB v3 | epub |
epub2 |
EPUB v2 | epub |
epub3 |
EPUB v3 | epub |
fb2 |
FictionBook2 | fb2 |
gfm |
GitHub-Flavored Markdown | md |
haddock |
Haddock markup | — |
html |
HTML | html |
html4 |
XHTML 1.0 Transitional | html |
html5 |
HTML | html |
icml |
InDesign ICML | — |
ipynb |
Jupyter notebook | ipynb |
jats |
JATS XML | xml |
zimwiki |
ZimWiki markup | — |
(The em-dash (—) indicates formats where formatExtensions does not define a default extension.)
Excluded Formats and Architecture Constraints
Despite being valid Pandoc outputs, pdf and revealjs are deliberately removed from the supported list at lines 74-78 of src/handlers/pandoc.ts. PDF generation requires external LaTeX engines or Chromium dependencies incompatible with the WebAssembly environment loaded from src/handlers/pandoc/pandoc.js. RevealJS output requires additional template assets and JavaScript dependencies that complicate the standalone conversion pipeline. All other formats returned by the runtime query() calls remain available.
Implementation Through FileFormat Objects
Each supported format is represented as a FileFormat object conforming to the interface in src/FormatHandler.ts. The handler constructs these objects using two static mappings:
formatNames: Maps technical identifiers likegfmordocbook5to human-readable labels.formatExtensions: Maps identifiers to their default file extensions, ensuring outputs receive correct suffixes.
This architecture decouples the presentation layer from Pandoc's internal naming conventions while maintaining type safety. The dynamic discovery mechanism ensures that any new formats added to the underlying Pandoc engine become available immediately without requiring updates to the static maps.
Programmatic Usage Example
To interact with these document conversion formats programmatically, initialize the handler and invoke the conversion method:
import { PandocHandler } from './src/handlers/pandoc';
const handler = new PandocHandler();
await handler.initialize(); // Queries input/output formats via query()
// Convert GitHub-Flavored Markdown to AsciiDoc
const result = await handler.convert({
input: "README.md",
inputFormat: "gfm",
outputFormat: "asciidoc"
});
The convert method validates both the source and target formats against the internally cached lists derived from the runtime queries before executing the transformation through the WebAssembly interface.
Summary
- The project exposes over 60 document conversion formats discovered dynamically from Pandoc's native capabilities via
query({ query: "input-formats" })andquery({ query: "output-formats" }). - Format metadata is defined in static maps at lines 8-87 of
src/handlers/pandoc.ts, providing human-readable names and default extensions. - PDF and RevealJS are intentionally filtered out at lines 74-78 to avoid dependency and asset complications.
- Each format is encapsulated as a
FileFormatobject per the interface insrc/FormatHandler.ts, enabling consistent handling across the application layer. - The WebAssembly-based Pandoc engine (
src/handlers/pandoc/pandoc.js) executes conversions after format validation against the dynamically built registry.
Frequently Asked Questions
Why are PDF and RevealJS excluded from the supported formats?
The handler explicitly removes pdf and revealjs from the available document conversion formats because PDF generation requires external LaTeX distributions or browser engines unavailable in the WebAssembly runtime, while RevealJS depends on external template files and JavaScript assets that complicate the standalone conversion environment.
How does the project discover new formats added to Pandoc?
The handler calls query({ query: "input-formats" }) and query({ query: "output-formats" }) against the compiled Pandoc module during initialization, building the supported list dynamically rather than using hardcoded arrays. This ensures new formats added to Pandoc become available automatically without modifying the TypeScript source code.
What determines the file extension for a converted document?
The formatExtensions map in src/handlers/pandoc.ts assigns default extensions (e.g., docx for Word, adoc for AsciiDoc) when creating FileFormat objects. If a format lacks a mapping in this static table, the extension is omitted and must be specified manually by the caller during the conversion process.
Can I convert between any two supported formats bidirectionally?
While the allFormats Set merges both input and output capabilities, bidirectional conversion is not guaranteed for every pair. The handler validates that the source format exists in the input-formats list and the target exists in the output-formats list before executing the transformation, respecting Pandoc's actual read/write limitations for each specific format.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →