WeChat Article Exporter Supported Formats: HTML, JSON, Excel, PDF, and More
The wechat-article-exporter supports seven distinct export formats: Excel (.xlsx), JSON, HTML, plain text (.txt), Markdown (.md), Microsoft Word (.docx), and PDF.
The wechat-article-exporter is a TypeScript-based archiving tool for WeChat public articles. According to the source code in utils/download/Exporter.ts, the library defines a strict ExportType union type that enumerates every available output format, routing each request to a specialized private method that handles format-specific conversion and file generation.
Complete List of Supported Export Formats
The ExportType union defined in the exporter core explicitly lists seven supported string literals:
type ExportType = 'excel' | 'json' | 'html' | 'txt' | 'markdown' | 'word' | 'pdf';
- Excel – Generates
.xlsxworkbooks containing article metadata and content - JSON – Exports structured data as machine-readable
.jsonfiles - HTML – Creates self-contained folders with
index.htmland bundled assets - TXT – Produces plain-text versions of article content
- Markdown – Converts rendered HTML to
.mdformat using the turndown library - Word – Generates
.docxfiles using thehtmlDocxbrowser API - PDF – Creates PDF documents via a server-side Puppeteer rendering endpoint
How Each Export Format Works
When you invoke Exporter.startExport(type), the code routes your request to a format-specific private method. Here is the technical implementation behind each format.
Excel (.xlsx)
The exportExcelFiles() method delegates to export2ExcelFile from utils/exporter.ts. This utility generates a standard Excel workbook containing article data, handling cell formatting and worksheet creation programmatically.
JSON
The exportJsonFiles() method calls export2JsonFile from utils/exporter.ts to serialize article objects and write them as formatted .json files to disk.
HTML
The exportHtmlFiles() method creates a dedicated folder per article, writes an index.html file, and bundles associated assets. This produces a fully browsable offline copy with rewritten resource links.
Plain Text (.txt)
The exportTxtFiles() method extracts raw text content from the article, stripping HTML tags and formatting, then writes the result to a .txt file.
Markdown (.md)
The exportMarkdownFiles() method first renders the article HTML, then passes it through turndown to convert semantic markup to Markdown syntax before writing the .md file.
Microsoft Word (.docx)
The exportWordFiles() method utilizes window.htmlDocx.asBlob to convert the rendered HTML into a Word-compatible binary blob, producing a .docx file that preserves formatting.
The exportPdfFiles() method POSTs the final HTML to the server-side endpoint /api/web/pdf/generate, which uses Puppeteer to render the content and returns a PDF blob. The client then writes this blob to disk as a .pdf file.
Core Implementation Files
The export functionality spans several key files in the repository:
utils/download/Exporter.ts– The core class that orchestrates all format-specific pipelines and defines theExportTypeunionutils/exporter.ts– Helper utilities that handle the actual binary writing for Excel and JSON formatsutils/download/BaseDownloader.ts– Base class providing concurrency, retry logic, and proxy handling used by the Exporterserver/api/web/pdf/generate.ts– Server-side API endpoint that converts HTML to PDF using Puppeteer
Usage Examples
To export articles programmatically, instantiate the Exporter class with an array of article URLs, then call startExport() with your desired format:
import { Exporter } from '~/utils/download/Exporter';
const articleUrls = [
'https://mp.weixin.qq.com/s?__biz=...',
'https://mp.weixin.qq.com/s?__biz=...'
];
const exporter = new Exporter(articleUrls);
// Export to Excel
await exporter.startExport('excel');
// Export to JSON
await exporter.startExport('json');
// Export to HTML (creates folder with assets)
await exporter.startExport('html');
// Export to Markdown
await exporter.startExport('markdown');
// Export to PDF (requires server endpoint)
await exporter.startExport('pdf');
Each call executes the full pipeline: fetching cached HTML, extracting remote resources, downloading assets concurrently, rewriting resource links, and writing the final output file(s).
Summary
- The wechat-article-exporter supports seven export formats: Excel, JSON, HTML, TXT, Markdown, Word, and PDF
- Format routing is handled by the
Exporterclass inutils/download/Exporter.tsvia theExportTypeunion type - Excel and JSON use helper utilities in
utils/exporter.ts - Markdown conversion relies on the turndown library
- Word generation uses the browser-based
htmlDocxAPI - PDF export requires a server-side Puppeteer endpoint at
/api/web/pdf/generate - The architecture supports concurrent downloads and automatic asset bundling for HTML exports
Frequently Asked Questions
What export formats does wechat-article-exporter support?
The tool supports seven formats: Excel (.xlsx), JSON, HTML (with bundled assets), plain text (.txt), Markdown (.md), Microsoft Word (.docx), and PDF. These are defined as string literals in the ExportType union in utils/download/Exporter.ts.
How does PDF generation work in wechat-article-exporter?
PDF export uses a hybrid client-server approach. The exportPdfFiles() method POSTs the rendered HTML to the server-side endpoint /api/web/pdf/generate, which runs Puppeteer to generate the PDF. The server returns the PDF blob, which the client then saves to disk.
Which class handles the export format routing?
The Exporter class in utils/download/Exporter.ts handles all format routing. When you call startExport(type), it dispatches to private methods like exportExcelFiles(), exportMarkdownFiles(), or exportPdfFiles() based on the type parameter you provide.
Can I export WeChat articles to Markdown format?
Yes. The exportMarkdownFiles() method converts article HTML to Markdown using the turndown library. This preserves document structure including headers, lists, and links while producing clean Markdown syntax suitable for version control or static site generators.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →