How the HTML Exporter Ensures 100% Style Preservation with Embedded Resources
The HTML exporter achieves 100% style preservation by discovering every external asset (images, CSS, background images), downloading them into a local IndexedDB cache, and surgically rewriting all URLs in the HTML to point to these embedded resources while keeping the original markup intact.
The wechat-article-exporter project reconstructs WeChat articles as self-contained, offline-capable HTML files. By implementing a three-phase pipeline that preserves every visual detail from fonts to background images, the HTML exporter ensures 100% style preservation with embedded resources.
How the Three-Phase Export Pipeline Works
The export process follows a strict sequence: resource discovery, cached downloading, and HTML normalisation. Because the exporter never removes or mutates the original markup—only replacing remote URLs with local paths—the rendered page looks identical to the source, including fonts, colours, layout, and background images.
Phase 1: Resource Discovery with extractResources()
The pipeline begins in utils/download/Exporter.ts where the extractResources() method parses each article’s HTML using the native DOMParser API. It scans for three categories of external dependencies:
- Image elements – checks both
srcanddata-srcattributes - Stylesheet links – captures every
<link rel="stylesheet">tag - CSS background images – uses a regular expression to extract URLs from inline style attributes containing
backgroundorbackground-imagedeclarations
// From utils/download/Exporter.ts
private async extractResources(): Promise<void> {
const parser = new DOMParser();
for (const url of this.urls) {
const article = await getArticleByLink(url);
const html = await cached.file.text();
const document = parser.parseFromString(html, 'text/html');
// ① Images
const imgs = document.querySelectorAll<HTMLImageElement>('img');
for (const img of imgs) {
const imgUrl = img.getAttribute('src') || img.getAttribute('data-src');
if (imgUrl) {
resources.push(imgUrl);
this.resources.add({ url: imgUrl, fakeid: article.fakeid });
}
}
// ② Stylesheets
const links = document.querySelectorAll<HTMLLinkElement>('link[rel="stylesheet"]');
for (const link of links) {
const url = link.href;
if (url) {
resources.push(url);
this.resources.add({ url, fakeid: article.fakeid });
}
}
// ③ Background images (CSS “url()”)
html.replaceAll(
/((?:background|background-image): url\((?:")?)((?:https?|\/\/)[^)]+?)((?:")?\))/gs,
(_, p1, url, p3) => {
resources.push(url);
this.resources.add({ url, fakeid: article.fakeid });
return `${p1}${url}${p3}`;
}
);
}
}
Lines 101–115 initialise the parser, while lines 124–131 handle images, 134–141 capture stylesheets, and 144–152 extract background-image URLs via regex.
Phase 2: Resource Download and Caching
Once discovered, each unique resource enters the downloadResourceTask() method (lines 332–364). The system employs a cached-first strategy: it checks getResourceCache() to see if the asset already exists in IndexedDB before initiating any network requests.
// From utils/download/Exporter.ts
private async downloadResourceTask(url: string, fakeid: string): Promise<void> {
this.pending.add(url);
const cached = await getResourceCache(url);
if (cached) { return; }
for (let attempt = 0; attempt < this.options.maxRetries; attempt++) {
const proxy = this.proxyManager.getBestProxy();
try {
const blob = await this.download(fakeid, url, proxy);
await updateResourceCache({ fakeid, url, file: blob });
this.pending.delete(url);
this.completed.add(url);
this.proxyManager.recordSuccess(proxy);
return;
} catch (error) {
await this.handleDownloadFailure(proxy, url, attempt, error);
}
}
}
This method uses the proxyManager to route requests, stores the resulting Blob in IndexedDB via updateResourceCache(), and tracks completion status to prevent duplicate downloads.
Phase 3: HTML Normalisation and Local Asset Mapping
The final phase occurs in exportHtmlFiles() (lines 380–424) and normalizeHtml(). For each article, the exporter:
- Creates a unique
urlmap(Map<string, string>) that associates remote URLs with local./assets/paths - Writes each cached resource to the filesystem using
mime.getExtension()to preserve original file types - Rewrites all references in the HTML to point to these local assets
// From utils/download/Exporter.ts
private async normalizeHtml(cachedHtml: HtmlAsset, html: string, urlmap: Map<string, string>) {
// Replace background‑image URLs
html = html.replaceAll(
/((?:background|background-image): url\((?:")?)((?:https?|\/\/)[^)]+?)((?:")?\))/gs,
(_, p1, url, p3) => urlmap.has(url) ? `${p1}${urlmap.get(url)}${p3}` : `${p1}${url}${p3}`
);
const parser = new DOMParser();
const document = parser.parseFromString(html, 'text/html');
// Replace <img> src attributes
const imgs = document.querySelectorAll<HTMLImageElement>('img');
for (const img of imgs) {
const imgUrl = img.getAttribute('src') || img.getAttribute('data-src');
if (imgUrl && urlmap.has(imgUrl)) img.src = urlmap.get(imgUrl)!;
}
// Replace <link rel="stylesheet"> href attributes
const links = document.querySelectorAll<HTMLLinkElement>('link[rel="stylesheet"]');
let localLinks = '';
for (const link of links) {
const linkUrl = link.href;
if (linkUrl && urlmap.has(linkUrl)) {
localLinks += `<link rel="stylesheet" href="${urlmap.get(linkUrl)}">\n`;
}
}
return `<!DOCTYPE html>
<html lang="zh_CN">
<head>
${localLinks}
</head>
<body>
${pageContentHTML}
</body>
</html>`;
}
Lines 665–700 handle background-image replacement, lines 664–672 rewrite image sources, and lines 778–786 reconstruct the stylesheet links.
Complete Code Example: Exporting Articles to Self-Contained HTML
Below is a runnable example showing how to initiate an export using the Exporter class from a Nuxt SPA environment:
import { Exporter } from '~/utils/download/Exporter';
// URLs of the articles you want to export
const urls = [
'https://mp.weixin.qq.com/s?__biz=MzU0...'
];
// Create an exporter instance
const exporter = new Exporter(urls);
// Start the HTML export – this will:
// 1️⃣ discover resources,
// 2️⃣ download them,
// 3️⃣ write assets & a self‑contained index.html.
exporter.startExport('html')
.then(() => console.log('✅ HTML export finished'))
.catch(err => console.error('❌ Export failed', err));
Executing this code opens the native directory picker, creates a folder per article, stores all assets under assets/, and writes a fully-styled index.html that can be viewed offline.
Key Files and Their Roles
utils/download/Exporter.ts– Core exporter class that orchestrates resource discovery, downloading, and HTML normalisationutils/download/BaseDownloader.ts– Shared download-queue logic used byExporterfor concurrent processingstore/v2/resource.ts– IndexedDB wrapper for storing downloaded binary blobsstore/v2/resource-map.ts– Cache that records which URLs belong to each specific articleshared/utils/helpers.ts– Utility functions includingfilterInvalidFilenameCharsfor safe filesystem operationsshared/utils/renderer.ts– Functions that transform raw WeChat CGI data into clean HTML before export
Summary
- Comprehensive asset discovery – The
extractResources()method inspects HTML elements, CSS links, and inline styles to find every external dependency. - Intelligent caching – Resources are stored in IndexedDB via
updateResourceCache()and only downloaded once, even across multiple export sessions. - Surgical URL rewriting – The
normalizeHtml()method creates aurlmapto replace remote URLs with local./assets/paths without altering the DOM structure. - 100% visual fidelity – Because the original markup remains untouched (only references change), the exported HTML preserves all original fonts, colours, layouts, and background images.
Frequently Asked Questions
How does the exporter handle CSS background images defined in inline styles?
The exporter uses a global regular expression /((?:background|background-image): url\((?:")?)((?:https?|\/\/)[^)]+?)((?:")?\))/gs in both the discovery and normalisation phases to extract and rewrite URLs embedded within style attributes. This ensures background images are downloaded and remapped just like standard <img> tags.
Does the HTML exporter modify the original article markup?
No. According to the implementation in utils/download/Exporter.ts, the exporter never removes or mutates the original markup except for the surgical replacement of remote URLs with local file paths. This approach guarantees that the rendered layout, including complex WeChat-specific styling, remains identical to the source.
What happens if a resource download fails?
The downloadResourceTask() method implements a retry mechanism controlled by this.options.maxRetries. If a download fails, it attempts to fetch the resource again using different proxies managed by proxyManager. Failed attempts are logged, and the system continues processing other assets to maximise the completeness of the export.
Can the exporter work offline after resources are cached?
Yes. Because the system checks getResourceCache() before initiating any network request, previously downloaded assets persist in IndexedDB. If you export the same article again or re-export during a session, the exporter uses the cached blobs immediately, requiring no network connectivity for those specific resources.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →