How the HTML Exporter Ensures 100% Style Preservation with Embedded Resources

The HTML exporter achieves 100% style preservation by discovering every external asset (images, CSS, background images), downloading them into a local IndexedDB cache, and surgically rewriting all URLs in the HTML to point to these embedded resources while keeping the original markup intact.

The wechat-article-exporter project reconstructs WeChat articles as self-contained, offline-capable HTML files. By implementing a three-phase pipeline that preserves every visual detail from fonts to background images, the HTML exporter ensures 100% style preservation with embedded resources.

How the Three-Phase Export Pipeline Works

The export process follows a strict sequence: resource discovery, cached downloading, and HTML normalisation. Because the exporter never removes or mutates the original markup—only replacing remote URLs with local paths—the rendered page looks identical to the source, including fonts, colours, layout, and background images.

Phase 1: Resource Discovery with extractResources()

The pipeline begins in utils/download/Exporter.ts where the extractResources() method parses each article’s HTML using the native DOMParser API. It scans for three categories of external dependencies:

  • Image elements – checks both src and data-src attributes
  • Stylesheet links – captures every <link rel="stylesheet"> tag
  • CSS background images – uses a regular expression to extract URLs from inline style attributes containing background or background-image declarations
// From utils/download/Exporter.ts
private async extractResources(): Promise<void> {
  const parser = new DOMParser();

  for (const url of this.urls) {
    const article = await getArticleByLink(url);
    const html = await cached.file.text();
    const document = parser.parseFromString(html, 'text/html');

    // ① Images
    const imgs = document.querySelectorAll<HTMLImageElement>('img');
    for (const img of imgs) {
      const imgUrl = img.getAttribute('src') || img.getAttribute('data-src');
      if (imgUrl) {
        resources.push(imgUrl);
        this.resources.add({ url: imgUrl, fakeid: article.fakeid });
      }
    }

    // ② Stylesheets
    const links = document.querySelectorAll<HTMLLinkElement>('link[rel="stylesheet"]');
    for (const link of links) {
      const url = link.href;
      if (url) {
        resources.push(url);
        this.resources.add({ url, fakeid: article.fakeid });
      }
    }

    // ③ Background images (CSS “url()”)
    html.replaceAll(
      /((?:background|background-image): url\((?:&quot;)?)((?:https?|\/\/)[^)]+?)((?:&quot;)?\))/gs,
      (_, p1, url, p3) => {
        resources.push(url);
        this.resources.add({ url, fakeid: article.fakeid });
        return `${p1}${url}${p3}`;
      }
    );
  }
}

Lines 101–115 initialise the parser, while lines 124–131 handle images, 134–141 capture stylesheets, and 144–152 extract background-image URLs via regex.

Phase 2: Resource Download and Caching

Once discovered, each unique resource enters the downloadResourceTask() method (lines 332–364). The system employs a cached-first strategy: it checks getResourceCache() to see if the asset already exists in IndexedDB before initiating any network requests.

// From utils/download/Exporter.ts
private async downloadResourceTask(url: string, fakeid: string): Promise<void> {
  this.pending.add(url);
  const cached = await getResourceCache(url);
  if (cached) { return; }

  for (let attempt = 0; attempt < this.options.maxRetries; attempt++) {
    const proxy = this.proxyManager.getBestProxy();
    try {
      const blob = await this.download(fakeid, url, proxy);
      await updateResourceCache({ fakeid, url, file: blob });
      this.pending.delete(url);
      this.completed.add(url);
      this.proxyManager.recordSuccess(proxy);
      return;
    } catch (error) {
      await this.handleDownloadFailure(proxy, url, attempt, error);
    }
  }
}

This method uses the proxyManager to route requests, stores the resulting Blob in IndexedDB via updateResourceCache(), and tracks completion status to prevent duplicate downloads.

Phase 3: HTML Normalisation and Local Asset Mapping

The final phase occurs in exportHtmlFiles() (lines 380–424) and normalizeHtml(). For each article, the exporter:

  1. Creates a unique urlmap (Map<string, string>) that associates remote URLs with local ./assets/ paths
  2. Writes each cached resource to the filesystem using mime.getExtension() to preserve original file types
  3. Rewrites all references in the HTML to point to these local assets
// From utils/download/Exporter.ts
private async normalizeHtml(cachedHtml: HtmlAsset, html: string, urlmap: Map<string, string>) {
  // Replace background‑image URLs
  html = html.replaceAll(
    /((?:background|background-image): url\((?:&quot;)?)((?:https?|\/\/)[^)]+?)((?:&quot;)?\))/gs,
    (_, p1, url, p3) => urlmap.has(url) ? `${p1}${urlmap.get(url)}${p3}` : `${p1}${url}${p3}`
  );

  const parser = new DOMParser();
  const document = parser.parseFromString(html, 'text/html');

  // Replace <img> src attributes
  const imgs = document.querySelectorAll<HTMLImageElement>('img');
  for (const img of imgs) {
    const imgUrl = img.getAttribute('src') || img.getAttribute('data-src');
    if (imgUrl && urlmap.has(imgUrl)) img.src = urlmap.get(imgUrl)!;
  }

  // Replace <link rel="stylesheet"> href attributes
  const links = document.querySelectorAll<HTMLLinkElement>('link[rel="stylesheet"]');
  let localLinks = '';
  for (const link of links) {
    const linkUrl = link.href;
    if (linkUrl && urlmap.has(linkUrl)) {
      localLinks += `<link rel="stylesheet" href="${urlmap.get(linkUrl)}">\n`;
    }
  }

  return `<!DOCTYPE html>
<html lang="zh_CN">
<head>
  ${localLinks}
</head>
<body>
  ${pageContentHTML}
</body>
</html>`;
}

Lines 665–700 handle background-image replacement, lines 664–672 rewrite image sources, and lines 778–786 reconstruct the stylesheet links.

Complete Code Example: Exporting Articles to Self-Contained HTML

Below is a runnable example showing how to initiate an export using the Exporter class from a Nuxt SPA environment:

import { Exporter } from '~/utils/download/Exporter';

// URLs of the articles you want to export
const urls = [
  'https://mp.weixin.qq.com/s?__biz=MzU0...'
];

// Create an exporter instance
const exporter = new Exporter(urls);

// Start the HTML export – this will:
//   1️⃣ discover resources,
//   2️⃣ download them,
//   3️⃣ write assets & a self‑contained index.html.
exporter.startExport('html')
  .then(() => console.log('✅ HTML export finished'))
  .catch(err => console.error('❌ Export failed', err));

Executing this code opens the native directory picker, creates a folder per article, stores all assets under assets/, and writes a fully-styled index.html that can be viewed offline.

Key Files and Their Roles

Summary

  • Comprehensive asset discovery – The extractResources() method inspects HTML elements, CSS links, and inline styles to find every external dependency.
  • Intelligent caching – Resources are stored in IndexedDB via updateResourceCache() and only downloaded once, even across multiple export sessions.
  • Surgical URL rewriting – The normalizeHtml() method creates a urlmap to replace remote URLs with local ./assets/ paths without altering the DOM structure.
  • 100% visual fidelity – Because the original markup remains untouched (only references change), the exported HTML preserves all original fonts, colours, layouts, and background images.

Frequently Asked Questions

How does the exporter handle CSS background images defined in inline styles?

The exporter uses a global regular expression /((?:background|background-image): url\((?:&quot;)?)((?:https?|\/\/)[^)]+?)((?:&quot;)?\))/gs in both the discovery and normalisation phases to extract and rewrite URLs embedded within style attributes. This ensures background images are downloaded and remapped just like standard <img> tags.

Does the HTML exporter modify the original article markup?

No. According to the implementation in utils/download/Exporter.ts, the exporter never removes or mutates the original markup except for the surgical replacement of remote URLs with local file paths. This approach guarantees that the rendered layout, including complex WeChat-specific styling, remains identical to the source.

What happens if a resource download fails?

The downloadResourceTask() method implements a retry mechanism controlled by this.options.maxRetries. If a download fails, it attempts to fetch the resource again using different proxies managed by proxyManager. Failed attempts are logged, and the system continues processing other assets to maximise the completeness of the export.

Can the exporter work offline after resources are cached?

Yes. Because the system checks getResourceCache() before initiating any network request, previously downloaded assets persist in IndexedDB. If you export the same article again or re-export during a session, the exporter uses the cached blobs immediately, requiring no network connectivity for those specific resources.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →