# How the HTML Exporter Ensures 100% Style Preservation with Embedded Resources

> Learn how our HTML exporter achieves 100% style preservation. It embeds all resources locally by caching external assets and rewriting URLs, ensuring your original markup remains perfectly intact.

- Repository: [公众号文章工具箱/wechat-article-exporter](https://github.com/wechat-article/wechat-article-exporter)
- Tags: how-to-guide
- Published: 2026-05-26

---

**The HTML exporter achieves 100% style preservation by discovering every external asset (images, CSS, background images), downloading them into a local IndexedDB cache, and surgically rewriting all URLs in the HTML to point to these embedded resources while keeping the original markup intact.**

The `wechat-article-exporter` project reconstructs WeChat articles as self-contained, offline-capable HTML files. By implementing a three-phase pipeline that preserves every visual detail from fonts to background images, the HTML exporter ensures 100% style preservation with embedded resources.

## How the Three-Phase Export Pipeline Works

The export process follows a strict sequence: **resource discovery**, **cached downloading**, and **HTML normalisation**. Because the exporter never removes or mutates the original markup—only replacing remote URLs with local paths—the rendered page looks identical to the source, including fonts, colours, layout, and background images.

### Phase 1: Resource Discovery with extractResources()

The pipeline begins in [`utils/download/Exporter.ts`](https://github.com/wechat-article/wechat-article-exporter/blob/main/utils/download/Exporter.ts) where the `extractResources()` method parses each article’s HTML using the native `DOMParser` API. It scans for three categories of external dependencies:

- **Image elements** – checks both `src` and `data-src` attributes
- **Stylesheet links** – captures every `<link rel="stylesheet">` tag
- **CSS background images** – uses a regular expression to extract URLs from inline style attributes containing `background` or `background-image` declarations

```typescript
// From utils/download/Exporter.ts
private async extractResources(): Promise<void> {
  const parser = new DOMParser();

  for (const url of this.urls) {
    const article = await getArticleByLink(url);
    const html = await cached.file.text();
    const document = parser.parseFromString(html, 'text/html');

    // ① Images
    const imgs = document.querySelectorAll<HTMLImageElement>('img');
    for (const img of imgs) {
      const imgUrl = img.getAttribute('src') || img.getAttribute('data-src');
      if (imgUrl) {
        resources.push(imgUrl);
        this.resources.add({ url: imgUrl, fakeid: article.fakeid });
      }
    }

    // ② Stylesheets
    const links = document.querySelectorAll<HTMLLinkElement>('link[rel="stylesheet"]');
    for (const link of links) {
      const url = link.href;
      if (url) {
        resources.push(url);
        this.resources.add({ url, fakeid: article.fakeid });
      }
    }

    // ③ Background images (CSS “url()”)
    html.replaceAll(
      /((?:background|background-image): url\((?:&quot;)?)((?:https?|\/\/)[^)]+?)((?:&quot;)?\))/gs,
      (_, p1, url, p3) => {
        resources.push(url);
        this.resources.add({ url, fakeid: article.fakeid });
        return `${p1}${url}${p3}`;
      }
    );
  }
}

```

*Lines 101–115 initialise the parser, while lines 124–131 handle images, 134–141 capture stylesheets, and 144–152 extract background-image URLs via regex.*

### Phase 2: Resource Download and Caching

Once discovered, each unique resource enters the `downloadResourceTask()` method (lines 332–364). The system employs a **cached-first strategy**: it checks `getResourceCache()` to see if the asset already exists in IndexedDB before initiating any network requests.

```typescript
// From utils/download/Exporter.ts
private async downloadResourceTask(url: string, fakeid: string): Promise<void> {
  this.pending.add(url);
  const cached = await getResourceCache(url);
  if (cached) { return; }

  for (let attempt = 0; attempt < this.options.maxRetries; attempt++) {
    const proxy = this.proxyManager.getBestProxy();
    try {
      const blob = await this.download(fakeid, url, proxy);
      await updateResourceCache({ fakeid, url, file: blob });
      this.pending.delete(url);
      this.completed.add(url);
      this.proxyManager.recordSuccess(proxy);
      return;
    } catch (error) {
      await this.handleDownloadFailure(proxy, url, attempt, error);
    }
  }
}

```

This method uses the `proxyManager` to route requests, stores the resulting `Blob` in IndexedDB via `updateResourceCache()`, and tracks completion status to prevent duplicate downloads.

### Phase 3: HTML Normalisation and Local Asset Mapping

The final phase occurs in `exportHtmlFiles()` (lines 380–424) and `normalizeHtml()`. For each article, the exporter:

1. Creates a unique `urlmap` (`Map<string, string>`) that associates remote URLs with local `./assets/` paths
2. Writes each cached resource to the filesystem using `mime.getExtension()` to preserve original file types
3. Rewrites all references in the HTML to point to these local assets

```typescript
// From utils/download/Exporter.ts
private async normalizeHtml(cachedHtml: HtmlAsset, html: string, urlmap: Map<string, string>) {
  // Replace background‑image URLs
  html = html.replaceAll(
    /((?:background|background-image): url\((?:&quot;)?)((?:https?|\/\/)[^)]+?)((?:&quot;)?\))/gs,
    (_, p1, url, p3) => urlmap.has(url) ? `${p1}${urlmap.get(url)}${p3}` : `${p1}${url}${p3}`
  );

  const parser = new DOMParser();
  const document = parser.parseFromString(html, 'text/html');

  // Replace <img> src attributes
  const imgs = document.querySelectorAll<HTMLImageElement>('img');
  for (const img of imgs) {
    const imgUrl = img.getAttribute('src') || img.getAttribute('data-src');
    if (imgUrl && urlmap.has(imgUrl)) img.src = urlmap.get(imgUrl)!;
  }

  // Replace <link rel="stylesheet"> href attributes
  const links = document.querySelectorAll<HTMLLinkElement>('link[rel="stylesheet"]');
  let localLinks = '';
  for (const link of links) {
    const linkUrl = link.href;
    if (linkUrl && urlmap.has(linkUrl)) {
      localLinks += `<link rel="stylesheet" href="${urlmap.get(linkUrl)}">\n`;
    }
  }

  return `<!DOCTYPE html>
<html lang="zh_CN">
<head>
  ${localLinks}
</head>
<body>
  ${pageContentHTML}
</body>
</html>`;
}

```

*Lines 665–700 handle background-image replacement, lines 664–672 rewrite image sources, and lines 778–786 reconstruct the stylesheet links.*

## Complete Code Example: Exporting Articles to Self-Contained HTML

Below is a runnable example showing how to initiate an export using the `Exporter` class from a Nuxt SPA environment:

```typescript
import { Exporter } from '~/utils/download/Exporter';

// URLs of the articles you want to export
const urls = [
  'https://mp.weixin.qq.com/s?__biz=MzU0...'
];

// Create an exporter instance
const exporter = new Exporter(urls);

// Start the HTML export – this will:
//   1️⃣ discover resources,
//   2️⃣ download them,
//   3️⃣ write assets & a self‑contained index.html.
exporter.startExport('html')
  .then(() => console.log('✅ HTML export finished'))
  .catch(err => console.error('❌ Export failed', err));

```

Executing this code opens the native directory picker, creates a folder per article, stores all assets under `assets/`, and writes a fully-styled [`index.html`](https://github.com/wechat-article/wechat-article-exporter/blob/main/index.html) that can be viewed offline.

## Key Files and Their Roles

- **[`utils/download/Exporter.ts`](https://github.com/wechat-article/wechat-article-exporter/blob/main/utils/download/Exporter.ts)** – Core exporter class that orchestrates resource discovery, downloading, and HTML normalisation
- **[`utils/download/BaseDownloader.ts`](https://github.com/wechat-article/wechat-article-exporter/blob/main/utils/download/BaseDownloader.ts)** – Shared download-queue logic used by `Exporter` for concurrent processing
- **[`store/v2/resource.ts`](https://github.com/wechat-article/wechat-article-exporter/blob/main/store/v2/resource.ts)** – IndexedDB wrapper for storing downloaded binary blobs
- **[`store/v2/resource-map.ts`](https://github.com/wechat-article/wechat-article-exporter/blob/main/store/v2/resource-map.ts)** – Cache that records which URLs belong to each specific article
- **[`shared/utils/helpers.ts`](https://github.com/wechat-article/wechat-article-exporter/blob/main/shared/utils/helpers.ts)** – Utility functions including `filterInvalidFilenameChars` for safe filesystem operations
- **[`shared/utils/renderer.ts`](https://github.com/wechat-article/wechat-article-exporter/blob/main/shared/utils/renderer.ts)** – Functions that transform raw WeChat CGI data into clean HTML before export

## Summary

- **Comprehensive asset discovery** – The `extractResources()` method inspects HTML elements, CSS links, and inline styles to find every external dependency.
- **Intelligent caching** – Resources are stored in IndexedDB via `updateResourceCache()` and only downloaded once, even across multiple export sessions.
- **Surgical URL rewriting** – The `normalizeHtml()` method creates a `urlmap` to replace remote URLs with local `./assets/` paths without altering the DOM structure.
- **100% visual fidelity** – Because the original markup remains untouched (only references change), the exported HTML preserves all original fonts, colours, layouts, and background images.

## Frequently Asked Questions

### How does the exporter handle CSS background images defined in inline styles?

The exporter uses a global regular expression `/((?:background|background-image): url\((?:&quot;)?)((?:https?|\/\/)[^)]+?)((?:&quot;)?\))/gs` in both the discovery and normalisation phases to extract and rewrite URLs embedded within `style` attributes. This ensures background images are downloaded and remapped just like standard `<img>` tags.

### Does the HTML exporter modify the original article markup?

No. According to the implementation in [`utils/download/Exporter.ts`](https://github.com/wechat-article/wechat-article-exporter/blob/main/utils/download/Exporter.ts), the exporter **never removes or mutates the original markup** except for the surgical replacement of remote URLs with local file paths. This approach guarantees that the rendered layout, including complex WeChat-specific styling, remains identical to the source.

### What happens if a resource download fails?

The `downloadResourceTask()` method implements a retry mechanism controlled by `this.options.maxRetries`. If a download fails, it attempts to fetch the resource again using different proxies managed by `proxyManager`. Failed attempts are logged, and the system continues processing other assets to maximise the completeness of the export.

### Can the exporter work offline after resources are cached?

Yes. Because the system checks `getResourceCache()` before initiating any network request, previously downloaded assets persist in IndexedDB. If you export the same article again or re-export during a session, the exporter uses the cached blobs immediately, requiring no network connectivity for those specific resources.