How Image and Video Resources Are Downloaded and Mapped During Export in WeChat Articles
The Exporter class drives a three-phase pipeline—discovery, concurrent cache-aware download, and URL remapping—that extracts every external asset from WeChat articles, stores them locally with UUID filenames, and rewrites HTML to reference the offline copies.
The wechat-article/wechat-article-exporter repository converts WeChat public account articles into self-contained HTML or ZIP archives. To ensure exported files display correctly without internet access, the system must identify all remote images, stylesheets, background images, and video assets, fetch them through proxy managers, and remap their URLs to relative local paths.
The Three-Step Export Pipeline
Located in utils/download/Exporter.ts, the Exporter class orchestrates asset handling through extractResources(), processExportQueue(), and exportHtmlFiles().
Step 1: Resource Discovery via DOM Parsing
The extractResources() method parses each article’s HTML to identify three categories of external dependencies:
- Image tags: Extracts
srcordata-srcattributes from<img>elements - Stylesheets: Captures
hrefvalues from<link rel="stylesheet">tags - CSS background images: Uses a global regex to capture URLs within inline
backgroundorbackground-imageproperty declarations
// utils/download/Exporter.ts → extractResources()
document.querySelectorAll<HTMLImageElement>('img')
.forEach(img => {
const imgUrl = img.getAttribute('src') || img.getAttribute('data-src');
if (imgUrl) {
this.resources.add({ url: imgUrl, fakeid: article.fakeid });
}
});
html.replaceAll(
/((?:background|background-image): url\((?:")?)((?:https?|\/\/)[^)]+?)((?:quot;)?\))/gs,
(_, p1, url, p3) => {
this.resources.add({ url, fakeid: article.fakeid });
return `${p1}${url}${p3}`;
});
Discovered URLs are stored in a Set to ensure uniqueness and simultaneously persisted to the resource-map table (store/v2/resource-map.ts), creating a per-article index of all external dependencies.
Step 2: Concurrent Download with Intelligent Caching
The processExportQueue() method implements a bounded concurrency pool (default limit of 5) to fetch assets without overwhelming network resources:
// utils/download/Exporter.ts → processExportQueue()
while (resources.length > 0 || activePromises.size > 0) {
while (activePromises.size < this.options.concurrency && resources.length > 0) {
const { url, fakeid } = resources.pop()!;
const promise = this.downloadResourceTask(url, fakeid);
activePromises.add(promise);
promise.finally(() => activePromises.delete(promise));
}
if (activePromises.size) await Promise.race(activePromises);
}
Each task invokes downloadResourceTask(), which first queries the resource cache (store/v2/resource.ts). If getResourceCache(url) returns a hit, the system reuses the stored Blob. On a cache miss, the method fetches the asset via the proxy manager, then persists the result using updateResourceCache():
// utils/download/Exporter.ts → downloadResourceTask()
const cached = await getResourceCache(url);
if (!cached) {
const blob = await this.download(fakeid, url, proxy);
await updateResourceCache({ fakeid, url, file: blob });
}
This caching layer ensures that duplicate assets across multiple articles or repeated exports are downloaded only once.
Step 3: Local File Generation and URL Remapping
When generating the final output, exportHtmlFiles() performs the conversion from remote URLs to local relative paths:
- Retrieves the article’s resource list from
getResourceMapCache() - Fetches each
Blobfrom the resource cache - Generates UUID-based filenames using
mime.getExtension()to determine the correct suffix - Writes files to
./assets/within the export directory - Builds a
Map<string,string>(urlmap) associating original remote URLs with local relative paths
// utils/download/Exporter.ts → exportHtmlFiles()
for (const resourceUrl of resourceMap.resources) {
const resource = await getResourceCache(resourceUrl);
if (!resource) continue;
const uuid = new Date().getTime() + Math.random().toString();
const ext = mime.getExtension(resource.file.type);
await this.writeFile(`${dirname}/assets/${uuid}.${ext}`, resource.file);
urlmap.set(resourceUrl, `./assets/${uuid}.${ext}`);
}
Finally, normalizeHtml() rewrites the DOM, replacing src attributes on images, href attributes on stylesheets, and CSS background URLs using the urlmap:
// utils/download/Exporter.ts → normalizeHtml()
imgs.forEach(img => {
const src = img.getAttribute('src') || img.getAttribute('data-src');
if (src && urlmap.has(src)) img.src = urlmap.get(src)!;
});
Video-Specific Asset Handling
Video content requires specialized parsing logic located in utils/index.ts. When the exporter encounters video share sections or embedded iframes, it extracts metadata from JavaScript variables embedded in the HTML:
// utils/index.ts → video handling
const mpVideoCoverUrlMatch = html.match(/window\.__mpVideoCoverUrl\s*=\s*'([^']*)'/);
const mpVideoTransInfoMatch = html.match(/window\.__mpVideoTransInfo\s*=\s*(\[[^\]]+\])/);
The system constructs a download queue containing the cover image and the highest-quality video file, then processes them through the same concurrent pool used for static assets:
await pool.downloads<string>(urls, async (url, proxy) => {
const videoData = await downloadAssetWithProxy<Blob>(url, proxy, false, 10);
const uuid = new Date().getTime() + Math.random().toString();
const ext = mime.getExtension(videoData.type);
zip.file(`assets/${uuid}.${ext}`, videoData);
videoURLMap.set(url, `./assets/${uuid}.${ext}`);
return videoData.size;
});
After download completion, the original iframe placeholder is replaced with a standard HTML5 video element referencing the locally mapped assets:
div.innerHTML = `<video src="${videoURLMap.get(videoUrl)}"
poster="${videoURLMap.get(poster)}"
controls style="width:100%;height:100%;"></video>`;
Key Implementation Files
| File | Purpose |
|---|---|
utils/download/Exporter.ts |
Core orchestration: resource discovery, concurrency management, URL rewriting |
store/v2/resource.ts |
IndexedDB table caching downloaded asset Blob objects keyed by URL |
store/v2/resource-map.ts |
IndexedDB table storing per-article lists of external resource URLs |
utils/index.ts |
Video-specific parsing, cover image extraction, and video tag injection |
types/video.d.ts |
Type definitions for video transcode info and page metadata |
Summary
- Discovery phase scans HTML for
<img>tags (checking bothsrcanddata-src), stylesheet links, and CSSbackground-imagedeclarations, storing findings inresource-map.ts - Download phase uses a bounded concurrency pool (default 5) to fetch assets, leveraging
resource.tscache to eliminate redundant network requests - Mapping phase generates UUID-based filenames with correct MIME extensions, writes assets to
./assets/, and constructs a lookup map to rewrite HTML references - Video handling parses JavaScript variables (
window.__mpVideoCoverUrl) to locate video files, downloads them through the shared pool, and injects<video>elements to replace iframe placeholders - All phases rely on IndexedDB-backed caching to ensure idempotent, resumable exports
Frequently Asked Questions
How does the exporter prevent downloading the same image multiple times?
The system checks the resource cache (store/v2/resource.ts) before initiating any network request. In downloadResourceTask(), the code calls getResourceCache(url); if a Blob exists for that URL, the exporter skips the download and uses the cached version immediately. This deduplication works across all articles sharing the same asset.
What concurrency limits exist to prevent rate limiting?
The processExportQueue() method enforces a default concurrency of 5 simultaneous downloads via a bounded promise pool. The pool uses Promise.race() to refill slots efficiently as individual downloads complete, maximizing throughput without overwhelming remote servers or the proxy manager.
How are video assets handled differently from static images?
Videos require parsing JavaScript variables from the HTML source to locate cover images and transcoded video files. Unlike images that retain their original <img> tags, videos replace iframe placeholders entirely: the system downloads the video and poster assets, builds a videoURLMap, and injects a new <video> element with src and poster attributes pointing to the local files.
What filename strategy prevents asset collisions in the export folder?
During exportHtmlFiles(), the exporter generates unique filenames by concatenating Date.getTime() with Math.random(), then appending the correct extension via mime.getExtension(blob.type). This UUID-style naming ensures that even identically named remote resources (e.g., image.jpg) receive distinct local filenames without overwriting each other.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →