How Instatic Super Import Handles HTML and Static Site Conversion
Instatic's Super Import system converts arbitrary HTML snippets or entire static-site ZIP archives into editable PageNode fragments through a rule-driven pipeline that parses DOM, harvests styles, strips unsafe content, and splices the result into the live page tree as an atomic, undoable transaction.
The Instatic CMS (available at CoreBunch/Instatic) provides a sophisticated Super Import capability that bridges the gap between legacy static sites and its node-based page editor. This system ingests raw HTML strings or complete static-site archives, transforming them into first-class PageNode fragments that integrate seamlessly with the existing page tree while preserving styling and media assets.
The Two-Phase Architecture
The Super Import pipeline splits processing between two coordinated subsystems: HTML Import for individual documents and Site Import for complete static sites.
Phase 1: HTML Import Core
Individual HTML processing lives in src/core/htmlImport/walkAndMap.ts, which exports the primary entry point importHtml(). This module takes an HTML string and returns a flat fragment of PageNode objects ready for insertion. The core mapping logic resides in src/core/htmlImport/rules.ts, which declares the rule-driven HTML_TO_MODULE_RULES table that maps DOM elements to specific modules (base.text, base.image, base.container, etc.). A catch-all * rule ensures every element maps to a node without fall-through gaps.
Phase 2: Site Import Orchestration
For full static-site conversion, src/core/siteImport/htmlPagePlan.ts orchestrates the per-page import process. The src/core/siteImport/adapter.ts file defines the contract used by the Super Import UI to handle ZIP archives or folders containing multiple HTML pages, linked CSS, and media assets. This layer aggregates CSS into synthetic per-page sources (<htmlPath>::inline) and surfaces conflicts like duplicate slugs or CSS selector collisions before any data writes occur.
The HTML Import Pipeline Deep Dive
The importHtml() function implements a six-step pure-function pipeline that processes raw HTML into CMS-ready data structures.
DOM Parsing and Safety Sanitization
The pipeline begins with parseHtml(source), which wraps new DOMParser().parseFromString(source, 'text/html') to create a document object. Immediately after parsing, three harvester functions execute:
harvestInlineStyles(doc): Records every element'sstyle="…"attribute valuescollectStyleCss(doc): Concatenates all<style>blocks before removalstripUnsafe(doc): Removes<script>tags, allon*event attributes, and unsafe<style>elements, returning aStripReportfor UI toast messages
Only CSS properties permitted by isEmittableProperty survive; unsafe properties drop from both inline styles and parsed <style> blocks.
Rule-Based Node Mapping
The walkAndMap(doc, inlineStyles) function traverses the cleaned DOM, matching each element against HTML_TO_MODULE_RULES. For each match, it creates a PageNode, attaches harvested inlineStyles, and records raw classIds. The result is a fragment object containing { nodes, rootIds, styleCss, stripped }, where nodes is a flat Record<string, PageNode>.
Atomic Insertion and Class Resolution
Once the fragment exists, insertImportedNodes(parentId, fragment, opts?) in src/admin/pages/site/store/slices/site/nodeActions.ts handles the actual tree insertion. This function:
- Resolves HTML class names to existing style-rule registry IDs or auto-creates bare classes
- Rewrites the node's
classIdsto the generated registry IDs - Registers new rules derived from
<style>blocks viacssToStyleRules - Executes everything inside a single
mutateActiveTreeAndSitecall, making the operation atomic and undoable
Static Site ZIP Import Workflow
When users drop a ZIP archive into the Super Import interface (handled by src/admin/modals/ImportHtml/ImportHtmlModal.tsx), the system executes the applyImport() function. This process:
- Iterates every HTML page in the archive
- Runs the HTML Import pipeline for each page
- Rewrites
url(…)references to uploaded media assets - Merges page-level
<body>metadata (classIds, safe props,inlineStyles) intobase.bodynodes - Displays conflict detection results before committing changes
The UI provides three entry points for these operations: the Spotlight Import HTML command (editor.importHtml in src/admin/spotlight/commands/importHtml.ts), the right-click Paste HTML here… context menu on container nodes, and the drag-and-drop Super Import zone for ZIP archives.
Code Implementation Examples
Import a raw HTML snippet using the core engine:
import { importHtml } from '@core/htmlImport';
const html = `<div class="hero"><h1>Hello</h1><style>.hero{color:red}</style></div>`;
const { nodes, rootIds, styleCss, stripped } = importHtml(html);
// nodes: Record<string, PageNode>
// styleCss: Raw CSS for later parsing
Insert the fragment into the active page tree via the editor store:
import { insertImportedNodes } from 'src/admin/pages/site/store/slices/site/nodeActions';
insertImportedNodes(parentId, {
nodes,
rootIds,
styleCss, // Consumer will call cssToStyleRules(...)
body: undefined, // Optional <body> metadata
});
Execute a complete static-site import from a ZIP file:
import { applyImport } from '@core/siteImport';
const zipFile = /* File object from drop event */;
await applyImport(zipFile);
// Runs per-page pipeline, resolves media, shows conflict UI
Summary
- Rule-driven mapping: The declarative
HTML_TO_MODULE_RULESinsrc/core/htmlImport/rules.tsguarantees every HTML element maps to a specific module without gaps. - Safety-first sanitization: The
stripUnsafe()function removes<script>tags and event handlers whileisEmittablePropertyfilters CSS properties before they enter the registry. - Atomic transactions: All import operations execute through
mutateActiveTreeAndSite, allowing single-step undo of complex multi-page imports. - Framework decoupling: The HTML import layer returns raw
styleCssrather than parsed CSS, letting the site-import layer apply breakpoint contexts and avoid circular dependencies between@core/htmlImportand@core/siteImport. - Conflict detection: The Super Import adapter surfaces slug collisions and selector clashes in a preview dialog before writing data.
Frequently Asked Questions
How does Instatic handle potentially dangerous HTML content during import?
The stripUnsafe() function in src/core/htmlImport/stripUnsafe.ts removes all <script> tags, on* event attributes, and unsafe <style> elements before DOM walking begins. Additionally, only CSS properties passing the isEmittableProperty check survive from inline styles and <style> blocks, preventing arbitrary code execution while preserving presentational styling.
Can I import an entire static website or only individual HTML snippets?
Both modes are supported. Individual snippets use importHtml() from @core/htmlImport, accessible via the Spotlight command or paste dialog. Complete static sites require the Super Import workflow via applyImport() in @core/siteImport, which processes ZIP archives containing multiple HTML files, linked CSS, and media assets through src/core/siteImport/htmlPagePlan.ts.
What happens to CSS classes when importing external HTML?
During insertImportedNodes(), the system walks the fragment's classIds and either reuses existing style-rule registry entries or auto-creates new bare classes. Raw HTML class names rewrite to generated registry IDs, and rules from <style> blocks register against these IDs, ensuring imported styling integrates with Instatic's design system.
Is the import operation reversible if something goes wrong?
Yes. The entire import—whether a single HTML paste or a full-site ZIP import—executes inside a single mutateActiveTreeAndSite transaction. This makes the operation atomic, allowing users to revert the complete import with a single undo action if conflicts or errors are discovered after insertion.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →