How Instatic Super Import Handles HTML and Static Site Conversion

Instatic's Super Import system converts arbitrary HTML snippets or entire static-site ZIP archives into editable PageNode fragments through a rule-driven pipeline that parses DOM, harvests styles, strips unsafe content, and splices the result into the live page tree as an atomic, undoable transaction.

The Instatic CMS (available at CoreBunch/Instatic) provides a sophisticated Super Import capability that bridges the gap between legacy static sites and its node-based page editor. This system ingests raw HTML strings or complete static-site archives, transforming them into first-class PageNode fragments that integrate seamlessly with the existing page tree while preserving styling and media assets.

The Two-Phase Architecture

The Super Import pipeline splits processing between two coordinated subsystems: HTML Import for individual documents and Site Import for complete static sites.

Phase 1: HTML Import Core

Individual HTML processing lives in src/core/htmlImport/walkAndMap.ts, which exports the primary entry point importHtml(). This module takes an HTML string and returns a flat fragment of PageNode objects ready for insertion. The core mapping logic resides in src/core/htmlImport/rules.ts, which declares the rule-driven HTML_TO_MODULE_RULES table that maps DOM elements to specific modules (base.text, base.image, base.container, etc.). A catch-all * rule ensures every element maps to a node without fall-through gaps.

Phase 2: Site Import Orchestration

For full static-site conversion, src/core/siteImport/htmlPagePlan.ts orchestrates the per-page import process. The src/core/siteImport/adapter.ts file defines the contract used by the Super Import UI to handle ZIP archives or folders containing multiple HTML pages, linked CSS, and media assets. This layer aggregates CSS into synthetic per-page sources (<htmlPath>::inline) and surfaces conflicts like duplicate slugs or CSS selector collisions before any data writes occur.

The HTML Import Pipeline Deep Dive

The importHtml() function implements a six-step pure-function pipeline that processes raw HTML into CMS-ready data structures.

DOM Parsing and Safety Sanitization

The pipeline begins with parseHtml(source), which wraps new DOMParser().parseFromString(source, 'text/html') to create a document object. Immediately after parsing, three harvester functions execute:

  • harvestInlineStyles(doc): Records every element's style="…" attribute values
  • collectStyleCss(doc): Concatenates all <style> blocks before removal
  • stripUnsafe(doc): Removes <script> tags, all on* event attributes, and unsafe <style> elements, returning a StripReport for UI toast messages

Only CSS properties permitted by isEmittableProperty survive; unsafe properties drop from both inline styles and parsed <style> blocks.

Rule-Based Node Mapping

The walkAndMap(doc, inlineStyles) function traverses the cleaned DOM, matching each element against HTML_TO_MODULE_RULES. For each match, it creates a PageNode, attaches harvested inlineStyles, and records raw classIds. The result is a fragment object containing { nodes, rootIds, styleCss, stripped }, where nodes is a flat Record<string, PageNode>.

Atomic Insertion and Class Resolution

Once the fragment exists, insertImportedNodes(parentId, fragment, opts?) in src/admin/pages/site/store/slices/site/nodeActions.ts handles the actual tree insertion. This function:

  1. Resolves HTML class names to existing style-rule registry IDs or auto-creates bare classes
  2. Rewrites the node's classIds to the generated registry IDs
  3. Registers new rules derived from <style> blocks via cssToStyleRules
  4. Executes everything inside a single mutateActiveTreeAndSite call, making the operation atomic and undoable

Static Site ZIP Import Workflow

When users drop a ZIP archive into the Super Import interface (handled by src/admin/modals/ImportHtml/ImportHtmlModal.tsx), the system executes the applyImport() function. This process:

  • Iterates every HTML page in the archive
  • Runs the HTML Import pipeline for each page
  • Rewrites url(…) references to uploaded media assets
  • Merges page-level <body> metadata (classIds, safe props, inlineStyles) into base.body nodes
  • Displays conflict detection results before committing changes

The UI provides three entry points for these operations: the Spotlight Import HTML command (editor.importHtml in src/admin/spotlight/commands/importHtml.ts), the right-click Paste HTML here… context menu on container nodes, and the drag-and-drop Super Import zone for ZIP archives.

Code Implementation Examples

Import a raw HTML snippet using the core engine:

import { importHtml } from '@core/htmlImport';

const html = `<div class="hero"><h1>Hello</h1><style>.hero{color:red}</style></div>`;
const { nodes, rootIds, styleCss, stripped } = importHtml(html);
// nodes: Record<string, PageNode>
// styleCss: Raw CSS for later parsing

Insert the fragment into the active page tree via the editor store:

import { insertImportedNodes } from 'src/admin/pages/site/store/slices/site/nodeActions';

insertImportedNodes(parentId, {
  nodes,
  rootIds,
  styleCss,          // Consumer will call cssToStyleRules(...)
  body: undefined,  // Optional <body> metadata
});

Execute a complete static-site import from a ZIP file:

import { applyImport } from '@core/siteImport';

const zipFile = /* File object from drop event */;
await applyImport(zipFile); 
// Runs per-page pipeline, resolves media, shows conflict UI

Summary

  • Rule-driven mapping: The declarative HTML_TO_MODULE_RULES in src/core/htmlImport/rules.ts guarantees every HTML element maps to a specific module without gaps.
  • Safety-first sanitization: The stripUnsafe() function removes <script> tags and event handlers while isEmittableProperty filters CSS properties before they enter the registry.
  • Atomic transactions: All import operations execute through mutateActiveTreeAndSite, allowing single-step undo of complex multi-page imports.
  • Framework decoupling: The HTML import layer returns raw styleCss rather than parsed CSS, letting the site-import layer apply breakpoint contexts and avoid circular dependencies between @core/htmlImport and @core/siteImport.
  • Conflict detection: The Super Import adapter surfaces slug collisions and selector clashes in a preview dialog before writing data.

Frequently Asked Questions

How does Instatic handle potentially dangerous HTML content during import?

The stripUnsafe() function in src/core/htmlImport/stripUnsafe.ts removes all <script> tags, on* event attributes, and unsafe <style> elements before DOM walking begins. Additionally, only CSS properties passing the isEmittableProperty check survive from inline styles and <style> blocks, preventing arbitrary code execution while preserving presentational styling.

Can I import an entire static website or only individual HTML snippets?

Both modes are supported. Individual snippets use importHtml() from @core/htmlImport, accessible via the Spotlight command or paste dialog. Complete static sites require the Super Import workflow via applyImport() in @core/siteImport, which processes ZIP archives containing multiple HTML files, linked CSS, and media assets through src/core/siteImport/htmlPagePlan.ts.

What happens to CSS classes when importing external HTML?

During insertImportedNodes(), the system walks the fragment's classIds and either reuses existing style-rule registry entries or auto-creates new bare classes. Raw HTML class names rewrite to generated registry IDs, and rules from <style> blocks register against these IDs, ensuring imported styling integrates with Instatic's design system.

Is the import operation reversible if something goes wrong?

Yes. The entire import—whether a single HTML paste or a full-site ZIP import—executes inside a single mutateActiveTreeAndSite transaction. This makes the operation atomic, allowing users to revert the complete import with a single undo action if conflicts or errors are discovered after insertion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →