How Instatic's Super Import Converts Raw HTML/CSS into Editable Nodes
Instatic's Super Import uses a deterministic pipeline in src/core/htmlImport/walkAndMap.ts to parse raw HTML, strip unsafe content through stripUnsafe, recursively map DOM elements to PageNode objects via walkAndMap, and extract CSS into the styleCss field for the editor's Style panel.
Instatic's visual editor manipulates a structured node tree to enable real-time page editing. The Super Import feature bridges static HTML documents and this editable ecosystem by transforming raw markup into native PageNode objects that the editor can manipulate immediately. According to the CoreBunch/Instatic source code, this conversion happens through a single public API that performs sanitization, parsing, and tree mapping in one atomic pass.
The Five-Stage Import Pipeline
The importHtml function exported from src/core/htmlImport/walkAndMap.ts processes HTML strings through five deterministic stages that ensure safe, accurate conversion to editable nodes.
Stage 1: Parsing the Raw Markup
The pipeline begins by invoking parseHtml to convert the raw HTML string into a DOM-like tree structure. This initial parse provides a traversable foundation for the sanitization and mapping steps that follow.
Stage 2: Sanitizing Unsafe Content
Before traversal, the stripUnsafe function removes potentially dangerous elements including <script> tags, on* event attributes, and dangerous URLs. This stage guarantees that imported fragments cannot execute code when rendered in the visual editor, and sets a stripped flag in the final result if any content was removed.
Stage 3: Walking and Mapping to PageNodes
The walkAndMap function recursively traverses the sanitized DOM tree, converting each element into a PageNode as defined in src/core/page-tree/types.ts. During this walk, the implementation extracts:
- HTML attributes → stored on
props.htmlAttributes - Inline styles → converted to camel-case and stored on
inlineStyles - Class names → transformed into
classIdsfor the selector registry - Text nodes → become
base.textnodes - Semantic elements → images become
base.image, links becomebase.link, and forms map to corresponding built-in modules
All children are processed recursively, preserving the original document nesting in the resulting node tree.
Stage 4: Extracting Page-Scoped CSS
While walking the DOM, the importer collects content from inline <style> blocks and <link rel="stylesheet"> references. These styles concatenate into the styleCss field of the ImportResult, which later persists as a page-scoped stylesheet and appears in the editor's Style panel.
Stage 5: Returning the Import Result
The function returns an ImportResult object containing:
nodes: AMap<string, PageNode>of all created nodesrootIds: Array identifiers for top-level nodes (typically the<body>element)styleCss: Concatenated CSS string for style extractionstripped: Boolean indicating whether unsafe content was removed
Integration with Editor Architecture
The importer integrates with Instatic's broader architecture through src/core/siteImport/htmlPagePlan.ts, which orchestrates the site-import pipeline. When users trigger imports via the admin interface—handled by src/admin/spotlight/commands/importHtml.ts—the system calls importHtml(htmlSource) and commits resulting nodes using standard editor mutations like insertNode and replaceNodeHtml.
Extracted CSS flows through src/core/publisher/cssCollector.ts to populate the style registry, ensuring that imported designs remain editable alongside natively created content.
Practical Implementation Example
The following TypeScript demonstrates importing raw HTML and inserting it into the active page tree:
import { importHtml } from '@core/htmlImport'
import { mutateActiveTree } from '@core/page-tree'
import { insertNode, addPageStyle } from 'src/admin/pages/site/store/slices/site/helpers'
const rawHtml = `
<html>
<head>
<style>.hero { background:#f0f0f0; }</style>
</head>
<body class="hero">
<h1>Hello world</h1>
<p data-note="example">A paragraph with <a href="/page">a link</a>.</p>
<img src="/photo.jpg" alt="Example" />
</body>
</html>
`
const { nodes, rootIds, styleCss, stripped } = importHtml(rawHtml)
// Insert each root node into the current page
rootIds.forEach(rootId => {
const node = nodes[rootId]
mutateActiveTree(state => {
insertNode({ parentId: state.activePageId, node })
})
})
// Persist extracted CSS to the page's style panel
addPageStyle({ pageId: 'current-page-id', css: styleCss })
This results in a base.body node containing base.text nodes for the heading and paragraph, a base.link module for the anchor tag, and a base.image node for the photograph, with the .hero class available in the Style panel for editing.
Summary
- Core Conversion: Instatic's Super Import relies on
importHtmlinsrc/core/htmlImport/walkAndMap.tsto transform HTML into editable nodes. - Safety First: The
stripUnsafefunction eliminates scripts and event handlers before DOM traversal begins. - Node Mapping: The
walkAndMaputility converts HTML elements to specific node types (base.text,base.image,base.link) while preserving attributes inprops.htmlAttributesand styles ininlineStyles. - CSS Preservation: Style blocks extract into the
styleCssfield for integration withsrc/core/publisher/cssCollector.tsand the visual editor's Style panel. - Seamless Integration: The
ImportResultinterface providesnodes,rootIds, and metadata for immediate insertion via standard mutation APIs likeinsertNode.
Frequently Asked Questions
What file contains the main importHtml function?
The core implementation lives in src/core/htmlImport/walkAndMap.ts. This file exports the importHtml function and contains the parseHtml, stripUnsafe, and walkAndMap utilities that power the conversion process.
How does Instatic handle unsafe HTML during import?
The stripUnsafe function removes all <script> tags, on* event handlers, and dangerous URLs before the DOM traversal begins. The ImportResult returns a stripped boolean flag indicating whether unsafe content was detected and removed, ensuring the editing surface remains secure.
Can imported CSS be edited within the Instatic visual editor?
Yes. The importer extracts CSS from <style> tags and external stylesheets into the styleCss field, which integrates with src/core/publisher/cssCollector.ts. This CSS appears in the editor's Style panel as page-scoped stylesheets that users can modify alongside Tailwind utilities.
What node types does the importer create from HTML elements?
The walkAndMap function maps standard HTML to specific module types: text nodes become base.text, images become base.image, links become base.link, and the body element transforms into base.body. Generic containers map to structural nodes while preserving original HTML attributes in props.htmlAttributes for complete styling control.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →