# How Lightpanda's DOM Tree Differs from Browser DOM Implementations

> Discover how Lightpanda's Zig-native DOM tree differs from browser DOM by leveraging union-based polymorphism and doubly-linked lists for efficient AI data extraction. Learn more.

- Repository: [Lightpanda/browser](https://github.com/lightpanda-io/browser)
- Tags: internals
- Published: 2026-03-14

---

**Lightpanda implements a deterministic, Zig-native DOM tree using union-based node polymorphism and doubly-linked child lists, eliminating C++ inheritance hierarchies and rendering pipelines to optimize for AI-driven semantic data extraction.**

Lightpanda is a headless browser engine written in Zig and hosted at `lightpanda-io/browser`, designed specifically for programmatic access and automated data extraction rather than visual rendering. While standard browsers like Chrome or Firefox rely on monolithic C++ engines with tightly coupled layout and paint systems, Lightpanda's DOM implementation treats the document as a pure data model, using flat memory structures and explicit ownership tracking to enable deterministic AI reasoning.

## Union-Based Node Architecture vs. C++ Inheritance

Mainstream browser engines use classical C++ inheritance where `Element`, `Text`, and `Document` extend a base `Node` class with virtual method dispatch. Lightpanda diverges significantly by implementing all node types within a single **flat union structure**.

In `src/browser/webapi/Node.zig` (lines 53-60), the `Node` struct contains a `_type` field that is a union of pointers to concrete implementations:

```zig
pub const Type = union(enum) {
    cdata: *CData,
    element: *Element,
    document: *Document,
    // ...
};
_type: Type,

```

This design avoids virtual function table overhead and ensures contiguous memory layout, making tree traversals cache-friendly and allocation patterns predictable. The entire node hierarchy lives in one compile-time-defined type rather than scattered C++ objects.

## Doubly-Linked Child Lists vs. Vector Storage

Where browsers typically store child nodes in resizable vectors (such as `nsINode::mChildren` in Gecko or `Node::children_` in Blink), Lightpanda uses a **doubly-linked list** from Zig's standard library to maintain sibling relationships.

Each `Node` in `src/browser/webapi/Node.zig` (lines 48-49) embeds the linkage directly:

```zig
_child_link: LinkedList.Node = .{},
_children: ?*Children = null,

```

This pattern enables O(1) subtree splicing during `appendChild` or `insertBefore` operations without reallocation or pointer invalidation. The `Children` wrapper manages the list head and tail, while individual nodes carry their own link metadata, mirroring the DOM specification's node list semantics without dynamic array resizing costs.

## First-Class Shadow DOM Implementation

Lightpanda treats **ShadowRoot** as a distinct Zig struct rather than a subclass of `DocumentFragment`. According to `src/browser/webapi/ShadowRoot.zig` (lines 27-68), shadow roots attach to elements via a dedicated hashmap lookup (`Element.ShadowRootLookup`), completely isolating shadow tree traversal from the primary document tree.

The implementation exposes standard methods like `mode`, `host`, and `getElementById` through dedicated bridge methods, while `isInShadowTree()` and `getRootNode({.composed = true})` handle composition. This separation allows Lightpanda to support shadow DOM semantics for semantic extraction without integrating complex style scoping or layout trees.

## Explicit Owner-Document Resolution

Standard browsers infer `ownerDocument` by walking a node's parent chain, which fails for detached nodes. Lightpanda solves this through an **explicit hash map** called `OwnerDocumentLookup` maintained at the page level, referenced in `src/browser/webapi/Node.zig` (lines 50-52).

Because Lightpanda frequently creates nodes off-document when preparing semantic snapshots, this lookup table provides constant-time resolution of document ownership without tree traversal, a necessity for the browser's AI-oriented use cases where nodes may exist outside the main DOM during processing.

## Lightweight Event and API Systems

Unlike browsers that implement full micro-task queues and observer patterns integrated with the JavaScript event loop, Lightpanda's `EventManager` in `src/browser/EventManager.zig` (lines 392-452) manually queues DOM mutation events when mutation methods are called. This lightweight approach skips the heavy synchronization required for visual rendering pipelines.

API exposure occurs through **compile-time bridge generation**. Each Zig type publishes a `js.Bridge` that maps native methods to JavaScript-visible properties using Zig's compile-time reflection, as seen in `src/browser/webapi/Node.zig` (lines 334-363). This replaces the WebIDL compilation pipelines used in Blink or WebKit, generating bindings directly from Zig struct definitions without external tooling.

## Zig Error Handling for DOM Exceptions

Rather than throwing JavaScript `DOMException` objects mapped from internal C++ error codes, Lightpanda uses **Zig's error system directly**. Operations like tree insertion return specific error values defined in `src/browser/webapi/Node.zig` (lines 299-311):

```zig
error.HierarchyError,  // Invalid node hierarchy
error.NotFound,        // Node not found in tree
// ...

```

These propagate through the JavaScript bridge and convert to appropriate exceptions at the boundary, eliminating the need for complex exception object allocation within the core DOM logic.

## Practical DOM Manipulation Examples

### Creating Documents and Elements

Document creation follows a factory pattern on the page object:

```zig
// Create a new empty document
var doc = try page.createDocument();

// Create <div id="root">
var div = try page.createElement("div");
try div.setAttribute("id", "root", page);

// Append the div to the document
try doc.appendChild(div, page);

```

These methods are implemented in `src/browser/Page.zig`, delegating to the underlying `Node` insertion logic.

### Tree Traversal

Traversal uses an explicit iterator pattern rather than recursion:

```zig
fn walk(node: *Node) void {
    const name = node.getNodeName(&page.buf);
    std.debug.print("{s}\n", .{name});

    var it = node.childrenIterator();
    while (it.next()) |child| {
        walk(child);
    }
}

walk(page.document);

```

The `childrenIterator` method (referenced around lines 60-70 in `Node.zig`) leverages the doubly-linked list structure for efficient iteration, while `getNodeName` (lines 232-236) handles node type dispatch through the union.

### Attaching Shadow Roots

Shadow DOM attachment follows the standard API but uses Lightpanda's internal hashmap linkage:

```zig
var host = try page.createElement("custom-element");
var shadow = try host.attachShadow("open", page);

var span = try page.createElement("span");
try span.setAttribute("class", "label", page);
try shadow.appendChild(span, page);

```

The `attachShadow` implementation in `src/browser/webapi/Element.zig` (lines 647-652) initializes the shadow root via `ShadowRoot.zig` (lines 44-52), registering it in the element's shadow lookup table.

### Serializing to JSON

For AI pipeline integration, nodes serialize directly to JSON:

```zig
var buf = std.json.Stringify.init(allocator);
defer buf.deinit();

try node.jsonStringify(&buf);
const json = buf.toString();

```

This method, found in `src/browser/webapi/Node.zig` (lines 1089-1095), traverses the union-based tree and outputs a structured representation without DOM-to-HTML string conversion overhead.

## Key Source Files

Understanding Lightpanda's DOM requires examining these specific implementation files:

- **`src/browser/webapi/Node.zig`** — Core DOM node definition, union type dispatch, doubly-linked child management, and JavaScript bridge interface (lines 334-363 for JsApi, lines 48-60 for structure).
- **`src/browser/webapi/Element.zig`** — Element-specific attribute handling and shadow root attachment logic (lines 647-652).
- **`src/browser/webapi/ShadowRoot.zig`** — Shadow DOM implementation with host linkage and mode handling (lines 27-68).
- **`src/browser/Page.zig`** — Factory for document creation and root node management.
- **`src/browser/EventManager.zig`** — Lightweight mutation event dispatch (lines 392-452).
- **`src/browser/SemanticTree.zig`** — Lightpanda-specific semantic extraction layer for AI processing.
- **`src/browser/webapi/js/bridge.zig`** — Compile-time JavaScript binding generation.
- **`src/testing.zig`** — Minimal HTML test runner for deterministic unit testing.

## Summary

- **Lightpanda uses Zig unions instead of C++ inheritance** for node polymorphism, eliminating virtual function overhead and ensuring memory contiguity.
- **Child nodes reside in doubly-linked lists** embedded within each node struct, enabling O(1) insertion and removal without dynamic array reallocation.
- **Shadow DOM is implemented as a separate struct** with hashmap-based host lookup rather than inheritance from `DocumentFragment`.
- **Owner document resolution uses an explicit hashmap** to handle detached nodes efficiently for semantic extraction workflows.
- **Error handling leverages Zig's error sets** directly, converting to JavaScript exceptions only at the API boundary.
- **No layout or rendering engine exists**; the DOM serves purely as a scriptable data model for AI agents and automated testing.

## Frequently Asked Questions

### How does Lightpanda handle DOM mutations without a layout engine?

Lightpanda queues mutation events manually through `EventManager.zig` (lines 392-452) when structural methods like `appendChild` are invoked. Because there is no visual rendering pipeline, mutations do not trigger style recalculation, layout, or paint operations, resulting in deterministic, immediate updates suitable for programmatic data extraction.

### Can Lightpanda's DOM implementation replace a full browser for testing?

Yes, Lightpanda ships with a minimal HTML runner (`testing.htmlRunner` in `src/testing.zig` and referenced in `Node.zig` lines 1891-1894) that drives the Zig DOM directly without browser chrome. This provides deterministic, fast unit testing for web components without the overhead of Web Platform Test harnesses or full browser initialization.

### Why does Lightpanda use a hashmap for ownerDocument lookup instead of parent chain walking?

Lightpanda frequently creates nodes off-document during semantic snapshot preparation for AI processing. Walking the parent chain fails for these detached nodes. The `OwnerDocumentLookup` hashmap in `Node.zig` (lines 50-52) provides constant-time resolution of document ownership regardless of a node's current tree position, which is essential for maintaining proper DOM semantics during complex extraction operations.

### Is Lightpanda's DOM API compatible with standard WebIDL specifications?

Lightpanda exposes JavaScript-compatible APIs through compile-time bridge generation (`js.Bridge` in `Node.zig` lines 334-363), which maps Zig methods to JavaScript properties. While it follows DOM specifications for node manipulation and shadow DOM, it does not implement the full WebIDL compiler toolchain used in Chromium or Firefox, focusing instead on the subset required for headless automation and semantic extraction.