# How Ladybird Loads and Parses HTML: Inside the Document Loading Pipeline

> Explore Ladybird's document loading pipeline. Learn how it creates a Document object, tokenizes input, and constructs the DOM tree to render HTML efficiently.

- Repository: [Ladybird/ladybird](https://github.com/LadybirdBrowser/ladybird)
- Tags: internals
- Published: 2026-03-05

---

**Ladybird creates a Document object, instantiates an HTMLParser that tokenizes input and runs a tree-construction dispatcher, then finalizes by firing `DOMContentLoaded` and `load` events.**

Ladybird is an independent open-source browser engine written in modern C++. Understanding the process for loading and parsing HTML in Ladybird reveals how the engine transforms raw bytes into a living DOM tree compliant with the HTML5 specification.

## The HTML Loading Pipeline Overview

When Ladybird receives a navigation request, the HTML loading pipeline executes through five distinct stages. The entry point resides in [`Libraries/LibWeb/DOM/DocumentLoading.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/Libraries/LibWeb/DOM/DocumentLoading.cpp), which orchestrates the entire lifecycle from empty document to interactive page.

The pipeline proceeds as follows:

1. **Document Creation** – `DOM::Document::create_and_initialize` constructs a new HTML Document and attaches the MIME type.
2. **Parser Selection** – The engine chooses between minimal document creation or full parser instantiation based on the URL and response body.
3. **Tokenization and Parsing** – `HTMLParser::run()` drives the tokenizer and tree-construction dispatcher.
4. **Insertion Mode Handling** – State machine handlers build the DOM while managing the stack of open elements.
5. **Finalization** – `HTMLParser::the_end()` executes spec-defined steps and fires load events.

## Document Creation and Initialization

The process begins at lines 71-73 of [`DocumentLoading.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/DocumentLoading.cpp) with a call to `DOM::Document::create_and_initialize`. This factory method accepts the document type (`HTML`), the MIME type string, and navigation parameters.

```cpp
auto document = TRY(DOM::Document::create_and_initialize(
    DOM::Document::Type::HTML, "text/html"_string, navigation_params));

```

This creates an empty document shell that will eventually house the parsed DOM tree. The method also establishes the document's origin and sets up the appropriate content type metadata required by the Fetch specification.

## Parser Selection and Instantiation

At lines 85-105 of [`DocumentLoading.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/DocumentLoading.cpp), Ladybird determines which parsing strategy to employ. For `about:blank` URLs with empty bodies, the engine takes a fast path that populates a minimal HTML structure without instantiating the full parser.

For standard documents, the engine constructs an `HTMLParser` using one of two factory methods defined in [`Libraries/LibWeb/HTML/Parser/HTMLParser.h`](https://github.com/LadybirdBrowser/ladybird/blob/main/Libraries/LibWeb/HTML/Parser/HTMLParser.h):

- **`HTMLParser::create_for_scripting`** – Used when the encoding is known in advance.
- **`HTMLParser::create_with_uncertain_encoding`** – Used when the encoding must be sniffed from the byte stream, taking an optional MIME-type hint.

```cpp
GC::Ref<HTMLParser> parser;
if (document->url_string() == "about:blank"_string && body_is_empty) {
    TRY(document->populate_with_html_head_and_body());
    HTMLParser::the_end(document);
} else {
    parser = HTMLParser::create_with_uncertain_encoding(document, data, mime_type);
}

```

The parser holds references to the **stack of open elements** and the **active formatting elements list**, which are crucial for handling misnested tags and auto-closing elements per the HTML5 specification.

## Tokenization and Tree Construction

The core parsing logic lives in [`Libraries/LibWeb/HTML/Parser/HTMLParser.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/Libraries/LibWeb/HTML/Parser/HTMLParser.cpp). The `run()` method (lines 206-260) implements the main event loop that drives the parsing process.

### The Tokenization Stage

First, the parser feeds the raw byte buffer into `HTMLTokenizer`, the lexical analyzer declared in [`HTMLTokenizer.h`](https://github.com/LadybirdBrowser/ladybird/blob/main/HTMLTokenizer.h). This component splits the input stream into distinct tokens—start tags, end tags, character data, comments, and DOCTYPE declarations.

### Insertion Mode Dispatch

After tokenization, the parser consults its **insertion mode** state machine to determine how to process each token. The insertion mode enum (defined at lines 74-78 of [`HTMLParser.h`](https://github.com/LadybirdBrowser/ladybird/blob/main/HTMLParser.h)) includes states such as `Initial`, `InHead`, `InBody`, `Text`, and `AfterBody`.

At lines 25-33 and 100-124 of [`HTMLParser.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/HTMLParser.cpp), the tree-construction dispatcher routes tokens to specialized handlers like `handle_in_head`, `handle_in_body`, and `handle_text`. These handlers:

- Insert elements into the DOM tree
- Manage the stack of open elements to track nesting depth
- Reconstruct the active formatting elements list when encountering phrasing content
- Handle foster parenting for misplaced table elements

```cpp
// Conceptual dispatch logic from HTMLParser.cpp
switch (insertion_mode) {
    case InsertionMode::InHead:
        handle_in_head(token);
        break;
    case InsertionMode::InBody:
        handle_in_body(token);
        break;
    // ... additional modes
}

```

## Finalization and Event Firing

When the tokenizer emits an end-of-file token, the parser executes its finalization sequence at lines 61-68 of [`HTMLParser.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/HTMLParser.cpp). The `flush_character_insertions()` method emits any remaining buffered text nodes, then `HTMLParser::the_end()` performs the "the-end" steps defined by the HTML specification.

This phase fires the `DOMContentLoaded` event once the DOM is fully constructed, followed by the `load` event after all subresources complete. The parser marks the document as finished and detaches itself, handing ownership to the document lifetime manager in [`Libraries/LibWeb/DOM/Document.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/Libraries/LibWeb/DOM/Document.cpp).

## Integration with the Rendering Pipeline

While [`DocumentLoading.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/DocumentLoading.cpp) handles the parsing logic, the UI layer in [`UI/Qt/BrowserWindow.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/UI/Qt/BrowserWindow.cpp) and [`UI/Qt/Tab.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/UI/Qt/Tab.cpp) creates `WebView` instances that observe the document's load lifecycle. Once `HTMLParser::the_end()` completes, the finished DOM tree flows into the rendering engine for layout and paint operations.

## Complete Code Example

The following snippet demonstrates the C++ API for manually parsing HTML content, mirroring the internal browser implementation:

```cpp
#include <LibWeb/DOM/Document.h>
#include <LibWeb/HTML/Parser/HTMLParser.h>
#include <LibWeb/Fetch/Infrastructure/HTTP/MIME.h>
#include <LibTextCodec/Decoder.h>

void parse_html_content(URL::URL const& url, ByteBuffer const& data, 
                        Optional<MimeSniff::MimeType> mime)
{
    // 1. Create an empty HTML Document
    auto document = MUST(DOM::Document::create_and_initialize(
        DOM::Document::Type::HTML, "text/html"_string, 
        HTML::NavigationParams {}));

    // 2. Build a parser that guesses encoding from the byte stream
    auto parser = HTMLParser::create_with_uncertain_encoding(document, data, mime);

    // 3. Run the parser to construct the DOM tree
    parser->run(url);

    // 4. Finalize and fire load events
    HTMLParser::the_end(document, parser);
}

```

This pattern appears throughout [`DocumentLoading.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/DocumentLoading.cpp) when processing network responses.

## Summary

- **Document initialization** occurs via `DOM::Document::create_and_initialize` in [`DocumentLoading.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/DocumentLoading.cpp), establishing the document type and metadata.
- **Parser instantiation** uses `HTMLParser::create_with_uncertain_encoding` for standard loads or `create_for_scripting` when encoding is known, with a fast path for `about:blank` documents.
- **Tokenization** happens in `HTMLTokenizer`, which feeds tokens to the tree-construction dispatcher.
- **Tree construction** follows the HTML5 insertion mode state machine, maintaining the stack of open elements and active formatting elements.
- **Finalization** executes through `HTMLParser::the_end`, firing `DOMContentLoaded` and `load` events before the document enters the interactive state.

## Frequently Asked Questions

### What triggers the HTML parsing process in Ladybird?

A navigation request triggers the HTML parsing process. When the browser receives a response, [`DocumentLoading.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/DocumentLoading.cpp) invokes `DOM::Document::create_and_initialize` to create the document shell, then instantiates `HTMLParser` to begin tokenization. The parser runs synchronously via the `run()` method until the input stream is exhausted.

### How does Ladybird handle character encoding detection?

Ladybird uses `HTMLParser::create_with_uncertain_encoding` when the encoding is unknown. This factory method accepts the raw byte buffer and an optional MIME-type hint, then employs encoding sniffing algorithms to determine the correct character set before tokenization begins. If the encoding is certain, `create_for_scripting` provides a faster path.

### What is the insertion mode state machine in Ladybird's HTML parser?

The insertion mode state machine tracks the parser's current context within the HTML document structure. Defined in [`HTMLParser.h`](https://github.com/LadybirdBrowser/ladybird/blob/main/HTMLParser.h) at lines 74-78, modes include `Initial`, `InHead`, `InBody`, and `Text`. Each mode has a dedicated handler in [`HTMLParser.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/HTMLParser.cpp) (such as `handle_in_body`) that determines how to process incoming tokens and manage the DOM construction rules.

### Where does the HTML parser notify the UI that loading is complete?

The parser calls `HTMLParser::the_end` at lines 61-68 of [`HTMLParser.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/HTMLParser.cpp) to signal completion. This method fires the `DOMContentLoaded` and `load` events on the Document object. The UI layer in [`UI/Qt/Tab.cpp`](https://github.com/LadybirdBrowser/ladybird/blob/main/UI/Qt/Tab.cpp) observes these events through the `WebView` component, which then updates the browser chrome and triggers the initial paint of the rendered page.