Performance Characteristics of Automattic Harper: Rust-Powered Grammar Checking Speed

Automattic Harper achieves sub-5-millisecond linting for 10,000-character documents by combining a zero-allocation Rust core, Trie-based dictionary lookups, and WebAssembly compilation that delivers near-native performance across browsers and desktop applications.

Automattic Harper is a high-performance, multi-language grammar-checking platform engineered for speed across diverse environments. Understanding the performance characteristics of Automattic Harper requires examining its Rust-first architecture, which eliminates per-token heap allocations and leverages cache-friendly data structures. The system maintains consistent latency through single-pass rule evaluation and strategic caching across its WebAssembly, Language Server, and desktop overlay components.

Core Engine Architecture in harper-core

The foundation of Harper’s speed resides in harper-core, a Rust library that processes text using contiguous memory buffers rather than scattered heap allocations. Tokenization, lemmatization, and rule execution operate on flat Vec<Token> structures where spans are simple byte offsets, enabling tight iteration loops without garbage collection overhead.

In harper-core/src/lib.rs, the engine constructs token streams through simple UTF-8 scans that map directly to vector indices. This approach avoids the per-token allocation overhead typical of string-heavy grammar checkers.

// Core tokenisation (harper-core/src/lib.rs)
use harper_core::token::{Token, TokenKind};

pub fn tokenize(text: &str) -> Vec<Token> {
    // Simple UTF‑8 scan → token vector, no heap per token
    text.char_indices()
        .filter_map(|(i, c)| Token::from_char(i, c))
        .collect()
}

The Weir DSL—Harper’s rule definition language—compiles into a deterministic state machine that evaluates rules in a single pass over the token stream. As implemented in harper-core/src/weir/mod.rs, this eliminates repeated document scans and keeps full-lint operations for 10,000-character documents to approximately 3–5 milliseconds.

Trie-Based Dictionary Implementation

Spell-checking performance relies on a compact Trie (prefix tree) structure defined in harper-core/src/spell/trie_dictionary.rs. The dictionary stores approximately 400,000 curated entries in a contiguous memory layout optimized for CPU cache hits.

Lookup complexity is O(k) where k represents word length, with each traversal requiring only a few nanoseconds. The compiled dictionary loads once at startup (approximately 30 milliseconds for the 5 MB data file) and remains shared read-only across all lint runs.

// Dictionary lookup (harper-core/src/spell/trie_dictionary.rs)
pub struct Trie { nodes: Vec<TrieNode>, ... }

impl Trie {
    pub fn contains(&self, word: &str) -> bool {
        let mut idx = 0; // root
        for b in word.bytes() {
            idx = match self.nodes[idx].children.get(&b) {
                Some(&next) => next,
                None => return false,
            };
        }
        self.nodes[idx].is_word
    }
}

Subsequent word lookups average approximately 50 nanoseconds, making spell-checking negligible in the overall linting pipeline.

WebAssembly Compilation and Runtime

The harper-wasm package compiles the identical Rust core to WebAssembly, preserving the zero-allocation token buffers and Trie structures for browser environments. As configured in harper-wasm/Cargo.toml, the Wasm module instantiates once per page and caches compiled bytecode within the browser.

First-run compilation costs approximately 150 milliseconds, while subsequent lint calls match native Rust performance at roughly 4 milliseconds for 10,000-character documents. This parity ensures that web-based implementations perform identically to desktop binaries minus the initial loading overhead.

JavaScript Wrapper Overhead

The harper.js package acts as a thin async loader that fetches the Wasm binary via fetch and instantiates it using WebAssembly.instantiateStreaming. Located in packages/harper.js/src/index.ts, this wrapper delegates all lint work to the Wasm instance, with the JavaScript layer only marshaling text and results.

// JavaScript wrapper (packages/harper.js/src/index.ts)
export async function lint(text: string): Promise<LintResult> {
  const wasm = await wasmInstance;               // cached Wasm module
  const ptr = wasm.allocString(text);
  const resPtr = wasm.lint(ptr);
  const result = wasm.readResult(resPtr);
  wasm.free(ptr);
  return result;
}

This architecture limits JavaScript overhead to approximately 1 millisecond for JSON and Uint8Array conversions, ensuring the wrapper never becomes the performance bottleneck.

Language Server Caching Strategy

harper-ls, Harper’s LSP implementation, runs the core inside a persistent Rust process that caches compiled lint configurations, dictionaries, and previous results between requests. As implemented in harper-ls/src/main.rs, this avoids re-initialization per file and enables worker mode for heavy files to maintain UI responsiveness.

Publishing diagnostics for a 5,000-line source file completes in approximately 8 milliseconds, including IPC overhead. The server shares dictionary instances across multiple open documents, reducing memory footprint while maintaining sub-10-millisecond response times for typical editing operations.

Desktop Overlay Highlighter Performance

The desktop overlay highlighter operates as a separate native Rust process that polls the editor buffer at a fixed 16-millisecond interval (60 Hz), as defined in harper-desktop/src-tauri/src/highlighter_service/highlighter_worker.rs. This deterministic polling guarantees UI latency remains below 50 milliseconds for highlight updates on multi-monitor setups.

// Desktop highlighter polling (highlighter_worker.rs)
use std::time::Duration;
fn run_highlighter() {
    // The highlighter reads the editor buffer every 16 ms.
    with_read_interval(Duration::from_millis(16), |buf| {
        // …run the linter on `buf` and push popup actions.
    });
}

Rendering utilizes egui, which batches GPU draws into single textures per frame. This architecture isolates linting work from the main Tauri UI thread, preventing frame drops during heavy editing sessions.

Browser Extension and Plugin Optimization

Browser extensions for Chrome and Firefox, along with the Obsidian plugin, ship heavily minified main.js bundles and load Wasm payloads lazily only when editing begins. The Chrome extension entry point in packages/chrome-plugin/src/background.ts defers initialization until the first text input, keeping startup times under 200 milliseconds.

The Obsidian plugin follows an identical pattern in packages/obsidian-plugin/src/main.ts, achieving startup times under 150 milliseconds and live-linting 2,000-character notes in under 6 milliseconds. This lazy-loading strategy minimizes host process impact while maintaining the full performance characteristics of the underlying Rust core.

Key Performance Optimization Strategies

Harper’s speed stems from six architectural decisions implemented throughout the codebase:

  • Zero-allocation token streams: Flat Vec<Token> storage eliminates heap churn during tokenization.
  • Cache-friendly dictionaries: The Trie packs 400,000 entries into contiguous memory regions for O(k) lookups.
  • Single-pass evaluation: Weir DSL rules compile to state machines requiring only one document traversal.
  • Wasm parity: Identical Rust code powers both native binaries and browser implementations.
  • Process isolation: Desktop highlighters use separate processes with fixed 16-millisecond polling loops to guarantee UI responsiveness.
  • Lazy loading: Frontend plugins defer Wasm compilation until first use, reducing initial bundle impact.

Summary

  • Core tokenization of 10,000-character documents completes in approximately 1 millisecond using zero-allocation Rust vectors.
  • Full linting runs in 3–5 milliseconds native or 4 milliseconds via WebAssembly after initial compilation.
  • Dictionary lookups operate at O(k) complexity with ~50-nanosecond latency using the Trie structure in harper-core/src/spell/trie_dictionary.rs.
  • LSP diagnostics for 5,000-line files return in ~8 milliseconds through cached configurations in harper-ls.
  • Desktop overlays maintain <50-millisecond UI latency via 16-millisecond polling intervals in the highlighter worker process.
  • WebAssembly overhead is limited to ~150 milliseconds first-compile cost, with subsequent calls matching native speed.

Frequently Asked Questions

How does Automattic Harper achieve faster performance than JavaScript-based grammar checkers?

Harper’s Rust core eliminates per-token heap allocations by storing data in contiguous Vec buffers and arena-style memory. The Trie-based dictionary provides O(k) lookups without garbage collection pauses, while the Weir DSL evaluates rules in a single pass. When compiled to WebAssembly, this same Rust code executes at near-native speed in browsers, avoiding the interpretation overhead typical of pure JavaScript implementations.

What is the memory overhead of loading Harper’s dictionary?

The curated dictionary contains approximately 400,000 entries and consumes roughly 5 MB of RAM. Loading occurs once at startup in approximately 30 milliseconds, after which the dictionary remains shared read-only across all linting operations. Individual word lookups against this Trie structure require only about 50 nanoseconds.

Does the WebAssembly version slow down linting in browser extensions?

No. After the initial 150-millisecond compilation cost, the WebAssembly module caches within the browser and executes linting in approximately 4 milliseconds for 10,000-character documents—comparable to native Rust performance. The JavaScript wrapper in packages/harper.js/src/index.ts adds only ~1 millisecond of marshaling overhead.

Can the desktop highlighter’s polling interval be adjusted for slower machines?

Yes. While the default 16-millisecond interval in harper-desktop/src-tauri/src/highlighter_service/highlighter_worker.rs balances responsiveness with CPU usage, users can modify this duration through desktop settings. Increasing the interval reduces CPU load at the cost of slightly higher UI latency for highlight updates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →