Overall Architecture of Automattic/harper: Modular Core, LSP, and WebAssembly Layers

Harper is organized as a modular monorepo that isolates a high-performance Rust core grammar engine from transport layers (LSP and WebAssembly) and presentation layers spanning editors, browsers, and desktop applications.

Harper is an open-source grammar and spell checker designed for speed and privacy. The overall architecture of Automattic/harper follows a strict three-layer pattern that separates text processing concerns from UI delivery, enabling the same Rust core to power everything from VS Code extensions to Tauri-based desktop apps without network dependencies.

Three-Layer Architecture

The project structure follows a logical stack defined in packages/web/src/routes/docs/contributors/architecture/+page.md, dividing functionality into Core, Transport, and Presentation layers.

Core Layer (harper-core)

The heart of the system resides in harper-core, located at harper-core/src/lib.rs. This crate defines the central types—Document, Token, Lint, and Linter—and implements the tokenization, parsing, and analysis pipelines for English text.

Key responsibilities include:

  • Language parsing: Markdown and plain-text parsers (e.g., harper-core::parsers::Markdown) emit streams of Token objects.
  • Rule execution: LintGroup and individual Linter implementations walk the token stream, applying patterns from harper-core::expr::* to detect grammar and spelling issues.
  • Overlap resolution: Utility functions like remove_overlaps and remove_lints_overlapping_expr de-duplicate conflicting suggestions before delivery.

Supporting crates provide optional features:

  • harper-brill: Part-of-speech (POS) tagging.
  • harper-thesaurus: Semantic synonym lookups.
  • harper-pos-utils: POS utilities.
  • harper-dictionary-wordlist: Built-in wordlists.

Transport Layer (harper-ls and harper-wasm)

The transport layer exposes the core engine to external environments through two primary interfaces:

harper-ls (harper-ls/src/main.rs) implements the Language Server Protocol (LSP), allowing IDEs like VS Code, Neovim, Helix, Emacs, and Sublime to communicate with the core over TCP or stdio. It acts as a thin wrapper that converts Lint objects into LSP diagnostics.

harper-wasm (harper-wasm/src/lib.rs) compiles harper-core to WebAssembly, enabling JavaScript environments to load the engine in browsers or Node.js. This crate serves as the bridge between Rust and the harper.js package.

Presentation Layer (Editors and Applications)

The presentation layer consumes the transport interfaces to deliver user-facing functionality:

  • harper.js: A TypeScript wrapper (packages/harper.js) that loads harper-wasm, manages configuration, and exposes a simple API (lint(text) → Lint[]). It powers the web demo and browser-based integrations.
  • Editor plugins: Lightweight wrappers in packages/vscode-plugin/src/extension.ts and similar directories translate LSP messages into UI actions.
  • harper-desktop: A Tauri-based application (harper-desktop/src-tauri/src/main.rs) bundling the core engine, a local LSP server, and a Svelte-Kit UI for offline Markdown editing.

Text Processing Pipeline

Regardless of the entry point, all text flows through a standardized five-step pipeline:

  1. Document Creation: Input text is transformed into a Document struct using constructors like Document::new_plain_english_curated().
  2. Tokenization: Language-specific parsers segment the document into Token streams.
  3. Lint Execution: The LintGroup iterates through rules, producing Lint objects that identify errors with spans and suggested fixes.
  4. Overlap Resolution: The system removes redundant diagnostics using overlap removal algorithms.
  5. Front-End Delivery: Results serialize to JSON (for WebAssembly) or LSP diagnostics (for harper-ls) depending on the transport layer.

Key Design Goals

The architecture optimizes for three primary constraints:

  • Memory efficiency: The engine maintains approximately 1/50th the memory footprint of LanguageTool.
  • Speed: Processing occurs in sub-millisecond timeframes for typical document sizes.
  • Privacy: All operations run locally unless the user explicitly configures a remote service; no network calls occur by default.

Code Examples

Using harper-core Directly in Rust

use harper_core::{Document, LintGroup, spell::FstDictionary, Dialect};

fn main() {
    // Build a curated dictionary and a US-English dialect
    let dict = FstDictionary::curated();
    let dialect = Dialect::American;

    // Create a lint group that includes all built-in rules
    let lint_group = LintGroup::new_curated(dict, dialect);

    // Parse some text
    let doc = Document::new_plain_english_curated("She dont like cats.");

    // Run the linter
    let lints = lint_group.lint(&doc);

    // Print the diagnostics
    for lint in lints {
        println!("{} → {}", lint.span, lint.message);
    }
}

Relevant source: harper-core/src/lib.rs (exports Document, LintGroup, etc.).

Running Harper via the Language Server (harper-ls)


# Start the server (listens on 127.0.0.1:4000)

harper-ls

Configure an LSP-compatible client (e.g., Neovim):

require('lspconfig').harper_ls.setup{
  cmd = {"harper-ls"},
  filetypes = {"markdown", "text"},
}

The client receives diagnostics automatically as you type.

Using harper.js in a Node Project

import { lint } from "harper.js";

(async () => {
  const text = "Its a beautiful day, but the greeter said hello.";
  const results = await lint(text); // returns an array of lint objects
  console.log(results);
})();

harper.js internally loads harper-wasm (see packages/harper.js for the wrapper).

Summary

  • Harper is a modular monorepo with a Rust core (harper-core) handling all grammar logic, separated from transport mechanisms.
  • harper-ls provides LSP integration for editors, while harper-wasm enables browser and Node.js usage via harper.js.
  • Data flow moves from Document creation through tokenization and linting, with overlap resolution ensuring clean diagnostics.
  • Design priorities emphasize low memory usage, sub-millisecond latency, and privacy-first local processing.

Frequently Asked Questions

What is the role of harper-core in the architecture?

harper-core is the foundational crate that defines all grammar-checking logic, including the Document, Token, and Lint types, as well as the Linter trait and rule implementations. All other components depend on this crate, as implemented in harper-core/src/lib.rs.

How does harper-ls communicate with editors?

harper-ls implements the Language Server Protocol (LSP) to communicate over stdio or TCP, translating between editor requests and the core's Lint outputs. Editor plugins in packages/vscode-plugin/ and similar directories act as thin clients that spawn the binary.

Can Harper run in a web browser?

Yes. The harper-wasm crate compiles harper-core to WebAssembly, which harper.js loads to provide a JavaScript API. This powers the web demo and any browser extensions without requiring a backend server.

What are the supporting crates in the Harper ecosystem?

Several specialized crates extend harper-core functionality: harper-brill handles part-of-speech tagging, harper-thesaurus provides synonym data, harper-pos-utils contains POS utilities, and harper-dictionary-wordlist manages built-in wordlists. These are optional features enabled as needed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →