Overall Architecture of Automattic/harper: Modular Core, LSP, and WebAssembly Layers
Harper is organized as a modular monorepo that isolates a high-performance Rust core grammar engine from transport layers (LSP and WebAssembly) and presentation layers spanning editors, browsers, and desktop applications.
Harper is an open-source grammar and spell checker designed for speed and privacy. The overall architecture of Automattic/harper follows a strict three-layer pattern that separates text processing concerns from UI delivery, enabling the same Rust core to power everything from VS Code extensions to Tauri-based desktop apps without network dependencies.
Three-Layer Architecture
The project structure follows a logical stack defined in packages/web/src/routes/docs/contributors/architecture/+page.md, dividing functionality into Core, Transport, and Presentation layers.
Core Layer (harper-core)
The heart of the system resides in harper-core, located at harper-core/src/lib.rs. This crate defines the central types—Document, Token, Lint, and Linter—and implements the tokenization, parsing, and analysis pipelines for English text.
Key responsibilities include:
- Language parsing: Markdown and plain-text parsers (e.g.,
harper-core::parsers::Markdown) emit streams ofTokenobjects. - Rule execution:
LintGroupand individualLinterimplementations walk the token stream, applying patterns fromharper-core::expr::*to detect grammar and spelling issues. - Overlap resolution: Utility functions like
remove_overlapsandremove_lints_overlapping_exprde-duplicate conflicting suggestions before delivery.
Supporting crates provide optional features:
harper-brill: Part-of-speech (POS) tagging.harper-thesaurus: Semantic synonym lookups.harper-pos-utils: POS utilities.harper-dictionary-wordlist: Built-in wordlists.
Transport Layer (harper-ls and harper-wasm)
The transport layer exposes the core engine to external environments through two primary interfaces:
harper-ls (harper-ls/src/main.rs) implements the Language Server Protocol (LSP), allowing IDEs like VS Code, Neovim, Helix, Emacs, and Sublime to communicate with the core over TCP or stdio. It acts as a thin wrapper that converts Lint objects into LSP diagnostics.
harper-wasm (harper-wasm/src/lib.rs) compiles harper-core to WebAssembly, enabling JavaScript environments to load the engine in browsers or Node.js. This crate serves as the bridge between Rust and the harper.js package.
Presentation Layer (Editors and Applications)
The presentation layer consumes the transport interfaces to deliver user-facing functionality:
harper.js: A TypeScript wrapper (packages/harper.js) that loadsharper-wasm, manages configuration, and exposes a simple API (lint(text) → Lint[]). It powers the web demo and browser-based integrations.- Editor plugins: Lightweight wrappers in
packages/vscode-plugin/src/extension.tsand similar directories translate LSP messages into UI actions. harper-desktop: A Tauri-based application (harper-desktop/src-tauri/src/main.rs) bundling the core engine, a local LSP server, and a Svelte-Kit UI for offline Markdown editing.
Text Processing Pipeline
Regardless of the entry point, all text flows through a standardized five-step pipeline:
- Document Creation: Input text is transformed into a
Documentstruct using constructors likeDocument::new_plain_english_curated(). - Tokenization: Language-specific parsers segment the document into
Tokenstreams. - Lint Execution: The
LintGroupiterates through rules, producingLintobjects that identify errors with spans and suggested fixes. - Overlap Resolution: The system removes redundant diagnostics using overlap removal algorithms.
- Front-End Delivery: Results serialize to JSON (for WebAssembly) or LSP diagnostics (for
harper-ls) depending on the transport layer.
Key Design Goals
The architecture optimizes for three primary constraints:
- Memory efficiency: The engine maintains approximately 1/50th the memory footprint of LanguageTool.
- Speed: Processing occurs in sub-millisecond timeframes for typical document sizes.
- Privacy: All operations run locally unless the user explicitly configures a remote service; no network calls occur by default.
Code Examples
Using harper-core Directly in Rust
use harper_core::{Document, LintGroup, spell::FstDictionary, Dialect};
fn main() {
// Build a curated dictionary and a US-English dialect
let dict = FstDictionary::curated();
let dialect = Dialect::American;
// Create a lint group that includes all built-in rules
let lint_group = LintGroup::new_curated(dict, dialect);
// Parse some text
let doc = Document::new_plain_english_curated("She dont like cats.");
// Run the linter
let lints = lint_group.lint(&doc);
// Print the diagnostics
for lint in lints {
println!("{} → {}", lint.span, lint.message);
}
}
Relevant source: harper-core/src/lib.rs (exports Document, LintGroup, etc.).
Running Harper via the Language Server (harper-ls)
# Start the server (listens on 127.0.0.1:4000)
harper-ls
Configure an LSP-compatible client (e.g., Neovim):
require('lspconfig').harper_ls.setup{
cmd = {"harper-ls"},
filetypes = {"markdown", "text"},
}
The client receives diagnostics automatically as you type.
Using harper.js in a Node Project
import { lint } from "harper.js";
(async () => {
const text = "Its a beautiful day, but the greeter said hello.";
const results = await lint(text); // returns an array of lint objects
console.log(results);
})();
harper.js internally loads harper-wasm (see packages/harper.js for the wrapper).
Summary
- Harper is a modular monorepo with a Rust core (
harper-core) handling all grammar logic, separated from transport mechanisms. harper-lsprovides LSP integration for editors, whileharper-wasmenables browser and Node.js usage viaharper.js.- Data flow moves from
Documentcreation through tokenization and linting, with overlap resolution ensuring clean diagnostics. - Design priorities emphasize low memory usage, sub-millisecond latency, and privacy-first local processing.
Frequently Asked Questions
What is the role of harper-core in the architecture?
harper-core is the foundational crate that defines all grammar-checking logic, including the Document, Token, and Lint types, as well as the Linter trait and rule implementations. All other components depend on this crate, as implemented in harper-core/src/lib.rs.
How does harper-ls communicate with editors?
harper-ls implements the Language Server Protocol (LSP) to communicate over stdio or TCP, translating between editor requests and the core's Lint outputs. Editor plugins in packages/vscode-plugin/ and similar directories act as thin clients that spawn the binary.
Can Harper run in a web browser?
Yes. The harper-wasm crate compiles harper-core to WebAssembly, which harper.js loads to provide a JavaScript API. This powers the web demo and any browser extensions without requiring a backend server.
What are the supporting crates in the Harper ecosystem?
Several specialized crates extend harper-core functionality: harper-brill handles part-of-speech tagging, harper-thesaurus provides synonym data, harper-pos-utils contains POS utilities, and harper-dictionary-wordlist manages built-in wordlists. These are optional features enabled as needed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →