Dependencies for Automattic/harper: Complete Cargo Workspace Guide

The Automattic/harper repository organizes over 20 crates into a Cargo workspace, where harper-core relies on libraries like fst, pulldown-cmark, and serde for text processing, while harper-ls extends this with tower-lsp-server and tokio for language-server functionality.

Harper is an open-source grammar and spell checker written in Rust. The project uses a workspace architecture defined in the root Cargo.toml to manage dependencies across its modular ecosystem, ranging from the core parsing engine to WebAssembly bindings and desktop applications.

Workspace Architecture Overview

The repository follows Rust's standard workspace pattern. The root Cargo.toml declares all member crates, ensuring version consistency and enabling intra-workspace linking via path dependencies.

Key workspace members include:

  • harper-core – The language-checking engine
  • harper-ls – Language-server protocol implementation
  • harper-cli – Command-line interface
  • harper-wasm – WebAssembly wrapper
  • harper-desktop/src-tauri – Desktop application backend
  • harper-brill – Brill tagger integration
  • harper-thesaurus – Optional synonym lookup

Each crate maintains its own Cargo.toml file, declaring both external crates.io dependencies and internal path references to sibling crates.

harper-core Dependencies

harper-core serves as the heart of the project. Located in harper-core/Cargo.toml, it combines finite-state transducers, markdown parsing, and Unicode handling to analyze text.

Data Structures and Algorithms

  • fst (0.4.7) – Finite-state transducers for fast spell-checking
  • hashbrown (0.16.1) – High-performance hash maps with serde support
  • smallvec (1.15.1) – Stack-allocated vectors for performance
  • lru (0.18.0) – Least-recently-used cache implementation
  • foldhash (0.2.0) – Fast hashing algorithm
  • trie-rs (0.4.2) – Trie data structures for string storage

Text Processing and Parsing

  • pulldown-cmark (0.13.3) – Markdown parser
  • unicode-blocks, unicode-script, unicode-width – Unicode categorization and width calculation
  • levenshtein_automata (0.2.1) – Approximate string matching with fst_automaton feature
  • regex (1.12.3) – Regular expression engine

Serialization and Utilities

  • serde / serde_json (1.0.228 / 1.0.150) – Data serialization with derive macros
  • thiserror (2.0.18) – Error handling macros
  • itertools (0.14.0) – Iterator extensions
  • strum / strum_macros (0.28.0) – Enum utilities
  • bitflags (2.11.0) – Bitflag handling with serde support

Internal and Optional Dependencies

  • harper-brill (path: "../harper-brill") – Brill-tagger integration at version 2.0.0
  • harper-thesaurus (path: "../harper-thesaurus", optional) – Synonym lookup enabled via the thesaurus feature
  • ammonia (4.1.2) – HTML sanitization
  • cached (0.59.0) – Memoization for expensive computations
  • zip (8.6.0) – ZIP archive handling with deflate support

Macro Crates

  • paste (1.0.14) – Token concatenation in macros
  • is-macro (0.3.6) – Macro detection utilities
  • blanket (0.4.0) – Helper macros for blanket trait implementations
  • ordered-float (5.3.0) – Float wrappers that implement Ord with serde support

Language Server and CLI Dependencies

harper-ls (LSP Server)

The harper-ls/Cargo.toml builds atop harper-core to provide IDE integration:

  • tower-lsp-server (0.22.1) – LSP protocol implementation
  • tokio (1.52.1) – Asynchronous runtime with selected features
  • clap (4.6.0) – Command-line argument parsing
  • dirs (6.0.0) – Platform-specific directory resolution
  • anyhow (1.0.102) – Flexible error handling
  • tracing / tracing-subscriber – Structured logging infrastructure
  • futures – Asynchronous programming utilities
  • globset – Glob pattern matching
  • resolve-path, open – Path resolution and file opening utilities

The crate also includes path dependencies to numerous Harper crates like harper-stats, harper-comments, and harper-typst.

harper-cli

The command-line interface defined in harper-cli/Cargo.toml maintains a lean dependency tree:

  • clap for CLI parsing
  • anyhow for error propagation
  • serde_json for output serialization
  • harper-core via path dependency

Platform-Specific Wrappers

WebAssembly (harper-wasm)

The harper-wasm/Cargo.toml targets browser environments with:

  • wasm-bindgen – Rust/JavaScript bindings
  • js-sys – JavaScript global bindings
  • wee-alloc – Small WebAssembly allocator

Desktop Application (harper-desktop)

Located at harper-desktop/src-tauri/Cargo.toml, the desktop backend combines:

  • tauri – Application framework
  • serde – Data serialization
  • anyhow – Error handling
  • harper-core – Core engine integration

Linguistic Components

How to Inspect the Dependency Tree

Cargo provides built-in tooling to visualize how crates interconnect. Run these commands from the repository root:


# Display complete dependency graph for entire workspace

cargo tree --all-features

# View only harper-core dependencies

cargo tree -p harper-core --all-features

# Check harper-ls with optional features enabled

cargo tree -p harper-ls --features concurrent,thesaurus

These commands resolve feature flags and show the exact versions locked in Cargo.lock, revealing how external libraries like tokio or fst integrate with internal workspace crates.

Summary

  • Automattic/harper uses a Cargo workspace to coordinate over 20 interconnected crates
  • harper-core forms the foundation, depending on text-processing crates including pulldown-cmark, fst, and regex, plus serialization via serde
  • harper-ls adds asynchronous infrastructure with tokio and tower-lsp-server for IDE integration
  • Platform wrappers introduce specialized dependencies: wasm-bindgen for WebAssembly and tauri for desktop
  • Internal path dependencies (e.g., harper-brill, harper-thesaurus) ensure version alignment across the workspace
  • Use cargo tree commands to audit the full dependency graph and feature resolution

Frequently Asked Questions

What crate provides the spell-checking algorithm in harper-core?

The fst crate (version 0.4.7) provides finite-state transducers for fast spell-checking, while levenshtein_automata (version 0.2.1) enables approximate string matching. These work alongside hashbrown for high-performance hash maps and trie-rs for prefix-tree storage of dictionary words.

Does Harper require tokio for command-line usage?

No, tokio (version 1.52.1) is only required when using harper-ls, the Language Server Protocol implementation. The harper-cli tool operates synchronously without an async runtime, depending only on clap, anyhow, and serde_json for its core functionality.

How does harper-wasm differ from harper-core in terms of dependencies?

While harper-core focuses on text-processing algorithms using standard Rust crates like regex and pulldown-cmark, harper-wasm introduces WebAssembly-specific dependencies: wasm-bindgen for JavaScript interoperability, js-sys for browser API bindings, and wee-alloc as a size-optimized memory allocator suitable for browser environments.

Where are dependency versions managed in Harper's Cargo workspace?

Dependency versions are managed hierarchically. The root Cargo.toml defines the workspace members list, while each subdirectory (e.g., harper-core/Cargo.toml, harper-ls/Cargo.toml) declares specific versions for its external crates. Internal dependencies use path = "../crate-name" syntax to link workspace members without version constraints, ensuring the compiler uses the local source code directly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →