Dependencies for Automattic/harper: Complete Cargo Workspace Guide
The Automattic/harper repository organizes over 20 crates into a Cargo workspace, where harper-core relies on libraries like fst, pulldown-cmark, and serde for text processing, while harper-ls extends this with tower-lsp-server and tokio for language-server functionality.
Harper is an open-source grammar and spell checker written in Rust. The project uses a workspace architecture defined in the root Cargo.toml to manage dependencies across its modular ecosystem, ranging from the core parsing engine to WebAssembly bindings and desktop applications.
Workspace Architecture Overview
The repository follows Rust's standard workspace pattern. The root Cargo.toml declares all member crates, ensuring version consistency and enabling intra-workspace linking via path dependencies.
Key workspace members include:
harper-core– The language-checking engineharper-ls– Language-server protocol implementationharper-cli– Command-line interfaceharper-wasm– WebAssembly wrapperharper-desktop/src-tauri– Desktop application backendharper-brill– Brill tagger integrationharper-thesaurus– Optional synonym lookup
Each crate maintains its own Cargo.toml file, declaring both external crates.io dependencies and internal path references to sibling crates.
harper-core Dependencies
harper-core serves as the heart of the project. Located in harper-core/Cargo.toml, it combines finite-state transducers, markdown parsing, and Unicode handling to analyze text.
Data Structures and Algorithms
- fst (0.4.7) – Finite-state transducers for fast spell-checking
- hashbrown (0.16.1) – High-performance hash maps with serde support
- smallvec (1.15.1) – Stack-allocated vectors for performance
- lru (0.18.0) – Least-recently-used cache implementation
- foldhash (0.2.0) – Fast hashing algorithm
- trie-rs (0.4.2) – Trie data structures for string storage
Text Processing and Parsing
- pulldown-cmark (0.13.3) – Markdown parser
- unicode-blocks, unicode-script, unicode-width – Unicode categorization and width calculation
- levenshtein_automata (0.2.1) – Approximate string matching with
fst_automatonfeature - regex (1.12.3) – Regular expression engine
Serialization and Utilities
- serde / serde_json (1.0.228 / 1.0.150) – Data serialization with derive macros
- thiserror (2.0.18) – Error handling macros
- itertools (0.14.0) – Iterator extensions
- strum / strum_macros (0.28.0) – Enum utilities
- bitflags (2.11.0) – Bitflag handling with serde support
Internal and Optional Dependencies
- harper-brill (path: "../harper-brill") – Brill-tagger integration at version 2.0.0
- harper-thesaurus (path: "../harper-thesaurus", optional) – Synonym lookup enabled via the
thesaurusfeature - ammonia (4.1.2) – HTML sanitization
- cached (0.59.0) – Memoization for expensive computations
- zip (8.6.0) – ZIP archive handling with deflate support
Macro Crates
- paste (1.0.14) – Token concatenation in macros
- is-macro (0.3.6) – Macro detection utilities
- blanket (0.4.0) – Helper macros for blanket trait implementations
- ordered-float (5.3.0) – Float wrappers that implement Ord with serde support
Language Server and CLI Dependencies
harper-ls (LSP Server)
The harper-ls/Cargo.toml builds atop harper-core to provide IDE integration:
- tower-lsp-server (0.22.1) – LSP protocol implementation
- tokio (1.52.1) – Asynchronous runtime with selected features
- clap (4.6.0) – Command-line argument parsing
- dirs (6.0.0) – Platform-specific directory resolution
- anyhow (1.0.102) – Flexible error handling
- tracing / tracing-subscriber – Structured logging infrastructure
- futures – Asynchronous programming utilities
- globset – Glob pattern matching
- resolve-path, open – Path resolution and file opening utilities
The crate also includes path dependencies to numerous Harper crates like harper-stats, harper-comments, and harper-typst.
harper-cli
The command-line interface defined in harper-cli/Cargo.toml maintains a lean dependency tree:
clapfor CLI parsinganyhowfor error propagationserde_jsonfor output serializationharper-corevia path dependency
Platform-Specific Wrappers
WebAssembly (harper-wasm)
The harper-wasm/Cargo.toml targets browser environments with:
- wasm-bindgen – Rust/JavaScript bindings
- js-sys – JavaScript global bindings
- wee-alloc – Small WebAssembly allocator
Desktop Application (harper-desktop)
Located at harper-desktop/src-tauri/Cargo.toml, the desktop backend combines:
- tauri – Application framework
- serde – Data serialization
- anyhow – Error handling
- harper-core – Core engine integration
Linguistic Components
- harper-brill (
harper-brill/Cargo.toml) usesserdeandserde_jsonfor model serialization - harper-thesaurus (
harper-thesaurus/Cargo.toml) handles optional synonym data loading with minimal dependencies
How to Inspect the Dependency Tree
Cargo provides built-in tooling to visualize how crates interconnect. Run these commands from the repository root:
# Display complete dependency graph for entire workspace
cargo tree --all-features
# View only harper-core dependencies
cargo tree -p harper-core --all-features
# Check harper-ls with optional features enabled
cargo tree -p harper-ls --features concurrent,thesaurus
These commands resolve feature flags and show the exact versions locked in Cargo.lock, revealing how external libraries like tokio or fst integrate with internal workspace crates.
Summary
- Automattic/harper uses a Cargo workspace to coordinate over 20 interconnected crates
- harper-core forms the foundation, depending on text-processing crates including
pulldown-cmark,fst, andregex, plus serialization viaserde - harper-ls adds asynchronous infrastructure with
tokioandtower-lsp-serverfor IDE integration - Platform wrappers introduce specialized dependencies:
wasm-bindgenfor WebAssembly andtaurifor desktop - Internal path dependencies (e.g.,
harper-brill,harper-thesaurus) ensure version alignment across the workspace - Use
cargo treecommands to audit the full dependency graph and feature resolution
Frequently Asked Questions
What crate provides the spell-checking algorithm in harper-core?
The fst crate (version 0.4.7) provides finite-state transducers for fast spell-checking, while levenshtein_automata (version 0.2.1) enables approximate string matching. These work alongside hashbrown for high-performance hash maps and trie-rs for prefix-tree storage of dictionary words.
Does Harper require tokio for command-line usage?
No, tokio (version 1.52.1) is only required when using harper-ls, the Language Server Protocol implementation. The harper-cli tool operates synchronously without an async runtime, depending only on clap, anyhow, and serde_json for its core functionality.
How does harper-wasm differ from harper-core in terms of dependencies?
While harper-core focuses on text-processing algorithms using standard Rust crates like regex and pulldown-cmark, harper-wasm introduces WebAssembly-specific dependencies: wasm-bindgen for JavaScript interoperability, js-sys for browser API bindings, and wee-alloc as a size-optimized memory allocator suitable for browser environments.
Where are dependency versions managed in Harper's Cargo workspace?
Dependency versions are managed hierarchically. The root Cargo.toml defines the workspace members list, while each subdirectory (e.g., harper-core/Cargo.toml, harper-ls/Cargo.toml) declares specific versions for its external crates. Internal dependencies use path = "../crate-name" syntax to link workspace members without version constraints, ensuring the compiler uses the local source code directly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →