How to Query Data in Automattic/harper: Core APIs and CLI Techniques
TLDR: Harper represents every text input as a Document containing a vector of Token objects, and the sealed TokenStringExt trait supplies macro-generated helpers such as first_verb(), iter_noun_indices(), and get_rel_slice() to query linguistic elements, while the harper-cli binary exposes the same engine through shell commands.
The Automattic/harper repository is a Rust-based grammar and spell checker that treats English text as a structured token stream. Learning how to query data in Automattic/harper means leveraging the Document type and the TokenStringExt trait to locate words, parts of speech, and character spans. Both the Rust API and the harper-cli front-end give you direct access to these query primitives.
Core Querying Concepts in harper-core
Harper’s query model rests on four interconnected primitives defined in the core crate.
Document
The Document type is the root container for tokenized text. Defined in harper-core/src/document.rs, it stores the raw character source as Lrc<[char]> and a vector of Token objects. It provides span-intersection logic, token-wise lookup, and parsing utilities that underpin every query operation.
TokenStringExt
The TokenStringExt trait, implemented in harper-core/src/token_string_ext.rs, is a sealed extension available only for Document and &[Token]. It supplies the primary query surface through macro-generated methods such as first_word(), iter_verb_indices(), get_rel(), and get_rel_slice(). Because the trait is sealed, you cannot implement it for external types; you consume it through the provided structs.
TokenKind
Each token carries a TokenKind enum variant defined in harper-core/src/token_kind.rs. The macro-generated is_<kind>() predicates feed directly into TokenStringExt helpers, letting you filter tokens by grammatical category without manual pattern matching.
harper-cli
For users who prefer not to write Rust, harper-cli/src/main.rs wraps the same Document and TokenStringExt machinery into command-line subcommands. The CLI constructs a Document internally and calls helpers like iter_word_indices() to emit results.
Query Patterns in the Rust API
These patterns are implemented in harper-core/src/token_string_ext.rs and consumed through Document instances.
Find the First or Last Token of a Kind
Use the auto-generated first_<kind>() and last_<kind>() methods to retrieve a single token by grammatical category.
use harper_core::{Document, TokenStringExt, parsers::PlainEnglish, spell::FstDictionary};
fn main() {
let doc = Document::new(
"The quick brown fox jumps over the lazy dog.",
&PlainEnglish,
&FstDictionary::curated(),
);
if let Some(first_verb) = doc.first_verb() {
println!("First verb: {}", first_verb.get_str(&doc.source));
}
}
This example calls first_verb() on the Document, which delegates to the TokenStringExt implementation in harper-core/src/token_string_ext.rs.
Iterate Over Indices or Token References
When you need to process every occurrence of a part of speech, use iter_<kind>_indices() or iter_<kind>s().
let noun_indices: Vec<_> = doc.iter_noun_indices().collect();
println!("Noun token indices: {:?}", noun_indices);
This pattern is common in custom linting rules that need to inspect every noun.
Relative Indexing with get_rel
Address tokens from the start or end of the stream with get_rel(). Pass 0 for the first token or -1 for the last.
let last_punct = doc.get_rel(-1);
Extract Token Slices with get_rel_slice
To capture a contiguous phrase, call get_rel_slice(start, end) on the document. This returns a subslice of tokens that you can flatten back into text via span().
use harper_core::{Document, TokenStringExt, parsers::PlainEnglish, spell::FstDictionary};
let doc = Document::new("She sold seashells by the seashore.", &PlainEnglish, &FstDictionary::curated());
let slice = doc.get_rel_slice(1, 3).unwrap();
let text = slice.span().unwrap().get_content_string(&doc.source);
println!("Extracted phrase: {}", text); // → "sold seashells"
The get_rel_slice helper is defined alongside get_rel in harper-core/src/token_string_ext.rs.
Span-Based Token Lookups
For editor integrations or UI highlights, compute which tokens intersect a character span using token_indices_intersecting(). This method lives in harper-core/src/document.rs.
use harper_core::{Document, TokenStringExt, parsers::PlainEnglish, spell::FstDictionary, Span};
let doc = Document::new(
"Hello, world! This is a test.",
&PlainEnglish,
&FstDictionary::curated(),
);
let start = doc.tokens()[2].span.start;
let end = doc.tokens()[5].span.end;
let span = Span::new(start, end);
let intersecting = doc.token_indices_intersecting(span);
println!("Tokens intersecting the span: {:?}", intersecting);
Query Data via the harper-cli Interface
The harper-cli binary exposes subcommands that build a Document and run the same TokenStringExt methods behind the scenes.
List Words by Frequency
The words subcommand tokenizes the file, calls iter_words(), and prints each word with its occurrence count.
harper-cli words path/to/file.txt
To limit output to the top entries, pipe the results:
harper-cli words README.md | head -n 20
Show Dictionary Metadata
Query the curated dictionary for frequency and spelling metadata without writing Rust.
harper-cli metadata --words apple banana
Emit Flat Dictionary Entries
External scripts can consume a flattened list of dictionary entries.
harper-cli words --metadata
These commands wire up FstDictionary::curated() and Document::new exactly as shown in harper-cli/src/main.rs.
Accessing Dictionary Metadata in Rust
You can also query the underlying dictionary directly. The FstDictionary::curated() method returns the built-in dictionary, and metadata_for() yields frequency data.
use harper_core::spell::FstDictionary;
let dict = FstDictionary::curated();
let meta = dict.metadata_for("harper").unwrap();
println!("Word: {}, Frequency: {}", meta.word, meta.frequency);
Summary
- Harper treats all input text as a
Documentcontaining a vector ofTokenobjects. - The sealed
TokenStringExttrait inharper-core/src/token_string_ext.rsprovides macro-generated query helpers likefirst_verb(),iter_noun_indices(),get_rel(), andget_rel_slice(). TokenKindpredicates inharper-core/src/token_kind.rspower the category-based filters.harper-core/src/document.rsadds span-based lookups viatoken_indices_intersecting().- The
harper-clifront-end inharper-cli/src/main.rswraps the same APIs into shell commands for word counts and dictionary queries.
Frequently Asked Questions
What is the entry point for querying text in Harper?
The entry point is the Document type in harper-core/src/document.rs. You construct a Document with Document::new(source, parser, dictionary) and then call methods from the TokenStringExt trait to search tokens.
Can I query Harper without writing Rust?
Yes. The harper-cli binary defined in harper-cli/src/main.rs exposes subcommands such as words and metadata that internally create a Document and run the same Rust query helpers for you.
How does TokenStringExt generate so many query methods?
The trait uses a macro in harper-core/src/token_string_ext.rs to automatically implement a family of methods—first_<kind>, last_<kind>, iter_<kind>_indices, and iter_<kind>s—for every variant of the TokenKind enum.
What is the best way to highlight a range of characters in an editor?
Use Document::token_indices_intersecting(span), where span is a Span struct created from character offsets. This returns the exact token indices that overlap the range, making it ideal for UI overlays.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →