# How to Query Data in Automattic/harper: Core APIs and CLI Techniques

> Query data in Automattic/harper using core APIs and CLI techniques. Explore how Harper represents text as Documents and Tokens, with helpers to query linguistic elements efficiently.

- Repository: [Automattic/harper](https://github.com/Automattic/harper)
- Tags: how-to-guide
- Published: 2026-07-27

---

**TLDR:** Harper represents every text input as a **`Document`** containing a vector of **`Token`** objects, and the sealed **`TokenStringExt`** trait supplies macro-generated helpers such as `first_verb()`, `iter_noun_indices()`, and `get_rel_slice()` to query linguistic elements, while the `harper-cli` binary exposes the same engine through shell commands.

The Automattic/harper repository is a Rust-based grammar and spell checker that treats English text as a structured token stream. Learning how to **query data in Automattic/harper** means leveraging the `Document` type and the `TokenStringExt` trait to locate words, parts of speech, and character spans. Both the Rust API and the `harper-cli` front-end give you direct access to these query primitives.

## Core Querying Concepts in harper-core

Harper’s query model rests on four interconnected primitives defined in the core crate.

### Document

The **`Document`** type is the root container for tokenized text. Defined in [`harper-core/src/document.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/document.rs), it stores the raw character source as `Lrc<[char]>` and a vector of `Token` objects. It provides span-intersection logic, token-wise lookup, and parsing utilities that underpin every query operation.

### TokenStringExt

The **`TokenStringExt`** trait, implemented in [`harper-core/src/token_string_ext.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/token_string_ext.rs), is a sealed extension available only for `Document` and `&[Token]`. It supplies the primary query surface through macro-generated methods such as `first_word()`, `iter_verb_indices()`, `get_rel()`, and `get_rel_slice()`. Because the trait is sealed, you cannot implement it for external types; you consume it through the provided structs.

### TokenKind

Each token carries a **`TokenKind`** enum variant defined in [`harper-core/src/token_kind.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/token_kind.rs). The macro-generated `is_<kind>()` predicates feed directly into `TokenStringExt` helpers, letting you filter tokens by grammatical category without manual pattern matching.

### harper-cli

For users who prefer not to write Rust, [`harper-cli/src/main.rs`](https://github.com/Automattic/harper/blob/main/harper-cli/src/main.rs) wraps the same `Document` and `TokenStringExt` machinery into command-line subcommands. The CLI constructs a `Document` internally and calls helpers like `iter_word_indices()` to emit results.

## Query Patterns in the Rust API

These patterns are implemented in [`harper-core/src/token_string_ext.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/token_string_ext.rs) and consumed through `Document` instances.

### Find the First or Last Token of a Kind

Use the auto-generated `first_<kind>()` and `last_<kind>()` methods to retrieve a single token by grammatical category.

```rust
use harper_core::{Document, TokenStringExt, parsers::PlainEnglish, spell::FstDictionary};

fn main() {
    let doc = Document::new(
        "The quick brown fox jumps over the lazy dog.",
        &PlainEnglish,
        &FstDictionary::curated(),
    );

    if let Some(first_verb) = doc.first_verb() {
        println!("First verb: {}", first_verb.get_str(&doc.source));
    }
}

```

This example calls `first_verb()` on the `Document`, which delegates to the `TokenStringExt` implementation in [`harper-core/src/token_string_ext.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/token_string_ext.rs).

### Iterate Over Indices or Token References

When you need to process every occurrence of a part of speech, use `iter_<kind>_indices()` or `iter_<kind>s()`.

```rust
let noun_indices: Vec<_> = doc.iter_noun_indices().collect();
println!("Noun token indices: {:?}", noun_indices);

```

This pattern is common in custom linting rules that need to inspect every noun.

### Relative Indexing with get_rel

Address tokens from the start or end of the stream with `get_rel()`. Pass `0` for the first token or `-1` for the last.

```rust
let last_punct = doc.get_rel(-1);

```

### Extract Token Slices with get_rel_slice

To capture a contiguous phrase, call `get_rel_slice(start, end)` on the document. This returns a subslice of tokens that you can flatten back into text via `span()`.

```rust
use harper_core::{Document, TokenStringExt, parsers::PlainEnglish, spell::FstDictionary};

let doc = Document::new("She sold seashells by the seashore.", &PlainEnglish, &FstDictionary::curated());

let slice = doc.get_rel_slice(1, 3).unwrap();
let text = slice.span().unwrap().get_content_string(&doc.source);
println!("Extracted phrase: {}", text); // → "sold seashells"

```

The `get_rel_slice` helper is defined alongside `get_rel` in [`harper-core/src/token_string_ext.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/token_string_ext.rs).

### Span-Based Token Lookups

For editor integrations or UI highlights, compute which tokens intersect a character span using `token_indices_intersecting()`. This method lives in [`harper-core/src/document.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/document.rs).

```rust
use harper_core::{Document, TokenStringExt, parsers::PlainEnglish, spell::FstDictionary, Span};

let doc = Document::new(
    "Hello, world! This is a test.",
    &PlainEnglish,
    &FstDictionary::curated(),
);

let start = doc.tokens()[2].span.start;
let end = doc.tokens()[5].span.end;
let span = Span::new(start, end);

let intersecting = doc.token_indices_intersecting(span);
println!("Tokens intersecting the span: {:?}", intersecting);

```

## Query Data via the harper-cli Interface

The `harper-cli` binary exposes subcommands that build a `Document` and run the same `TokenStringExt` methods behind the scenes.

### List Words by Frequency

The `words` subcommand tokenizes the file, calls `iter_words()`, and prints each word with its occurrence count.

```bash
harper-cli words path/to/file.txt

```

To limit output to the top entries, pipe the results:

```bash
harper-cli words README.md | head -n 20

```

### Show Dictionary Metadata

Query the curated dictionary for frequency and spelling metadata without writing Rust.

```bash
harper-cli metadata --words apple banana

```

### Emit Flat Dictionary Entries

External scripts can consume a flattened list of dictionary entries.

```bash
harper-cli words --metadata

```

These commands wire up `FstDictionary::curated()` and `Document::new` exactly as shown in [`harper-cli/src/main.rs`](https://github.com/Automattic/harper/blob/main/harper-cli/src/main.rs).

## Accessing Dictionary Metadata in Rust

You can also query the underlying dictionary directly. The `FstDictionary::curated()` method returns the built-in dictionary, and `metadata_for()` yields frequency data.

```rust
use harper_core::spell::FstDictionary;

let dict = FstDictionary::curated();
let meta = dict.metadata_for("harper").unwrap();
println!("Word: {}, Frequency: {}", meta.word, meta.frequency);

```

## Summary

- Harper treats all input text as a **`Document`** containing a vector of **`Token`** objects.
- The sealed **`TokenStringExt`** trait in [`harper-core/src/token_string_ext.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/token_string_ext.rs) provides macro-generated query helpers like `first_verb()`, `iter_noun_indices()`, `get_rel()`, and `get_rel_slice()`.
- **`TokenKind`** predicates in [`harper-core/src/token_kind.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/token_kind.rs) power the category-based filters.
- **[`harper-core/src/document.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/document.rs)** adds span-based lookups via `token_indices_intersecting()`.
- The **`harper-cli`** front-end in [`harper-cli/src/main.rs`](https://github.com/Automattic/harper/blob/main/harper-cli/src/main.rs) wraps the same APIs into shell commands for word counts and dictionary queries.

## Frequently Asked Questions

### What is the entry point for querying text in Harper?

The entry point is the **`Document`** type in [`harper-core/src/document.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/document.rs). You construct a `Document` with `Document::new(source, parser, dictionary)` and then call methods from the `TokenStringExt` trait to search tokens.

### Can I query Harper without writing Rust?

Yes. The `harper-cli` binary defined in [`harper-cli/src/main.rs`](https://github.com/Automattic/harper/blob/main/harper-cli/src/main.rs) exposes subcommands such as `words` and `metadata` that internally create a `Document` and run the same Rust query helpers for you.

### How does TokenStringExt generate so many query methods?

The trait uses a macro in [`harper-core/src/token_string_ext.rs`](https://github.com/Automattic/harper/blob/main/harper-core/src/token_string_ext.rs) to automatically implement a family of methods—`first_<kind>`, `last_<kind>`, `iter_<kind>_indices`, and `iter_<kind>s`—for every variant of the **`TokenKind`** enum.

### What is the best way to highlight a range of characters in an editor?

Use `Document::token_indices_intersecting(span)`, where `span` is a `Span` struct created from character offsets. This returns the exact token indices that overlap the range, making it ideal for UI overlays.