What Database Types Does Automattic/harper Support?

Harper is a pure-Rust grammar-checking engine that operates entirely in-memory without requiring any external database, though the repository includes optional Node.js database clients for example integrations.

Automattic/harper is an open-source grammar and spell-checking engine written in Rust. While the core linting functionality is completely database-agnostic, the repository contains optional database integrations for specific tooling and demonstration purposes. Understanding how Harper handles data storage helps developers integrate it correctly into their applications without unnecessary dependencies.

Core Architecture: No Database Required

Harper's primary engine, located in harper-core, functions as a pure-Rust grammar-checking engine that works entirely in-memory. The system does not persist grammar rules or dictionary data in external storage systems, making it highly portable and lightweight.

In-Memory Dictionary Implementation

The core spell-checking functionality relies on an in-memory dictionary compiled directly into the binary. In harper-core/src/spell/dictionary.rs, the dictionary implementation loads all word lists and rule data at runtime without database dependencies. This design ensures Harper remains self-contained and performsant across different deployment environments.

Optional Database Support in the Repository

While Harper's core engine requires no database, the repository includes several optional components that ship with database client libraries. These appear as optional runtime dependencies in pnpm-lock.yaml (lines approximately 6550‑6634) and serve tooling, CI, or integration demonstration purposes only.

SQLite Integration

The repository lists better-sqlite3, expo-sqlite, and sqlite3 in pnpm-lock.yaml under Node-side packages. These provide lightweight, file-based storage options for example scripts or test harnesses that need to persist text before processing it through Harper.

MySQL and PlanetScale Support

For server-side Node.js environments, the repository includes @planetscale/database. This MySQL-compatible client appears in integration examples demonstrating how Harper can be called from applications connected to PlanetScale or other MySQL databases.

PostgreSQL and Neon Support

The @neondatabase/serverless package in pnpm-lock.yaml enables example code to run against PostgreSQL-compatible serverless endpoints. As documented in packages/web/src/routes/docs/harperjs/configurerules/+page.md, storage decisions—including database selection—are explicitly left to the consumer of harper.js.

Database Vocabulary in Harper's Dictionary

Harper recognizes database-related terminology during spell-checking. The file harper-core/dictionary.dict contains comprehensive vocabulary including BadgerDB, CockroachDB, CouchDB, and RDBMS. This ensures Harper correctly handles technical database terms when linting software documentation or technical writing.

Practical Integration Examples

The following examples demonstrate how to use the optional database packages with Harper in a Node.js environment. These are standalone implementations for persisting text that you later submit to Harper for analysis.

SQLite Example with better-sqlite3

// Install: pnpm add better-sqlite3
const Database = require('better-sqlite3');
const db = new Database('example.db');

// Simple table for storing user-submitted text that Harper will lint
db.exec(`
  CREATE TABLE IF NOT EXISTS submissions (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    body TEXT NOT NULL
  );
`);

// Insert a new submission
const stmt = db.prepare('INSERT INTO submissions (body) VALUES (?)');
stmt.run('Harper provides great grammar suggestions.');

PlanetScale MySQL Example

// Install: pnpm add @planetscale/database
const { connect } = require('@planetscale/database');

const client = connect({
  host: 'aws.connect.psdb.io',
  username: 'my_user',
  password: 'my_password',
});

// Example query – fetch recent text entries
async function fetchTexts() {
  const rows = await client.execute('SELECT body FROM submissions ORDER BY id DESC LIMIT 10');
  console.log(rows.rows);
}

fetchTexts();

Neon PostgreSQL Example

// Install: pnpm add @neondatabase/serverless
const { Pool } = require('@neondatabase/serverless');

const pool = new Pool({
  connectionString: process.env.DATABASE_URL, // e.g. postgres://user:pass@…/dbname
});

// Store a new piece of text
async function storeText(text) {
  await pool.query('INSERT INTO submissions (body) VALUES ($1)', [text]);
}

storeText('Harper works fine with PostgreSQL!');

These snippets illustrate persistence layers only; they do not affect Harper's linting engine operation.

Storage in Harper Desktop

According to harper-desktop/README.md, the desktop application stores configuration as JSON files rather than using a database. This aligns with Harper's philosophy of minimal external dependencies and file-based configuration management.

Summary

  • Harper's core engine (harper-core) operates entirely in-memory without database dependencies
  • The repository includes optional SQLite, MySQL (PlanetScale), and PostgreSQL (Neon) clients for example integrations
  • Database terminology like CockroachDB and RDBMS appears in harper-core/dictionary.dict for spell-checking accuracy
  • Harper Desktop uses JSON files for configuration storage, as documented in the source repository
  • You can integrate Harper with any database type or run it completely database-free

Frequently Asked Questions

Does Harper require a database to function?

No. Harper's core grammar and spell-checking engine works entirely in-memory according to the source code in harper-core/src/spell/dictionary.rs. All dictionaries and rule data are compiled into the binary and loaded at runtime. You can deploy and run Harper without any database connection.

What database types can I use with Harper?

You can use any database type you prefer. The repository provides optional examples using SQLite (better-sqlite3), MySQL-compatible databases via PlanetScale (@planetscale/database), and PostgreSQL via Neon (@neondatabase/serverless). These appear in pnpm-lock.yaml as optional dependencies for demonstration purposes, not core requirements.

Why does Harper's dictionary include database terms?

The harper-core/dictionary.dict file contains technical vocabulary including BadgerDB, CockroachDB, CouchDB, and RDBMS. This allows Harper to recognize and correctly spell-check database-related terminology when analyzing technical documentation and software prose, ensuring accurate linting of technical content.

How does Harper Desktop store user data?

Harper Desktop stores configuration data as JSON files rather than in a database, as documented in harper-desktop/README.md. This file-based approach maintains consistency with Harper's lightweight, minimal-dependency architecture and avoids external storage requirements.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →