# How to Configure Tesseract for Offline and Air-Gapped Environments with LiteParse

> Configure Tesseract for offline and air-gapped systems using LiteParse. Learn to pre-download .traineddata files and set the tessdata path for secure OCR operations.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-06

---

**To run Tesseract OCR offline with LiteParse, pre-download the required `.traineddata` files and expose their local directory using the `--tessdata-path` CLI flag, the `TESSDATA_PREFIX` environment variable, or the equivalent `tessdata_path` configuration option in your code.**

LiteParse from the `run-llama/liteparse` repository bundles a built-in **Tesseract OCR engine** that normally downloads **language data** on first use. In **offline** or **air-gapped** environments without internet access, this automatic download fails, so you must provide **language files** locally and point LiteParse to them. This guide explains the exact resolution logic and configuration options based on the current source code.

## Why Offline Configuration Is Required

In [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs), LiteParse instantiates a `TesseractOcrEngine` when OCR is enabled and no external HTTP OCR server URL is supplied (line 156). The engine's `recognize` method in [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs) calls `ensure_traineddata` to obtain the required `<lang>.traineddata` file. If the file is absent, `ensure_traineddata` attempts an **HTTP download** from the official `tessdata_best` repository, which will fail without network connectivity and produce a clear error. To prevent this failure, you must place the **language data** on-disk before running LiteParse.

## How LiteParse Resolves Tesseract Data Directories

The `TesseractOcrEngine` resolves the `tessdata` directory through a strict **three-level fallback** defined in [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs) (lines 63-70):

1. **Explicit path** — the value passed via the `tessdata_path` constructor argument or configuration.
2. **Environment variable** — the **`TESSDATA_PREFIX`** environment variable if it is set.
3. **Default fallback** — a platform-specific directory:
   - Linux: `~/.tesseract-rs/tessdata`
   - macOS: `~/Library/Application Support/tesseract-rs/tessdata`
   - Windows: `%APPDATA%\tesseract-rs\tessdata`

If the required `.traineddata` file is not found in the resolved directory, LiteParse attempts a remote download. In an air-gapped deployment, you must ensure the file exists at the resolved path so the engine never reaches the network call.

## Preparing Local Tesseract Language Files

Before running LiteParse offline, download the necessary **language files** from the official Tesseract data repository at `https://github.com/tesseract-ocr/tessdata`. Place the files—such as `eng.traineddata` or `fra.traineddata`—into a single local directory that will serve as your **`tessdata` path**. Verify that the directory contains every **`.traineddata`** file your workload expects before starting the parser.

## Configuration Methods for Offline Use

You can expose the local **`tessdata` directory** to LiteParse through multiple interfaces. Choose the method that matches your deployment target.

### CLI and Environment Variable

The command-line interface accepts a **`--tessdata-path`** flag that maps directly to `LiteParseConfig.tessdata_path` in [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs) (line 13). Alternatively, set the **`TESSDATA_PREFIX`** environment variable.

```bash

# Using the environment variable

export TESSDATA_PREFIX=/opt/tessdata
lit parse sample.pdf --ocr-language eng

# Using the command-line flag

lit parse sample.pdf --tessdata-path /opt/tessdata --ocr-language eng

```

### Node.js Binding

When using the Node.js library, pass the `tessdataPath` option in the constructor configuration object:

```javascript
import { LiteParse } from 'liteparse';

const parser = new LiteParse({
  ocrEnabled: true,
  ocrLanguage: 'eng',
  tessdataPath: '/opt/tessdata',
});

await parser.parse('sample.pdf');

```

### Python Binding

In Python, set the `tessdata_path` attribute on `LiteParseConfig` before constructing the parser:

```python
from liteparse import LiteParse, LiteParseConfig

cfg = LiteParseConfig()
cfg.ocr_enabled = True
cfg.ocr_language = "eng"
cfg.tessdata_path = "/opt/tessdata"

parser = LiteParse(cfg)
result = parser.parse("sample.pdf")

```

### Rust Library

If you are using the Rust crate directly, you can manually instantiate a `TesseractOcrEngine` with a custom data directory. This is only necessary if you bypass the automatic configuration in `LiteParseConfig`:

```rust
use liteparse::ocr::tesseract::TesseractOcrEngine;
use std::sync::Arc;

let engine = Arc::new(TesseractOcrEngine::new(Some("/opt/tessdata".into())));
// Pass `engine` via `LiteParse::with_ocr_engine` if you want to override the default.

```

## Summary

- LiteParse's built-in Tesseract engine downloads language data automatically by default, which fails in air-gapped networks.
- The engine resolves `tessdata` via an explicit path, the `TESSDATA_PREFIX` environment variable, or a platform-specific default directory.
- To operate offline, pre-download the required `.traineddata` files from the official Tesseract repository and place them in a local directory.
- Expose that directory using `--tessdata-path`, `TESSDATA_PREFIX`, or the equivalent `tessdata_path` config field in Node.js, Python, or Rust.

## Frequently Asked Questions

### What happens if LiteParse cannot find the local Tesseract language data?

If the required `<lang>.traineddata` file is missing from the resolved directory, LiteParse calls `ensure_traineddata`, which attempts to download the file over HTTP. In an offline environment, this network request fails and produces a clear error, halting OCR processing.

### Can I use the same local tessdata directory across all LiteParse language bindings?

Yes. The `tessdata` directory is a standard Tesseract format. Once you populate a directory with the correct `.traineddata` files, you can reference it from the CLI, Node.js, Python, or Rust interfaces using the equivalent `tessdata_path` or `--tessdata-path` option.

### Does LiteParse require any other external resources when running Tesseract offline?

No. After the `.traineddata` files are available locally and the data directory is correctly configured via `TESSDATA_PREFIX` or `--tessdata-path`, the Tesseract engine inside LiteParse operates without any network access. No additional runtime downloads are required.

### Where does LiteParse look for language data if I do not set a custom path?

According to the source code in [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs), if no explicit path or environment variable is provided, LiteParse falls back to a platform-specific default directory. On Linux, this is `~/.tesseract-rs/tessdata`. On macOS, it uses `~/Library/Application Support/tesseract-rs/tessdata`, while Windows resolves `%APPDATA%\tesseract-rs\tessdata`.