# Does Firecrawl pdf-inspector Support PDF Forms? A Complete Technical Guide

> Discover if Firecrawl pdf-inspector supports PDF forms. Learn how it extracts AcroForm field values into TextItems for seamless markdown conversion.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Yes, Firecrawl pdf-inspector fully supports PDF forms by extracting interactive AcroForm field values and converting them into positioned TextItems that flow through the standard markdown conversion pipeline.**

The `firecrawl/pdf-inspector` Rust library treats PDF form data as first-class citizens. Whether you're processing tax documents, applications, or contracts with fillable fields, the tool automatically detects, extracts, and formats form content alongside regular text.

## How PDF Form Extraction Works in pdf-inspector

The form handling pipeline operates through three coordinated components that traverse the PDF's AcroForm hierarchy and integrate results into the final output.

### The Core Extraction Function: `extract_form_fields`

In [`src/extractor/links.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/links.rs), the `extract_form_fields` function (lines 121-127) does the heavy lifting. It walks the `/AcroForm → /Fields` dictionary structure, retrieves each widget's bounding rectangle, and captures the current field value:

```rust
// From src/extractor/links.rs (lines 121-127)
// Returns Vec<TextItem> where each form field becomes a positioned text element

```

For each field encountered, the function constructs a `TextItem` with `item_type` set to `ItemType::FormField(name, value)`. This design choice preserves spatial information—form fields maintain their page position just like regular text, ensuring proper reading order in the final output.

### Integration with the Main Extraction Flow

The [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) file orchestrates the merge. Lines 25-29 call `extract_form_fields` and concatenate the results with standard text extraction:

```rust
// Pseudocode representation of src/extractor/mod.rs lines 25-29
let mut items = extract_text_elements(&doc, page)?;
let form_items = extract_form_fields(&doc, page)?;
items.extend(form_items);

```

This unified approach means **form fields require no special post-processing**—they participate in the same sorting, deduplication, and formatting logic as extracted text.

## Using PDF Form Support: Three Methods

### Method 1: Command-Line with `pdf2md`

The `pdf2md` binary automatically includes form fields in both output formats:

```bash

# Standard markdown output (form fields appear as plain text)

pdf2md my_form.pdf

# Structured JSON with explicit form field typing

pdf2md my_form.pdf --json

```

JSON output includes a `type: "form_field"` discriminator, making downstream parsing straightforward.

### Method 2: Programmatic Rust API

For native Rust integration, use `process_pdf_with_options` with default settings (form extraction is enabled by default):

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::process_mode::ProcessOptions;

// Form extraction enabled by default
let opts = ProcessOptions::default();
let result = process_pdf_with_options("my_form.pdf", opts)?;

// Inspect form field items
for itm in result.items {
    if let ItemType::FormField(name, value) = itm.item_type {
        println!("Form field \"{name}\" = \"{value}\"");
    }
}

```

The `ItemType::FormField(name, value)` variant is defined in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs), providing structured access to both the field's **name** (dictionary key) and **value** (current content).

### Method 3: Python Bindings

Python users access the same functionality through the N-API wrapper exposed in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs):

```python
from pdf_inspector import PDFInspector

inspector = PDFInspector()
doc = inspector.process_pdf("my_form.pdf")  # Returns dict with "items" list

for item in doc["items"]:
    if item["type"] == "form_field":
        print(f'{item["name"]}: {item["value"]}')

```

The Python interface normalizes Rust enums into string discriminators and flattens nested structures for idiomatic dictionary access.

## What Form Types Are Supported?

pdf-inspector handles **AcroForm** fields—the standard PDF specification for interactive forms. This includes:

- **Text fields** (single-line and multi-line)
- **Checkboxes and radio buttons** (extracted as selected values)
- **Dropdown menus and list boxes** (current selection captured)
- **Signature fields** (field presence noted; signature validation not performed)

Form field **appearance streams** are not rendered. The tool extracts the **underlying value**, not the visual representation. For signatures, this means you receive the field name and placeholder status, not a cryptographic validation or image bitmap.

## Key Source Files for Form Handling

| File | Responsibility |
|------|---------------|
| [`src/extractor/links.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/links.rs) | `extract_form_fields` implementation; AcroForm hierarchy traversal |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Entry point that merges form items with text extraction results |
| [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) | `ItemType::FormField` and `TextItem` structure definitions |
| [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | CLI interface passing form extraction to underlying engine |
| [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) | Python N-API bindings exposing form support |

## Limitations and Edge Cases

- **XFA forms** (Adobe's XML Forms Architecture) are not supported—only legacy AcroForm structures
- **JavaScript-driven form calculations** execute at PDF runtime; pdf-inspector captures the **stored value**, not dynamically computed results
- **Field formatting** (number masks, date formats) is not applied—raw values are returned

Positioning accuracy depends on the PDF's `/Rect` entries. Malformed form dictionaries may produce misaligned output.

## Summary

- **Firecrawl pdf-inspector supports PDF forms through automatic AcroForm extraction** enabled by default
- The `extract_form_fields` function in [`src/extractor/links.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/links.rs) walks field hierarchies and converts entries to positioned `TextItem`s
- Form data flows through standard markdown/JSON pipelines without special configuration
- Rust, CLI, and Python interfaces all expose complete form functionality
- Only field **values** and **positions** are extracted—visual styling and JavaScript behavior are not captured

## Frequently Asked Questions

### Does pdf-inspector require manual configuration to extract PDF forms?

No. Form extraction is **enabled by default** in all processing modes. The `ProcessOptions::default()` constructor in Rust, the bare `pdf2md` invocation in CLI, and `PDFInspector.process_pdf()` in Python all automatically include form fields in output.

### Can I distinguish form fields from regular text in the output?

Yes. In JSON mode (`--json` flag or Python dict output), each item has a `type` field set to `"form_field"` for interactive elements versus `"text"` for standard content. The Rust API exposes this through pattern matching on `ItemType::FormField(name, value)`.

### Does pdf-inspector extract form field labels or only values?

Only **values** and **internal field names**. PDFs do not reliably encode human-readable labels in the AcroForm structure—labels are typically static text positioned near fields. The tool captures this static text as regular text items, leaving correlation to downstream processing.

### Is digital signature field content preserved?

Signature fields appear as `FormField` entries with their **name** and **current value** (often a placeholder string). The **cryptographic signature blob** and **visual appearance** are not extracted—pdf-inspector treats signatures as opaque form values without validation logic.