Does Firecrawl pdf-inspector Support PDF Forms? A Complete Technical Guide
Yes, Firecrawl pdf-inspector fully supports PDF forms by extracting interactive AcroForm field values and converting them into positioned TextItems that flow through the standard markdown conversion pipeline.
The firecrawl/pdf-inspector Rust library treats PDF form data as first-class citizens. Whether you're processing tax documents, applications, or contracts with fillable fields, the tool automatically detects, extracts, and formats form content alongside regular text.
How PDF Form Extraction Works in pdf-inspector
The form handling pipeline operates through three coordinated components that traverse the PDF's AcroForm hierarchy and integrate results into the final output.
The Core Extraction Function: extract_form_fields
In src/extractor/links.rs, the extract_form_fields function (lines 121-127) does the heavy lifting. It walks the /AcroForm → /Fields dictionary structure, retrieves each widget's bounding rectangle, and captures the current field value:
// From src/extractor/links.rs (lines 121-127)
// Returns Vec<TextItem> where each form field becomes a positioned text element
For each field encountered, the function constructs a TextItem with item_type set to ItemType::FormField(name, value). This design choice preserves spatial information—form fields maintain their page position just like regular text, ensuring proper reading order in the final output.
Integration with the Main Extraction Flow
The src/extractor/mod.rs file orchestrates the merge. Lines 25-29 call extract_form_fields and concatenate the results with standard text extraction:
// Pseudocode representation of src/extractor/mod.rs lines 25-29
let mut items = extract_text_elements(&doc, page)?;
let form_items = extract_form_fields(&doc, page)?;
items.extend(form_items);
This unified approach means form fields require no special post-processing—they participate in the same sorting, deduplication, and formatting logic as extracted text.
Using PDF Form Support: Three Methods
Method 1: Command-Line with pdf2md
The pdf2md binary automatically includes form fields in both output formats:
# Standard markdown output (form fields appear as plain text)
pdf2md my_form.pdf
# Structured JSON with explicit form field typing
pdf2md my_form.pdf --json
JSON output includes a type: "form_field" discriminator, making downstream parsing straightforward.
Method 2: Programmatic Rust API
For native Rust integration, use process_pdf_with_options with default settings (form extraction is enabled by default):
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::process_mode::ProcessOptions;
// Form extraction enabled by default
let opts = ProcessOptions::default();
let result = process_pdf_with_options("my_form.pdf", opts)?;
// Inspect form field items
for itm in result.items {
if let ItemType::FormField(name, value) = itm.item_type {
println!("Form field \"{name}\" = \"{value}\"");
}
}
The ItemType::FormField(name, value) variant is defined in src/types.rs, providing structured access to both the field's name (dictionary key) and value (current content).
Method 3: Python Bindings
Python users access the same functionality through the N-API wrapper exposed in src/python.rs:
from pdf_inspector import PDFInspector
inspector = PDFInspector()
doc = inspector.process_pdf("my_form.pdf") # Returns dict with "items" list
for item in doc["items"]:
if item["type"] == "form_field":
print(f'{item["name"]}: {item["value"]}')
The Python interface normalizes Rust enums into string discriminators and flattens nested structures for idiomatic dictionary access.
What Form Types Are Supported?
pdf-inspector handles AcroForm fields—the standard PDF specification for interactive forms. This includes:
- Text fields (single-line and multi-line)
- Checkboxes and radio buttons (extracted as selected values)
- Dropdown menus and list boxes (current selection captured)
- Signature fields (field presence noted; signature validation not performed)
Form field appearance streams are not rendered. The tool extracts the underlying value, not the visual representation. For signatures, this means you receive the field name and placeholder status, not a cryptographic validation or image bitmap.
Key Source Files for Form Handling
| File | Responsibility |
|---|---|
src/extractor/links.rs |
extract_form_fields implementation; AcroForm hierarchy traversal |
src/extractor/mod.rs |
Entry point that merges form items with text extraction results |
src/types.rs |
ItemType::FormField and TextItem structure definitions |
src/bin/pdf2md.rs |
CLI interface passing form extraction to underlying engine |
src/python.rs |
Python N-API bindings exposing form support |
Limitations and Edge Cases
- XFA forms (Adobe's XML Forms Architecture) are not supported—only legacy AcroForm structures
- JavaScript-driven form calculations execute at PDF runtime; pdf-inspector captures the stored value, not dynamically computed results
- Field formatting (number masks, date formats) is not applied—raw values are returned
Positioning accuracy depends on the PDF's /Rect entries. Malformed form dictionaries may produce misaligned output.
Summary
- Firecrawl pdf-inspector supports PDF forms through automatic AcroForm extraction enabled by default
- The
extract_form_fieldsfunction insrc/extractor/links.rswalks field hierarchies and converts entries to positionedTextItems - Form data flows through standard markdown/JSON pipelines without special configuration
- Rust, CLI, and Python interfaces all expose complete form functionality
- Only field values and positions are extracted—visual styling and JavaScript behavior are not captured
Frequently Asked Questions
Does pdf-inspector require manual configuration to extract PDF forms?
No. Form extraction is enabled by default in all processing modes. The ProcessOptions::default() constructor in Rust, the bare pdf2md invocation in CLI, and PDFInspector.process_pdf() in Python all automatically include form fields in output.
Can I distinguish form fields from regular text in the output?
Yes. In JSON mode (--json flag or Python dict output), each item has a type field set to "form_field" for interactive elements versus "text" for standard content. The Rust API exposes this through pattern matching on ItemType::FormField(name, value).
Does pdf-inspector extract form field labels or only values?
Only values and internal field names. PDFs do not reliably encode human-readable labels in the AcroForm structure—labels are typically static text positioned near fields. The tool captures this static text as regular text items, leaving correlation to downstream processing.
Is digital signature field content preserved?
Signature fields appear as FormField entries with their name and current value (often a placeholder string). The cryptographic signature blob and visual appearance are not extracted—pdf-inspector treats signatures as opaque form values without validation logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →