Troubleshooting PDF Extraction Failures with Hiring Agent: 6 Common Issues and Fixes

Enable debug logging and verify PyMuPDF compatibility, LLM credentials, and template integrity to resolve most PDF extraction failures in the Hiring Agent pipeline.

Hiring Agent converts résumé PDFs into structured JSON by chaining three core components: PyMuPDF-based text extraction, section-wise LLM parsing, and Pydantic model validation. When any step fails, PDFHandler returns None and logs specific diagnostic messages that pinpoint the root cause. Understanding these failure points allows you to quickly restore extraction functionality without digging through the entire pdf.py orchestration logic.

Understanding the Extraction Pipeline

The conversion process flows through distinct stages defined in pdf.py. First, PDFHandler.extract_text_from_pdf opens the file using PyMuPDF and invokes to_markdown from pymupdf_rag.py to generate markdown-style text. This content then passes through section-specific LLM prompts initialized via initialize_llm_provider, with responses cleaned by extract_json_from_response and transformed by transform_parsed_data. Finally, the dictionaries convert to Pydantic models (Basics, Work, etc.) and assemble into a JSONResume object.

When failures occur, they typically surface in one of six specific areas: file access, markdown conversion, LLM configuration, template rendering, JSON decoding, or model validation.

Most file-level errors occur during the initial extraction phase in pdf.py lines 48-54, where the system logs both exceptions and character counts of extracted text.

FileNotFoundError at line 49

This error indicates the path passed to extract_json_from_pdf does not resolve to an existing file. Verify whether you provided an absolute or relative path from the correct working directory. Ensure the file exists and the process has read permissions before invoking the handler.

Empty String Returns

When extract_text_from_pdf returns an empty string, the PDF is likely corrupted, password-protected, or contains only scanned images without text layers. Open the file in a standard PDF viewer to verify integrity. If protected, remove the password or re-save the document through a PDF repair tool before reprocessing.

PyMuPDF Library Exceptions

Incompatible PyMuPDF versions or missing system dependencies (like libxml2) cause exceptions during file opening. Check your installed version with pip show pymupdf and upgrade to the latest stable release using pip install -U pymupdf. This resolves most compatibility issues with newer PDF specifications.

Markdown Conversion Problems

The to_markdown function in pymupdf_rag.py (called at lines 54-57) can fail when processing PDFs containing unusual fonts, complex layouts, or embedded images that confuse the layout algorithm.

If extraction produces garbled text or raises layout errors, insert temporary debug logging immediately after the to_markdown call to inspect raw markdown output. For documents with heavy OCR requirements, pre-process the PDF using Tesseract or update pymupdf_rag.py to a newer implementation that better handles OCR layers before feeding content to the Hiring Agent pipeline.

LLM Provider Misconfiguration

The LLM provider initializes once during PDFHandler.__init__ via initialize_llm_provider(DEFAULT_MODEL) at lines 43-46. Failures here typically stem from missing credentials or invalid model references.

Missing API Keys

The default configuration expects a GEMINI_API_KEY environment variable. Verify your environment setup:

echo $GEMINI_API_KEY
python -c "import os; print(os.getenv('GEMINI_API_KEY'))"

If empty, export the key or add it to a .env file following the repository's .env.example pattern.

Invalid Model Parameters

Ensure the model name exists in MODEL_PARAMETERS. The provider call wraps inside _call_llm_for_section, where exceptions bubble up and log at lines 33-34. Check these log entries specifically for authentication or model access errors.

Prompt Rendering Errors

Each résumé section relies on Jinja-style templates loaded by TemplateManager from the prompts/templates/ directory. If render_template returns None due to missing files or syntax errors, the LLM call skips and logs "❌ Failed to render … template" at lines 80-86.

Verify template integrity by running a quick sanity check:

from prompts.template_manager import TemplateManager

tm = TemplateManager()
print(tm.render_template("basics", text_content="dummy"))

If this prints None, reinstall the repository or restore the missing template files to ensure the prompts/templates/ directory matches the expected structure.

JSON Decoding Failures

After the LLM responds, extract_json_from_response cleans the text and slices between the first { and last }. If the model outputs malformed JSON, json.loads raises JSONDecodeError caught at lines 27-30, causing the section to discard.

Enable verbose logging to inspect the raw response_text printed in the catch block:

import logging
logging.basicConfig(level=logging.DEBUG)

If the LLM consistently misformats output, adjust the system_message template in your prompt configuration to explicitly demand valid JSON with enforced schema constraints.

Pydantic Model Construction Errors

Even valid JSON can fail when converted to JSONResume if required fields are missing. The error logs at lines 21-22 in pdf.py indicate validation failures during the final assembly stage.

Verify that parsed dictionaries contain mandatory keys: basics, work, education, and other required sections defined in models.py. If certain sections are optional, ensure the corresponding Pydantic fields use Optional typing to prevent validation errors during the JSONResume object construction at lines 71-77 and 111-119.

Debugging with Logging

Hiring Agent uses Python's standard logging module with line-specific debug statements. Configure debug-level logging to surface exact failure points:

import logging
logging.basicConfig(level=logging.DEBUG)

This reveals the specific line numbers in pdf.py where extraction halts, making it trivial to map errors to the troubleshooting sections above.

Complete Working Example

The following snippet demonstrates proper error handling and debugging techniques:

import logging
from pdf import PDFHandler

# Enable detailed logging to surface failure points

logging.basicConfig(level=logging.DEBUG)

handler = PDFHandler()

# Successful extraction

resume = handler.extract_json_from_pdf("samples/resume.pdf")
if resume:
    print("✅ Extraction succeeded")
else:
    print("❌ Extraction failed – check logs for specific line references")

# Inspecting raw markdown for conversion issues

raw_text = handler.extract_text_from_pdf("samples/resume.pdf")
print("--- Raw markdown preview ---")
print(raw_text[:500])

Summary

  • File issues manifest as FileNotFoundError or empty strings at pdf.py lines 48-54; verify paths and PDF integrity.
  • Markdown conversion fails on complex layouts in pymupdf_rag.py; debug by inspecting raw output.
  • LLM configuration errors occur at initialization (lines 43-46); confirm GEMINI_API_KEY and model names.
  • Template errors return None from render_template (lines 79-82); ensure prompts/templates/ exists.
  • JSON decoding fails at lines 27-30 when LLM output is malformed; fix via prompt engineering or response cleaning.
  • Pydantic validation errors log at lines 21-22; ensure all required fields in models.py are present or marked Optional.

Frequently Asked Questions

Why does Hiring Agent return None instead of throwing an exception?

The PDFHandler deliberately catches exceptions at each pipeline stage—file opening, markdown conversion, LLM calls, and JSON parsing—to allow graceful degradation. It logs the specific error with line numbers in pdf.py (lines 21-22, 27-30, 33-34) and returns None to signal failure without crashing the calling process. Check the logs immediately after the None return to identify which component failed.

How do I fix "JSONDecodeError" when parsing LLM responses?

This error occurs in extract_json_from_response when the LLM returns text outside valid JSON boundaries or malformed syntax. The code attempts to slice between the first { and last }, but inconsistent model output breaks this logic. Enable logging.DEBUG to view the raw response_text at lines 28-30, then modify the system prompt in your template to explicitly request JSON-only output without markdown code blocks.

What should I check if PyMuPDF raises exceptions during extraction?

First verify the PyMuPDF version compatibility with your system libraries using pip show pymupdf. Upgrade to the latest version with pip install -U pymupdf to resolve most font rendering and layout engine issues. If errors persist, the PDF may contain corrupted streams or require OCR preprocessing—attempt to re-save the file through a standard PDF reader to normalize the structure before processing through extract_text_from_pdf.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →