# Troubleshooting PDF Extraction Failures with Hiring Agent: 6 Common Issues and Fixes

> Troubleshoot PDF extraction failures in Hiring Agent. Learn 6 common issues and quick fixes including debug logging, PyMuPDF compatibility, LLM credentials and template checks.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-01

---

**Enable debug logging and verify PyMuPDF compatibility, LLM credentials, and template integrity to resolve most PDF extraction failures in the Hiring Agent pipeline.**

Hiring Agent converts résumé PDFs into structured JSON by chaining three core components: PyMuPDF-based text extraction, section-wise LLM parsing, and Pydantic model validation. When any step fails, `PDFHandler` returns `None` and logs specific diagnostic messages that pinpoint the root cause. Understanding these failure points allows you to quickly restore extraction functionality without digging through the entire [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) orchestration logic.

## Understanding the Extraction Pipeline

The conversion process flows through distinct stages defined in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py). First, `PDFHandler.extract_text_from_pdf` opens the file using PyMuPDF and invokes `to_markdown` from [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) to generate markdown-style text. This content then passes through section-specific LLM prompts initialized via `initialize_llm_provider`, with responses cleaned by `extract_json_from_response` and transformed by `transform_parsed_data`. Finally, the dictionaries convert to Pydantic models (`Basics`, `Work`, etc.) and assemble into a `JSONResume` object.

When failures occur, they typically surface in one of six specific areas: file access, markdown conversion, LLM configuration, template rendering, JSON decoding, or model validation.

## File-Related Issues

Most file-level errors occur during the initial extraction phase in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) lines 48-54, where the system logs both exceptions and character counts of extracted text.

**FileNotFoundError at line 49**

This error indicates the path passed to `extract_json_from_pdf` does not resolve to an existing file. Verify whether you provided an absolute or relative path from the correct working directory. Ensure the file exists and the process has read permissions before invoking the handler.

**Empty String Returns**

When `extract_text_from_pdf` returns an empty string, the PDF is likely corrupted, password-protected, or contains only scanned images without text layers. Open the file in a standard PDF viewer to verify integrity. If protected, remove the password or re-save the document through a PDF repair tool before reprocessing.

**PyMuPDF Library Exceptions**

Incompatible PyMuPDF versions or missing system dependencies (like `libxml2`) cause exceptions during file opening. Check your installed version with `pip show pymupdf` and upgrade to the latest stable release using `pip install -U pymupdf`. This resolves most compatibility issues with newer PDF specifications.

## Markdown Conversion Problems

The `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (called at lines 54-57) can fail when processing PDFs containing unusual fonts, complex layouts, or embedded images that confuse the layout algorithm.

If extraction produces garbled text or raises layout errors, insert temporary debug logging immediately after the `to_markdown` call to inspect raw markdown output. For documents with heavy OCR requirements, pre-process the PDF using Tesseract or update [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) to a newer implementation that better handles OCR layers before feeding content to the Hiring Agent pipeline.

## LLM Provider Misconfiguration

The LLM provider initializes once during `PDFHandler.__init__` via `initialize_llm_provider(DEFAULT_MODEL)` at lines 43-46. Failures here typically stem from missing credentials or invalid model references.

**Missing API Keys**

The default configuration expects a `GEMINI_API_KEY` environment variable. Verify your environment setup:

```bash
echo $GEMINI_API_KEY
python -c "import os; print(os.getenv('GEMINI_API_KEY'))"

```

If empty, export the key or add it to a `.env` file following the repository's `.env.example` pattern.

**Invalid Model Parameters**

Ensure the model name exists in `MODEL_PARAMETERS`. The provider call wraps inside `_call_llm_for_section`, where exceptions bubble up and log at lines 33-34. Check these log entries specifically for authentication or model access errors.

## Prompt Rendering Errors

Each résumé section relies on Jinja-style templates loaded by `TemplateManager` from the `prompts/templates/` directory. If `render_template` returns `None` due to missing files or syntax errors, the LLM call skips and logs "❌ Failed to render … template" at lines 80-86.

Verify template integrity by running a quick sanity check:

```python
from prompts.template_manager import TemplateManager

tm = TemplateManager()
print(tm.render_template("basics", text_content="dummy"))

```

If this prints `None`, reinstall the repository or restore the missing template files to ensure the `prompts/templates/` directory matches the expected structure.

## JSON Decoding Failures

After the LLM responds, `extract_json_from_response` cleans the text and slices between the first `{` and last `}`. If the model outputs malformed JSON, `json.loads` raises `JSONDecodeError` caught at lines 27-30, causing the section to discard.

Enable verbose logging to inspect the raw `response_text` printed in the catch block:

```python
import logging
logging.basicConfig(level=logging.DEBUG)

```

If the LLM consistently misformats output, adjust the `system_message` template in your prompt configuration to explicitly demand valid JSON with enforced schema constraints.

## Pydantic Model Construction Errors

Even valid JSON can fail when converted to `JSONResume` if required fields are missing. The error logs at lines 21-22 in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) indicate validation failures during the final assembly stage.

Verify that parsed dictionaries contain mandatory keys: `basics`, `work`, `education`, and other required sections defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py). If certain sections are optional, ensure the corresponding Pydantic fields use `Optional` typing to prevent validation errors during the `JSONResume` object construction at lines 71-77 and 111-119.

## Debugging with Logging

Hiring Agent uses Python's standard `logging` module with line-specific debug statements. Configure debug-level logging to surface exact failure points:

```python
import logging
logging.basicConfig(level=logging.DEBUG)

```

This reveals the specific line numbers in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) where extraction halts, making it trivial to map errors to the troubleshooting sections above.

## Complete Working Example

The following snippet demonstrates proper error handling and debugging techniques:

```python
import logging
from pdf import PDFHandler

# Enable detailed logging to surface failure points

logging.basicConfig(level=logging.DEBUG)

handler = PDFHandler()

# Successful extraction

resume = handler.extract_json_from_pdf("samples/resume.pdf")
if resume:
    print("✅ Extraction succeeded")
else:
    print("❌ Extraction failed – check logs for specific line references")

# Inspecting raw markdown for conversion issues

raw_text = handler.extract_text_from_pdf("samples/resume.pdf")
print("--- Raw markdown preview ---")
print(raw_text[:500])

```

## Summary

- **File issues** manifest as `FileNotFoundError` or empty strings at [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) lines 48-54; verify paths and PDF integrity.
- **Markdown conversion** fails on complex layouts in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py); debug by inspecting raw output.
- **LLM configuration** errors occur at initialization (lines 43-46); confirm `GEMINI_API_KEY` and model names.
- **Template errors** return `None` from `render_template` (lines 79-82); ensure `prompts/templates/` exists.
- **JSON decoding** fails at lines 27-30 when LLM output is malformed; fix via prompt engineering or response cleaning.
- **Pydantic validation** errors log at lines 21-22; ensure all required fields in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) are present or marked Optional.

## Frequently Asked Questions

### Why does Hiring Agent return None instead of throwing an exception?

The `PDFHandler` deliberately catches exceptions at each pipeline stage—file opening, markdown conversion, LLM calls, and JSON parsing—to allow graceful degradation. It logs the specific error with line numbers in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 21-22, 27-30, 33-34) and returns `None` to signal failure without crashing the calling process. Check the logs immediately after the `None` return to identify which component failed.

### How do I fix "JSONDecodeError" when parsing LLM responses?

This error occurs in `extract_json_from_response` when the LLM returns text outside valid JSON boundaries or malformed syntax. The code attempts to slice between the first `{` and last `}`, but inconsistent model output breaks this logic. Enable `logging.DEBUG` to view the raw `response_text` at lines 28-30, then modify the system prompt in your template to explicitly request JSON-only output without markdown code blocks.

### What should I check if PyMuPDF raises exceptions during extraction?

First verify the PyMuPDF version compatibility with your system libraries using `pip show pymupdf`. Upgrade to the latest version with `pip install -U pymupdf` to resolve most font rendering and layout engine issues. If errors persist, the PDF may contain corrupted streams or require OCR preprocessing—attempt to re-save the file through a standard PDF reader to normalize the structure before processing through `extract_text_from_pdf`.