How to Debug LLM Extraction Failures and JSON Parsing Errors in the Hiring Agent
To debug LLM extraction failures in the interviewstreet/hiring-agent repository, inspect the raw LLM response before sanitization, test the extract_json_from_response utility in isolation to verify markdown stripping, and implement defensive error handling around json.loads to capture malformed payloads with full context.
The hiring agent processes résumé PDFs by sending section-specific prompts to a Large Language Model and parsing the returned JSON. When the LLM wraps output in markdown fences, truncates responses, or returns explanatory text instead of valid JSON, the extraction pipeline throws decoding errors. Understanding the exact flow from pdf.py through llm_utils.py allows you to pinpoint whether failures stem from the model output, the sanitization logic, or schema validation.
Understanding the Extraction Pipeline
The extraction workflow is deliberately split into three discrete stages to enable independent debugging. The primary orchestration happens in pdf.py, which calls the LLM and delegates JSON cleaning to shared utilities.
Core Modules and Entry Points
pdf.py– Containsextract_json_from_pdf(lines 111‑116), which orchestrates the LLM call and invokes the sanitizer. Error logging occurs at line 127.llm_utils.py– Housesextract_json_from_response(lines 13‑33), the critical routine that strips markdown fences and returns a pure JSON string.evaluator.py– Mirrors the same pattern inevaluate_resume(lines 25‑48), using the same utilities to request structured résumé evaluations.
This architecture isolates each transformation—raw text, sanitized string, and parsed object—so you can log or unit-test them separately.
Typical Failure Flow
When extraction fails, the error propagates through this sequence:
extract_json_from_pdf → calls LLM → populates response_text
→ extract_json_from_response → strips fences → returns cleaned string
→ json.loads → raises JSONDecodeError → caught in pdf.py
The log line surfacing the failure resides in pdf.py at line 127, displaying "❌ Error during PDF to JSON extraction" when the final parsing step fails.
Common Causes of LLM JSON Extraction Failures
Markdown Code Block Wrappers
The LLM frequently wraps JSON inside markdown fences (```json … ```) or adds conversational text before the object. If extract_json_from_response fails to strip these wrappers using its regex logic (lines 31‑33 in llm_utils.py), json.loads immediately throws a JSONDecodeError.
Truncated or Partial Responses
Large PDFs or strict token limits can cause the LLM to return a response cut off mid-object. The incomplete string passes through the sanitizer but fails during parsing because braces are unbalanced or values are missing.
Prompt Template Mismatches
Prompt templates live in prompts/templates/ and are managed by template_manager.py. If a template references a section name that does not match the PDF’s actual headings, the LLM may return a plain text message like "I couldn't find the information" instead of valid JSON, causing the parser to fail.
Step-by-Step Debugging Techniques
Inspect the Raw LLM Response
Before any cleaning occurs, capture the exact string returned by the model. The _call_llm_for_method method in pdf.py is deliberately exposed for this purpose, returning the raw output before post-processing.
from hiring_agent.pdf import PDFHandler
handler = PDFHandler()
raw = handler._call_llm_for_method(
resume_text="John Doe …",
prompt="Extract the work‑experience section as JSON."
)
print("=== RAW LLM RESPONSE ===")
print(raw)
Examining this output reveals whether the model added explanatory prose, markdown fences, or truncated the JSON object.
Test the JSON Sanitizer in Isolation
Once you have a problematic response, test extract_json_from_response independently to verify the regex cleaning logic.
from hiring_agent.llm_utils import extract_json_from_response
bad_reply = """Here is the result:\n```json
{
"work": [
{"company": "Acme", "title": "Engineer"}
]
```"""
clean_json = extract_json_from_response(bad_reply)
print("=== CLEAN JSON ===")
print(clean_json)
If clean_json still contains stray characters, inspect the regex implementation in llm_utils.py lines 31‑33 and adjust the pattern to handle edge cases like missing closing backticks.
Implement Defensive Error Handling
Wrap the parsing logic to capture the full context when json.loads fails. Replace the existing parsing block in pdf.py with a helper that logs the original and cleaned strings.
import json
from hiring_agent.llm_utils import extract_json_from_response
def safe_parse(reply: str):
try:
clean = extract_json_from_response(reply)
return json.loads(clean)
except json.JSONDecodeError as exc:
print("❌ JSON parsing failed:")
print("Original reply:", reply)
print("After cleaning:", clean)
print("Error details:", exc)
raise
# Usage after LLM call
parsed = safe_parse(raw_response)
This pattern preserves the state of the data at each transformation stage, making it trivial to identify whether the LLM or the sanitizer introduced the error.
Validate Against the Pydantic Schema
After successful parsing, verify that the dictionary conforms to the expected structure using the JSONResume model defined in models.py.
from hiring_agent.models import JSONResume
def validate_resume(data: dict) -> JSONResume:
try:
return JSONResume(**data)
except Exception as e:
print("⚠️ Validation error:", e)
raise
# After json.loads(...)
resume_obj = validate_resume(parsed)
Running this validation in a test harness quickly reveals field name mismatches or missing required keys that the LLM omitted, distinguishing between JSON syntax errors and schema compliance failures.
Key Source Files for Troubleshooting
| File | Debugging Relevance |
|---|---|
pdf.py |
Core extraction logic, LLM invocation, and the primary error handler at line 127. |
llm_utils.py |
Contains the extract_json_from_response sanitizer; inspect here when markdown fences persist. |
evaluator.py |
Secondary usage site for the extraction utilities; compare behaviors when pdf.py behaves unexpectedly. |
github.py |
Also uses extract_json_from_response; useful for verifying consistency across different agent modules. |
prompts/template_manager.py |
Holds Jinja2 templates; errors here often manifest as LLM responses lacking JSON structure. |
models.py |
Defines the JSONResume Pydantic schema; reference this to confirm expected field names and types. |
Summary
- Trace the data flow through
pdf.py→llm_utils.py→json.loadsto isolate which stage corrupts the payload. - Capture raw LLM output using
_call_llm_for_methodbefore any sanitization occurs. - Test
extract_json_from_responseindependently to verify regex patterns correctly strip markdown fences and trailing text. - Wrap parsers in try-except blocks that log both the original and cleaned strings to diagnose
JSONDecodeErrororigins. - Validate outputs against
JSONResumeinmodels.pyto distinguish between malformed JSON and schema mismatches.
Frequently Asked Questions
Why does json.loads fail even when the LLM returns what looks like valid JSON?
The LLM often embeds JSON inside markdown code blocks (```json) or adds conversational text before or after the object. If extract_json_from_response in llm_utils.py fails to strip these wrappers completely, the resulting string contains invalid characters that cause json.loads to raise a JSONDecodeError. Always inspect the raw response to confirm whether sanitization is working correctly.
How can I test the extraction logic without processing a full PDF?
Import PDFHandler and call the private _call_llm_for_method with a hardcoded résumé text string and prompt. This method returns the raw LLM response without triggering the full pipeline, allowing you to iterate on prompt engineering or test the sanitizer without file I/O overhead.
What causes the "Error during PDF to JSON extraction" log message?
This message originates from the exception handler in pdf.py at line 127. It fires when extract_json_from_response returns a string that json.loads cannot parse, typically due to truncated responses from large PDFs, malformed markdown fences, or template mismatches that cause the LLM to return plain text instead of JSON.
Where should I add logging to catch extraction errors early?
Insert logging immediately after the LLM call in pdf.py to capture response_text, and again inside extract_json_from_response in llm_utils.py to log the cleaned string before it reaches json.loads. Logging at these two points provides a clear before-and-after comparison when debugging sanitization failures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →