How to Debug and Inspect LLM Responses During the Resume Evaluation Process
You can debug LLM responses in the hiring-agent repository by intercepting calls in the ResumeEvaluator class, logging outputs from llm_utils.py, and validating JSON extraction before schema parsing.
The resume evaluation pipeline in the interviewstreet/hiring-agent repository relies on structured LLM outputs to score candidates against specific criteria. When responses fail schema validation or return unexpected formats, developers need visibility into the prompt construction, raw API responses, and JSON parsing stages to diagnose issues effectively.
Understanding the Resume Evaluation Architecture
The evaluation flow centers on the ResumeEvaluator class defined in evaluator.py. When evaluate_resume() is invoked, it orchestrates three critical phases: prompt rendering, LLM invocation, and structured data extraction.
The pipeline first renders evaluation criteria from prompts/templates/resume_evaluation_criteria.jinja and system instructions from prompts/templates/resume_evaluation_system_message.jinja. These templates combine with candidate data to form the complete message payload sent to the LLM provider.
Inspecting Prompt Construction Before Sending
To debug issues with model comprehension or instruction following, inspect the rendered prompts immediately before the LLM call. The evaluate_resume() method builds the final prompt string using Jinja2 template rendering.
Add logging or breakpoints immediately after template rendering to capture the exact text sent to the provider:
# In evaluator.py, inside evaluate_resume()
from pathlib import Path
# Log rendered templates for inspection
criteria_template = Path("prompts/templates/resume_evaluation_criteria.jinja").read_text()
system_template = Path("prompts/templates/resume_evaluation_system_message.jinja").read_text()
print(f"System prompt length: {len(system_template)}")
print(f"Criteria prompt: {criteria_template[:500]}...") # Truncate for readability
Monitoring Raw LLM Responses
The concrete provider implementation—either OllamaProvider or GeminiProvider—is instantiated by initialize_llm_provider() in llm_utils.py. These providers expose a chat() method that accepts the model name, message list, and a format argument enforcing JSON structured output based on the EvaluationData schema.
To capture raw responses before processing, wrap the provider's chat() method:
# Debugging wrapper in llm_utils.py or evaluator.py
def debug_chat_call(provider, messages, format):
raw_response = provider.chat(
model="gemini-pro", # or your configured model
messages=messages,
format=format
)
# Inspect the raw string before JSON extraction
with open("debug_llm_response.json", "w") as f:
f.write(str(raw_response))
return raw_response
Debugging JSON Extraction and Schema Validation
After receiving the LLM output, the pipeline passes the response through extract_json_from_response() in llm_utils.py. This utility strips markdown code-block fences (```json ... ```) and removes optional wrapping characters to isolate the pure JSON payload.
Common failures occur when models return explanatory text outside the JSON blocks or use irregular fence formatting. To inspect this stage:
# In llm_utils.py, modify extract_json_from_response()
def extract_json_from_response(response: str) -> dict:
print(f"Raw response preview: {response[:200]}...")
# Original stripping logic
cleaned = response.strip()
if cleaned.startswith("```json"):
cleaned = cleaned[7:]
if cleaned.endswith("```"):
cleaned = cleaned[:-3]
print(f"Cleaned JSON string: {cleaned[:200]}...")
return json.loads(cleaned)
Validating Provider-Specific Behavior
Different providers handle the format parameter differently. GeminiProvider may enforce JSON mode at the API level, while OllamaProvider might rely on system prompting for structure. When debugging provider-specific issues:
- OllamaProvider: Check that the local model supports the
jsonformat mode and that theEvaluationDataschema is properly serialized in the request. - GeminiProvider: Verify that the
response_mime_typeis set toapplication/jsonin the generation config.
# Provider initialization inspection
provider = initialize_llm_provider("gemini") # or "ollama"
print(f"Provider type: {type(provider).__name__}")
print(f"Supports structured output: {hasattr(provider, 'supports_json_schema')}")
Summary
- Prompt inspection: Log outputs from
resume_evaluation_criteria.jinjaandresume_evaluation_system_message.jinjabefore sending to the LLM. - Raw response capture: Intercept the return value from
chat()inOllamaProviderorGeminiProviderto examine pre-parsing output. - JSON extraction: Add instrumentation to
extract_json_from_response()inllm_utils.pyto diagnose markdown fence stripping failures. - Provider differences: Verify that
initialize_llm_provider()returns the expected concrete class and that theformatargument matches provider capabilities.
Frequently Asked Questions
Where is the resume evaluation logic located?
The core logic resides in the ResumeEvaluator class within evaluator.py. This class coordinates template rendering, LLM provider initialization via initialize_llm_provider() in llm_utils.py, and JSON extraction through extract_json_from_response().
How do I view the raw LLM response before JSON parsing?
Wrap the provider's chat() method call to capture the raw string output before it reaches extract_json_from_response(). Write this output to a debug file or stdout to inspect markdown fences, whitespace issues, or explanatory text that prevents valid JSON parsing.
What causes JSON parsing errors in resume evaluation?
Parsing errors typically occur when the LLM returns markdown code blocks (```json ... ```) that extract_json_from_response() fails to strip completely, or when the model outputs conversational text alongside the JSON payload. Verify that the format parameter in the chat() call strictly enforces the EvaluationData schema.
Can I switch between Ollama and Gemini providers for debugging?
Yes. Modify the configuration passed to initialize_llm_provider() in llm_utils.py to return either OllamaProvider or GeminiProvider. Each provider handles the EvaluationData schema differently, so switching providers can help isolate whether issues stem from model behavior or provider-specific JSON enforcement.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →