How to Debug and Inspect LLM Responses During the Resume Evaluation Process

You can debug LLM responses in the hiring-agent repository by intercepting calls in the ResumeEvaluator class, logging outputs from llm_utils.py, and validating JSON extraction before schema parsing.

The resume evaluation pipeline in the interviewstreet/hiring-agent repository relies on structured LLM outputs to score candidates against specific criteria. When responses fail schema validation or return unexpected formats, developers need visibility into the prompt construction, raw API responses, and JSON parsing stages to diagnose issues effectively.

Understanding the Resume Evaluation Architecture

The evaluation flow centers on the ResumeEvaluator class defined in evaluator.py. When evaluate_resume() is invoked, it orchestrates three critical phases: prompt rendering, LLM invocation, and structured data extraction.

The pipeline first renders evaluation criteria from prompts/templates/resume_evaluation_criteria.jinja and system instructions from prompts/templates/resume_evaluation_system_message.jinja. These templates combine with candidate data to form the complete message payload sent to the LLM provider.

Inspecting Prompt Construction Before Sending

To debug issues with model comprehension or instruction following, inspect the rendered prompts immediately before the LLM call. The evaluate_resume() method builds the final prompt string using Jinja2 template rendering.

Add logging or breakpoints immediately after template rendering to capture the exact text sent to the provider:


# In evaluator.py, inside evaluate_resume()

from pathlib import Path

# Log rendered templates for inspection

criteria_template = Path("prompts/templates/resume_evaluation_criteria.jinja").read_text()
system_template = Path("prompts/templates/resume_evaluation_system_message.jinja").read_text()

print(f"System prompt length: {len(system_template)}")
print(f"Criteria prompt: {criteria_template[:500]}...")  # Truncate for readability

Monitoring Raw LLM Responses

The concrete provider implementation—either OllamaProvider or GeminiProvider—is instantiated by initialize_llm_provider() in llm_utils.py. These providers expose a chat() method that accepts the model name, message list, and a format argument enforcing JSON structured output based on the EvaluationData schema.

To capture raw responses before processing, wrap the provider's chat() method:


# Debugging wrapper in llm_utils.py or evaluator.py

def debug_chat_call(provider, messages, format):
    raw_response = provider.chat(
        model="gemini-pro",  # or your configured model

        messages=messages,
        format=format
    )
    # Inspect the raw string before JSON extraction

    with open("debug_llm_response.json", "w") as f:
        f.write(str(raw_response))
    return raw_response

Debugging JSON Extraction and Schema Validation

After receiving the LLM output, the pipeline passes the response through extract_json_from_response() in llm_utils.py. This utility strips markdown code-block fences (```json ... ```) and removes optional wrapping characters to isolate the pure JSON payload.

Common failures occur when models return explanatory text outside the JSON blocks or use irregular fence formatting. To inspect this stage:


# In llm_utils.py, modify extract_json_from_response()

def extract_json_from_response(response: str) -> dict:
    print(f"Raw response preview: {response[:200]}...")
    
    # Original stripping logic

    cleaned = response.strip()
    if cleaned.startswith("```json"):
        cleaned = cleaned[7:]
    if cleaned.endswith("```"):
        cleaned = cleaned[:-3]
    
    print(f"Cleaned JSON string: {cleaned[:200]}...")
    return json.loads(cleaned)

Validating Provider-Specific Behavior

Different providers handle the format parameter differently. GeminiProvider may enforce JSON mode at the API level, while OllamaProvider might rely on system prompting for structure. When debugging provider-specific issues:

  • OllamaProvider: Check that the local model supports the json format mode and that the EvaluationData schema is properly serialized in the request.
  • GeminiProvider: Verify that the response_mime_type is set to application/json in the generation config.

# Provider initialization inspection

provider = initialize_llm_provider("gemini")  # or "ollama"

print(f"Provider type: {type(provider).__name__}")
print(f"Supports structured output: {hasattr(provider, 'supports_json_schema')}")

Summary

  • Prompt inspection: Log outputs from resume_evaluation_criteria.jinja and resume_evaluation_system_message.jinja before sending to the LLM.
  • Raw response capture: Intercept the return value from chat() in OllamaProvider or GeminiProvider to examine pre-parsing output.
  • JSON extraction: Add instrumentation to extract_json_from_response() in llm_utils.py to diagnose markdown fence stripping failures.
  • Provider differences: Verify that initialize_llm_provider() returns the expected concrete class and that the format argument matches provider capabilities.

Frequently Asked Questions

Where is the resume evaluation logic located?

The core logic resides in the ResumeEvaluator class within evaluator.py. This class coordinates template rendering, LLM provider initialization via initialize_llm_provider() in llm_utils.py, and JSON extraction through extract_json_from_response().

How do I view the raw LLM response before JSON parsing?

Wrap the provider's chat() method call to capture the raw string output before it reaches extract_json_from_response(). Write this output to a debug file or stdout to inspect markdown fences, whitespace issues, or explanatory text that prevents valid JSON parsing.

What causes JSON parsing errors in resume evaluation?

Parsing errors typically occur when the LLM returns markdown code blocks (```json ... ```) that extract_json_from_response() fails to strip completely, or when the model outputs conversational text alongside the JSON payload. Verify that the format parameter in the chat() call strictly enforces the EvaluationData schema.

Can I switch between Ollama and Gemini providers for debugging?

Yes. Modify the configuration passed to initialize_llm_provider() in llm_utils.py to return either OllamaProvider or GeminiProvider. Each provider handles the EvaluationData schema differently, so switching providers can help isolate whether issues stem from model behavior or provider-specific JSON enforcement.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →