# How to Debug and Inspect LLM Responses During the Resume Evaluation Process

> Learn to debug and inspect LLM responses for resume evaluation. Intercept calls, log LLM outputs, and validate JSON extraction in the hiring-agent repository.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-06-27

---

**You can debug LLM responses in the hiring-agent repository by intercepting calls in the `ResumeEvaluator` class, logging outputs from [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py), and validating JSON extraction before schema parsing.**

The resume evaluation pipeline in the `interviewstreet/hiring-agent` repository relies on structured LLM outputs to score candidates against specific criteria. When responses fail schema validation or return unexpected formats, developers need visibility into the prompt construction, raw API responses, and JSON parsing stages to diagnose issues effectively.

## Understanding the Resume Evaluation Architecture

The evaluation flow centers on the `ResumeEvaluator` class defined in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py). When `evaluate_resume()` is invoked, it orchestrates three critical phases: prompt rendering, LLM invocation, and structured data extraction.

The pipeline first renders evaluation criteria from `prompts/templates/resume_evaluation_criteria.jinja` and system instructions from `prompts/templates/resume_evaluation_system_message.jinja`. These templates combine with candidate data to form the complete message payload sent to the LLM provider.

## Inspecting Prompt Construction Before Sending

To debug issues with model comprehension or instruction following, inspect the rendered prompts immediately before the LLM call. The `evaluate_resume()` method builds the final prompt string using Jinja2 template rendering.

Add logging or breakpoints immediately after template rendering to capture the exact text sent to the provider:

```python

# In evaluator.py, inside evaluate_resume()

from pathlib import Path

# Log rendered templates for inspection

criteria_template = Path("prompts/templates/resume_evaluation_criteria.jinja").read_text()
system_template = Path("prompts/templates/resume_evaluation_system_message.jinja").read_text()

print(f"System prompt length: {len(system_template)}")
print(f"Criteria prompt: {criteria_template[:500]}...")  # Truncate for readability

```

## Monitoring Raw LLM Responses

The concrete provider implementation—either `OllamaProvider` or `GeminiProvider`—is instantiated by `initialize_llm_provider()` in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py). These providers expose a `chat()` method that accepts the model name, message list, and a `format` argument enforcing JSON structured output based on the `EvaluationData` schema.

To capture raw responses before processing, wrap the provider's `chat()` method:

```python

# Debugging wrapper in llm_utils.py or evaluator.py

def debug_chat_call(provider, messages, format):
    raw_response = provider.chat(
        model="gemini-pro",  # or your configured model

        messages=messages,
        format=format
    )
    # Inspect the raw string before JSON extraction

    with open("debug_llm_response.json", "w") as f:
        f.write(str(raw_response))
    return raw_response

```

## Debugging JSON Extraction and Schema Validation

After receiving the LLM output, the pipeline passes the response through `extract_json_from_response()` in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py). This utility strips markdown code-block fences (`` ```json ... ``` ``) and removes optional wrapping characters to isolate the pure JSON payload.

Common failures occur when models return explanatory text outside the JSON blocks or use irregular fence formatting. To inspect this stage:

```python

# In llm_utils.py, modify extract_json_from_response()

def extract_json_from_response(response: str) -> dict:
    print(f"Raw response preview: {response[:200]}...")
    
    # Original stripping logic

    cleaned = response.strip()
    if cleaned.startswith("```json"):
        cleaned = cleaned[7:]
    if cleaned.endswith("```"):
        cleaned = cleaned[:-3]
    
    print(f"Cleaned JSON string: {cleaned[:200]}...")
    return json.loads(cleaned)

```

## Validating Provider-Specific Behavior

Different providers handle the `format` parameter differently. `GeminiProvider` may enforce JSON mode at the API level, while `OllamaProvider` might rely on system prompting for structure. When debugging provider-specific issues:

- **OllamaProvider**: Check that the local model supports the `json` format mode and that the `EvaluationData` schema is properly serialized in the request.
- **GeminiProvider**: Verify that the `response_mime_type` is set to `application/json` in the generation config.

```python

# Provider initialization inspection

provider = initialize_llm_provider("gemini")  # or "ollama"

print(f"Provider type: {type(provider).__name__}")
print(f"Supports structured output: {hasattr(provider, 'supports_json_schema')}")

```

## Summary

- **Prompt inspection**: Log outputs from `resume_evaluation_criteria.jinja` and `resume_evaluation_system_message.jinja` before sending to the LLM.
- **Raw response capture**: Intercept the return value from `chat()` in `OllamaProvider` or `GeminiProvider` to examine pre-parsing output.
- **JSON extraction**: Add instrumentation to `extract_json_from_response()` in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) to diagnose markdown fence stripping failures.
- **Provider differences**: Verify that `initialize_llm_provider()` returns the expected concrete class and that the `format` argument matches provider capabilities.

## Frequently Asked Questions

### Where is the resume evaluation logic located?

The core logic resides in the `ResumeEvaluator` class within [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py). This class coordinates template rendering, LLM provider initialization via `initialize_llm_provider()` in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py), and JSON extraction through `extract_json_from_response()`.

### How do I view the raw LLM response before JSON parsing?

Wrap the provider's `chat()` method call to capture the raw string output before it reaches `extract_json_from_response()`. Write this output to a debug file or stdout to inspect markdown fences, whitespace issues, or explanatory text that prevents valid JSON parsing.

### What causes JSON parsing errors in resume evaluation?

Parsing errors typically occur when the LLM returns markdown code blocks (`` ```json ... ``` ``) that `extract_json_from_response()` fails to strip completely, or when the model outputs conversational text alongside the JSON payload. Verify that the `format` parameter in the `chat()` call strictly enforces the `EvaluationData` schema.

### Can I switch between Ollama and Gemini providers for debugging?

Yes. Modify the configuration passed to `initialize_llm_provider()` in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) to return either `OllamaProvider` or `GeminiProvider`. Each provider handles the `EvaluationData` schema differently, so switching providers can help isolate whether issues stem from model behavior or provider-specific JSON enforcement.