# How the Caching Mechanism Works in Development Mode in Hiring-Agent

> Discover how Hiring Agent's caching mechanism speeds up development. Learn how JSON files and the DEVELOPMENT_MODE flag optimize PDF parsing and GitHub API calls by caching intermediate results.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: internals
- Published: 2026-07-08

---

**During development, the Hiring-Agent project stores intermediate results as JSON files on disk, checking for cached data before performing expensive operations like PDF parsing or GitHub API calls when the `DEVELOPMENT_MODE` flag is enabled.**

The `interviewstreet/hiring-agent` repository uses a file-based caching mechanism to accelerate iterative development workflows. When the `DEVELOPMENT_MODE` flag is set to `True` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py), the system persists parsed resume data and GitHub API responses to the `cache/` directory. This allows developers to skip redundant processing and API calls on subsequent runs while maintaining deterministic output.

## Resume Processing Cache in score.py

The resume evaluation pipeline in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) implements a deterministic caching layer for PDF parsing results. This avoids the overhead of repeatedly extracting text and structure from the same resume files during development cycles.

### Cache Path Construction and Lookup

Before parsing a PDF, the code constructs a cache filename based on the input file's basename. The pattern `cache/resumecache_<pdf-basename>.json` ensures each resume maps to a unique cache entry.

According to the source in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 216-219), the path is built dynamically:

```python
cache_filename = f"cache/resumecache_{os.path.basename(pdf_path)}.json"

```

When `DEVELOPMENT_MODE` is active, the system checks for this file's existence before invoking the PDF parser (lines 227-231). If found, the JSON is deserialized into a `JSONResume` object, bypassing the expensive parsing logic entirely.

### Handling Corrupted Cache Files

The implementation includes robust error handling for invalid cache data. If JSON decoding fails, [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) logs a warning, deletes the corrupted file, and falls back to full PDF processing (lines 237-244). This prevents stale or malformed caches from blocking development workflows.

### Writing Cache After Successful Parsing

Upon successful PDF extraction, the system creates the `cache/` directory if needed and persists the structured resume data as JSON (lines 259-261). This write-back strategy ensures the next execution run loads instantly from disk rather than re-parsing the document.

## GitHub API Response Cache in github.py

External API calls to GitHub are similarly cached to avoid rate limiting and network latency during development. The [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) module wraps all API requests with a deterministic caching layer.

### Deterministic Filename Generation

The helper function `_create_cache_filename` (lines 18-25) generates cache keys by sanitizing the API URL and appending a hash of request parameters. This produces filenames like `cache/gh_githubcache_<url-parts>[_<param-hash>].json`, ensuring unique storage per endpoint and parameter combination.

### Cache Retrieval and Fallback Logic

Before executing an HTTP request, the code checks for an existing cache file when `DEVELOPMENT_MODE` is enabled (lines 35-42). If present, the JSON response is loaded and returned immediately, simulating a successful API call without network overhead.

When cache files contain invalid JSON, the system logs a warning, removes the defective file, and proceeds with a live API request (lines 44-50). This graceful degradation ensures development isn't halted by cache corruption.

### Persisting Successful Responses

After receiving a 200 OK response from GitHub, the response body is written to the cache directory (lines 104-107). Any exceptions during the write operation are logged but non-fatal, allowing the application to continue even if disk persistence fails (line 111).

## Configuration and Activation

The caching behavior is controlled by the `DEVELOPMENT_MODE` boolean defined in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py) (line 6). By default, this is hard-coded to `True` for local development environments.

```python

# From config.py line 6

DEVELOPMENT_MODE = True

```

Production deployments should override this flag to disable caching and ensure fresh data processing. This can be achieved by modifying the configuration file or injecting environment variables before application startup.

## Practical Code Examples

The following examples demonstrate cache utilization in typical development scenarios.

**Running the resume evaluator with automatic caching:**

```python
from score import main

# First run parses the PDF and writes cache/resumecache_jane_doe_resume.pdf.json

result = main("samples/jane_doe_resume.pdf")

# Subsequent runs load the JSONResume object from disk instantly

result_cached = main("samples/jane_doe_resume.pdf")

```

**Fetching GitHub data with response caching:**

```python
from github import request_github_api

# Initial call contacts the live API and caches the response

status, data = request_github_api(
    api_url="https://api.github.com/users/interviewstreet",
    params=None,
)

# Second call reads from cache/gh_githubcache_api_github_com_users_interviewstreet.json

status_cached, data_cached = request_github_api(
    api_url="https://api.github.com/users/interviewstreet",
    params=None,
)

```

**Disabling cache for production execution:**

```bash

# Override the development flag to force fresh processing

export DEVELOPMENT_MODE=False
python -m score samples/jane_doe_resume.pdf

```

## Summary

- The **development mode caching mechanism** in Hiring-Agent relies on JSON file storage under the `cache/` directory, activated by the `DEVELOPMENT_MODE` flag in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py).
- **Resume parsing results** are cached in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) using filenames derived from PDF basenames, with automatic fallback to full parsing when caches are missing or corrupted.
- **GitHub API responses** are cached in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) using deterministic filenames generated from URL and parameter hashes, preventing redundant network requests.
- Both implementations feature **graceful degradation**, deleting invalid cache files and proceeding with live processing when JSON decoding fails.
- Cache writes occur **after successful operations**, ensuring subsequent development runs benefit from persisted intermediate state while maintaining data consistency.

## Frequently Asked Questions

### Where is the development mode cache stored in Hiring-Agent?

All cached data is stored as JSON files in the `cache/` directory at the project root. Resume caches follow the pattern `resumecache_<pdf-basename>.json`, while GitHub API caches use the prefix `gh_githubcache_` followed by URL-derived identifiers.

### How does the system handle corrupted cache files?

When JSON decoding fails during cache loading, both [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 237-244) and [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 44-50) log a warning message, delete the corrupted file, and automatically fall back to live processing or API requests. This ensures that cache corruption never blocks development workflows.

### Can I disable caching without modifying the source code?

While `DEVELOPMENT_MODE` is hard-coded to `True` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py), production deployments typically override this value via environment variables or configuration management tools. Setting this flag to `False` disables all cache lookups and writes, forcing fresh PDF parsing and API calls on every execution.

### Does the cache mechanism affect production performance?

The caching layer is designed specifically for development mode and should be disabled in production by setting `DEVELOPMENT_MODE = False`. When disabled, the system bypasses all disk I/O related to caching, ensuring production runs process fresh data without filesystem overhead.