# How Development Mode Caching Works for Resumes and GitHub Data in the Hiring Agent Repository

> Learn how hiring agent caching in development mode speeds up processing resumes and GitHub data by storing results locally in JSON files, avoiding repeated parsing and API calls.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: internals
- Published: 2026-06-27

---

**When `DEVELOPMENT_MODE` is set to `True` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py), the hiring-agent repository stores parsed resume PDFs and GitHub API responses as JSON files in a local `cache/` directory, bypassing expensive re-parsing and HTTP requests on subsequent runs.**

The interviewstreet/hiring-agent project accelerates iterative development through an intelligent caching layer that activates exclusively in development environments. By toggling the `DEVELOPMENT_MODE` environment variable in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py), developers can persist intermediate processing results for resume PDFs and GitHub API data, eliminating redundant I/O operations and network latency during debugging and feature development.

## Resume PDF Caching Implementation

The resume processing pipeline in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) implements a robust file-based caching mechanism that serializes extracted `JSONResume` objects to disk.

### Cache File Naming Convention

The system generates deterministic cache filenames based on the input PDF's basename. In [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 215–217), the code constructs the path as follows:

```python
cache_filename = f"cache/resumecache_{os.path.basename(pdf_path).replace('.pdf', '')}.json"

```

This creates files like [`cache/resumecache_john_doe.json`](https://github.com/interviewstreet/hiring-agent/blob/main/cache/resumecache_john_doe.json) for an input file named `john_doe.pdf`, ensuring each resume maintains its own isolated cache entry.

### Loading from Cache with Validation

When `DEVELOPMENT_MODE` is enabled, the application checks for existing cache files before invoking the PDF extraction engine. The loading logic in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 226–231) attempts deserialization directly into a `JSONResume` object:

```python
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached data from {cache_filename}")
    cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
    loaded_resume = JSONResume(**cached_data)

```

If the cache file is corrupted, contains invalid JSON, or fails schema validation, the system automatically purges the invalid file and falls back to full PDF processing (lines 237–240):

```python
except Exception as e:
    print(f"⚠️ Warning: Invalid cache file {cache_filename}: {e}")
    os.remove(cache_filename)

```

### Writing Extracted Data to Cache

Following a successful PDF extraction, the pipeline persists the `JSONResume` dictionary representation to disk only when development mode is active. The write operation in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 259–266) ensures the cache directory exists before serialization:

```python
if DEVELOPMENT_MODE:
    os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
    Path(cache_filename).write_text(json.dumps(resume.dict()), encoding="utf-8")

```

## GitHub Data Caching Mechanism

The GitHub integration module ([`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)) implements an analogous caching strategy for API responses, preventing redundant HTTP requests during development iterations.

### Deterministic Cache Key Generation

Cache filenames incorporate the request URL and query parameters to create unique identifiers for each endpoint. As implemented in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 18–25):

```python
filename = f"cache/gh_githubcache_{url_parts}_{param_str}.json"

```

This generates paths like [`cache/gh_githubcache_users_octocat_.json`](https://github.com/interviewstreet/hiring-agent/blob/main/cache/gh_githubcache_users_octocat_.json) for a user profile request, ensuring distinct API calls do not collide.

### Cache Retrieval and Fallback

The fetch logic prioritizes cached responses over live network calls. In [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 35–42), the system returns HTTP 200 status codes for cache hits to maintain interface consistency:

```python
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached GitHub data from {cache_filename}")
    cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
    return 200, cached_data

```

Similar to the resume pipeline, corrupted GitHub cache files trigger automatic removal and fresh API requests (lines 44–49):

```python
except Exception as e:
    print(f"⚠️ Warning: Error reading cache file {cache_filename}: {e}")
    os.remove(cache_filename)

```

### Persisting API Responses

Successful API responses (status code 200) are serialized to JSON immediately after retrieval. The persistence logic in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 104–107) writes only valid responses:

```python
if DEVELOPMENT_MODE and status_code == 200:
    os.makedirs("cache", exist_ok=True)
    Path(cache_filename).write_text(json.dumps(data), encoding="utf-8")

```

## Configuration and Environment Setup

The caching behavior is controlled by a single boolean flag in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py). Setting `DEVELOPMENT_MODE = True` enables both the resume and GitHub caching layers simultaneously. When this flag is `False`, the application bypasses all cache read and write operations, ensuring production environments always process fresh data.

## Practical Code Examples

### Processing Resumes with Cache Acceleration

The following pattern demonstrates how subsequent calls to `evaluate_resume` leverage the cache automatically:

```python
from score import evaluate_resume

pdf_path = "samples/jane_doe_resume.pdf"

# First execution parses the PDF and writes to cache/resumecache_jane_doe_resume.json

result1 = evaluate_resume(pdf_path)

# Second execution loads directly from cache, skipping PDF extraction

result2 = evaluate_resume(pdf_path)

```

### Fetching GitHub Data with Local Persistence

Similarly, GitHub API calls utilize the cache layer transparently:

```python
from github import fetch_github_data

# Initial call performs HTTP request and caches the response

status, data = fetch_github_data("https://api.github.com/users/octocat")

# Subsequent calls read from cache/gh_githubcache_users_octocat_.json

status, data = fetch_github_data("https://api.github.com/users/octocat")

```

Both examples require `DEVELOPMENT_MODE = True` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py) to activate caching behavior.

## Summary

- **Development mode caching** is controlled by the `DEVELOPMENT_MODE` boolean in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py) and applies to both resume processing and GitHub API interactions.
- **Resume caches** are stored as `cache/resumecache_{filename}.json` and contain serialized `JSONResume` objects, with automatic validation and cleanup of corrupted files.
- **GitHub caches** use the naming pattern `cache/gh_githubcache_{url}_{params}.json` to store API responses, returning status code 200 for cache hits to maintain consistent interfaces.
- **Automatic recovery** occurs when cache files are invalid—both modules delete corrupted files and fall back to fresh processing or HTTP requests.
- **On-demand directory creation** ensures the `cache/` folder is created automatically via `os.makedirs(..., exist_ok=True)` when writing the first cache entry.

## Frequently Asked Questions

### What happens if a cache file becomes corrupted?

Both [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) and [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) implement defensive error handling that catches deserialization exceptions, prints a warning message, deletes the invalid cache file using `os.remove()`, and proceeds with fresh data extraction or API requests. This ensures corrupted caches never block development workflow.

### How do I completely disable caching for production deployments?

Set `DEVELOPMENT_MODE = False` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py). When this flag is disabled, the application skips all cache existence checks and write operations, forcing live PDF parsing and GitHub API calls on every execution.

### Where exactly are cache files stored?

All cache files reside in a `cache/` directory relative to the execution path. Resume caches follow the pattern `cache/resumecache_*.json`, while GitHub caches use `cache/gh_githubcache_*.json`. The directories are created automatically on first write via `os.makedirs("cache", exist_ok=True)`.

### Does caching bypass JSONResume schema validation?

No. When loading from cache in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py), the system instantiates a `JSONResume(**cached_data)` object, which enforces the Pydantic schema validation. If the cached JSON violates the schema, the instantiation raises an exception, triggering cache deletion and re-processing of the original PDF.