# How the Development Mode Caching Mechanism Works in InterviewStreet's Hiring Agent

> Discover how InterviewStreet's Hiring Agent uses development mode caching to speed up résumé processing and GitHub API calls, storing parsed data as JSON files and bypassing re-runs.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: internals
- Published: 2026-07-06

---

**When `DEVELOPMENT_MODE` is set to `True` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py), the interviewstreet/hiring-agent repository activates a lightweight file-based cache that stores parsed résumé data and GitHub API responses as JSON files, automatically bypassing expensive re-processing and HTTP calls on subsequent runs while silently invalidating corrupted entries.**

The development mode caching mechanism is designed to accelerate iterative development and testing by avoiding redundant computation and external API requests. According to the interviewstreet/hiring-agent source code, this system operates exclusively when the environment flag is enabled, ensuring production deployments always fetch fresh data.

## Resume Parsing Cache in score.py

The first component of the development mode caching mechanism handles expensive PDF parsing operations. In [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py), parsed résumé data is serialized to JSON and stored for rapid retrieval across multiple scoring runs.

### Cache Location and File Naming

Cache files follow the pattern `resumecache_*.json` and reside in the `cache/` directory. The filename is generated deterministically based on the input PDF filename, ensuring that identical résumés map to identical cache entries.

### Reading and Writing Logic

At the start of the processing pipeline (around lines 227-260 in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py)), the code checks for existing cache entries before invoking the parser:

```python

# development‑mode cache check

if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached data from {cache_filename}")
    try:
        cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
        loaded_resume = JSONResume(**cached_data)
        cache_loaded = True
    except Exception as e:
        print(f"⚠️ Warning: Invalid cache file {cache_filename}: {e}")
        print("Ignoring cache and reprocessing PDF...")
        os.remove(cache_filename)

```

When parsing completes successfully and no cache was loaded, the system writes the result:

```python
if not cache_loaded:
    # ... parsing logic ...

    if resume_data_is_valid:
        os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
        Path(cache_filename).write_text(
            json.dumps(resume_dict, ensure_ascii=False, indent=2), encoding="utf-8"
        )

```

### Cache Invalidation Strategy

If `json.loads()` raises an exception due to malformed JSON or disk corruption, the code immediately removes the offending file using `os.remove(cache_filename)` and falls back to re-processing the original PDF. This defensive pattern ensures that transient disk errors or manual file edits never break the development workflow.

## GitHub API Response Cache in github.py

The second cached component prevents redundant HTTP requests to the GitHub REST API. In [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), successful API responses are persisted to disk and served on subsequent identical requests.

### Request Interception and Cache Keys

Before executing any HTTP call, the code generates a deterministic cache filename using `_create_cache_filename(api_url, params)` (lines 35-50). If a matching file exists, the network request is skipped entirely:

```python
cache_filename = _create_cache_filename(api_url, params)
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached GitHub data from {cache_filename}")
    try:
        cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
        if cached_data:
            return 200, cached_data
    except Exception as e:
        print(f"⚠️ Warning: Error reading cache file {cache_filename}: {e}")
        os.remove(cache_filename)

```

### Persisting Successful Responses

Only successful responses (`status_code == 200`) are cached to avoid storing error payloads. The write operation creates the `cache/` directory on demand:

```python
if DEVELOPMENT_MODE and status_code == 200:
    os.makedirs("cache", exist_ok=True)
    Path(cache_filename).write_text(
        json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8"
    )

```

## Configuration and Activation

The entire caching layer is gated by a single boolean flag defined in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py). By default, `DEVELOPMENT_MODE = True` enables the cache, while setting it to `False` disables all file I/O operations related to caching. This centralized configuration ensures no cache code paths execute in production environments.

## Production vs. Development Behavior

When `DEVELOPMENT_MODE` is disabled, none of the `if DEVELOPMENT_MODE …` conditionals evaluate to true. Consequently:

- **Résumé parsing** always processes the PDF from scratch
- **GitHub API calls** always execute fresh HTTP requests
- **No cache files** are read, written, or deleted

This strict separation guarantees that production deployments reflect current data, eliminating the risk of stale cache entries affecting candidate scoring results.

## Summary

- The **development mode caching mechanism** is controlled by the `DEVELOPMENT_MODE` flag in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py) and is enabled by default for local development.
- **Resume parsing results** are cached as `resumecache_*.json` files in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py), with automatic re-processing triggered if the JSON is corrupted.
- **GitHub API responses** are stored as `gh_githubcache_*.json` files in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), intercepting network requests only when valid cache exists.
- Cache filenames are generated deterministically from input parameters, ensuring consistent cache hits across identical inputs.
- Corrupted cache files are automatically deleted and regenerated, preventing development workflow interruptions.
- In production mode, all caching logic is bypassed entirely, ensuring fresh data and API responses.

## Frequently Asked Questions

### How do I disable the development mode caching mechanism?

Set `DEVELOPMENT_MODE = False` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py). When this flag is disabled, all cache checks in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) and [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) are skipped, forcing the application to parse PDFs from scratch and execute fresh GitHub API requests on every run.

### What happens if a cache file becomes corrupted?

The code handles corruption defensively. In both [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) and [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), if `json.loads()` raises an exception when reading a cache file, the exception is caught, a warning is printed to the console, `os.remove(cache_filename)` deletes the corrupted file, and the system falls back to the original data source (PDF parsing or HTTP request).

### Where are the cache files stored?

Cache files are stored in the `cache/` directory at the project root. The directory is created automatically on demand via `os.makedirs(..., exist_ok=True)` when the first cache write occurs. Resume caches follow the pattern `resumecache_*.json`, while GitHub API caches use `gh_githubcache_*.json`.

### Is the cache used in production deployments?

No. The development mode caching mechanism is explicitly development-only. When `DEVELOPMENT_MODE` is `False` (the recommended setting for production), none of the cache read or write operations execute, ensuring that production instances always process fresh résumé data and current GitHub API responses.