# Caching System Behavior in Development Mode: How interviewstreet/hiring-agent Optimizes Local Development

> Discover how interviewstreet/hiring-agent optimizes development with its caching system. Learn how it speeds up processes and remains inactive in production.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: internals
- Published: 2026-07-05

---

**When `DEVELOPMENT_MODE` is set to `True` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py), the repository activates a lightweight file-based cache that speeds up résumé parsing and GitHub API calls by storing JSON responses locally, while automatically removing corrupted files and remaining completely inactive in production.**

The `interviewstreet/hiring-agent` repository implements a development-only caching mechanism to accelerate iterative testing and local development workflows. This system leverages a boolean flag defined in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py) to toggle between persistent file-based caching during development and direct API access in production. Understanding this **caching system's behavior in development mode** helps developers optimize their local setup while ensuring production deployments always fetch fresh data.

## How Development Mode Activates the Cache

The cache is gated by the `DEVELOPMENT_MODE` constant defined in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py). When this flag evaluates to `True`, the application checks for existing cache files before executing expensive operations like PDF parsing or HTTP requests.

## Resume Parsing Cache in score.py

In [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py), the system caches parsed résumé data to avoid reprocessing PDFs on every run.

### Loading Cached Résumés

At the start of the scoring process, the code checks for an existing cache file:

```python

# development‑mode cache check

if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached data from {cache_filename}")
    try:
        cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
        loaded_resume = JSONResume(**cached_data)
        cache_loaded = True
    except Exception as e:
        print(f"⚠️ Warning: Invalid cache file {cache_filename}: {e}")
        print("Ignoring cache and reprocessing PDF...")
        os.remove(cache_filename)

```

### Writing New Cache Entries

After successful parsing, the system persists the structured data:

```python
if not cache_loaded:
    # ... parsing logic ...

    if resume_data_is_valid:
        os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
        Path(cache_filename).write_text(
            json.dumps(resume_dict, ensure_ascii=False, indent=2), encoding="utf-8"
        )

```

## GitHub API Cache in github.py

The [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) module implements similar short-circuiting for GitHub REST API calls to prevent rate limiting and reduce latency during development.

### Short-Circuiting HTTP Requests

Before making network requests, the code attempts to load cached responses:

```python
cache_filename = _create_cache_filename(api_url, params)
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached GitHub data from {cache_filename}")
    try:
        cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
        if cached_data:
            return 200, cached_data
    except Exception as e:
        print(f"⚠️ Warning: Error reading cache file {cache_filename}: {e}")
        os.remove(cache_filename)

```

### Persisting API Responses

Successful HTTP 200 responses are written to disk for subsequent runs:

```python
if DEVELOPMENT_MODE and status_code == 200:
    os.makedirs("cache", exist_ok=True)
    Path(cache_filename).write_text(
        json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8"
    )

```

## Cache File Management and Invalidation

Cache files follow deterministic naming conventions based on input parameters—PDF filenames for résumé parsing and URL parameters for GitHub requests. The system handles cache corruption gracefully by catching exceptions during load operations, printing diagnostic warnings, and deleting invalid files before falling back to fresh processing.

## Production Behavior

When `DEVELOPMENT_MODE` is set to `False`, all `if DEVELOPMENT_MODE` conditional blocks are skipped entirely. In this state, the application never checks for cache files, never writes new entries, and always processes PDFs from scratch while making live GitHub API calls.

## Summary

- **Development mode** (`DEVELOPMENT_MODE = True` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py)) enables file-based JSON caching for both résumé parsing and GitHub API calls.
- **[`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py)** caches parsed résumé data as `resumecache_*.json` files to eliminate redundant PDF processing.
- **[`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)** stores GitHub REST API responses as `gh_githubcache_*.json` files to prevent unnecessary network requests.
- Cache files are named deterministically based on input parameters and stored in a `cache/` directory created on demand via `os.makedirs(..., exist_ok=True)`.
- Corrupted cache files trigger automatic deletion and reprocessing, with warnings printed to the console.
- **Production mode** completely bypasses the caching layer, ensuring fresh data on every execution.

## Frequently Asked Questions

### How do I enable or disable the caching system?

Set `DEVELOPMENT_MODE` to `True` or `False` in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py). When enabled, the system caches résumé parses and GitHub API responses locally; when disabled, all cache logic is bypassed and fresh data is fetched every time.

### What happens if a cache file becomes corrupted?

If `json.loads()` raises an exception while reading a cache file in either [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) or [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), the system prints a warning message, deletes the corrupted file using `os.remove(cache_filename)`, and proceeds to reprocess the PDF or reissue the API request.

### Where are cache files stored and how are they named?

Cache files are stored in a `cache/` directory created automatically when needed. Files are named deterministically based on the input—`resumecache_*.json` for résumés (based on PDF filename) and `gh_githubcache_*.json` for GitHub calls (based on URL and parameters)—ensuring identical inputs map to identical cache files.

### Is the caching system safe to use in production?

No, the caching system is designed exclusively for development convenience. When `DEVELOPMENT_MODE` is `False` (production), the conditional blocks containing cache logic are never executed, ensuring the application always retrieves current data and never relies on potentially stale local files.