# How Caching Works in the Hiring Agent Pipeline: A Complete Developer Guide

> Learn how the Hiring Agent pipeline uses file-based caching with the cache/ directory to optimize local development runs. Discover how parsed resumes and GitHub API data are stored via JSON for faster access.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-07-09

---

**The Hiring Agent pipeline implements a file-based caching layer under the `cache/` directory that only activates when `DEVELOPMENT_MODE` is enabled, storing JSON representations of parsed resumes and GitHub API responses to speed up repeated local runs.**

The `interviewstreet/hiring-agent` repository uses a development-only caching mechanism to eliminate redundant PDF parsing and API calls during local iteration. Understanding how caching in the Hiring Agent pipeline functions helps you optimize your development workflow while ensuring production deployments always fetch fresh data.

## Resume Processing Cache in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py)

The pipeline caches expensive resume extraction operations to avoid re-parsing PDFs on every run.

### Cache File Structure and Naming

When processing a PDF file, the system generates a cache filename based on the input document:

```python
cache_filename = f"cache/resumecache_{os.path.basename(pdf_path).replace('.pdf', '')}.json"

```

This creates files like [`cache/resumecache_john_doe.json`](https://github.com/interviewstreet/hiring-agent/blob/main/cache/resumecache_john_doe.json) directly under the project root.

### Cache Read and Write Logic

In [`main/score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/score.py) (lines 215–227 and 259–321), the pipeline checks for existing cache before performing extraction:

```python
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached data from {cache_filename}")
    cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
    loaded_resume = JSONResume(**cached_data)
    cache_loaded = True

```

If no cache exists, the code extracts the resume data and persists it to disk:

```python
if not cache_loaded:
    # Extraction logic occurs here...

    os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
    Path(cache_filename).write_text(json.dumps(resume_dict, ensure_ascii=False, indent=2))

```

## GitHub API Caching in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)

The [`main/github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/github.py) module implements a separate cache for external API calls to prevent rate limiting and reduce latency during development.

### Dynamic Filename Generation

The `_create_cache_filename` function (lines 18–42) constructs unique filenames based on the API endpoint and parameters:

```python
def _create_cache_filename(api_url: str, params: dict = None) -> str:
    url_parts = "_".join(api_url.split("/")[3:])  # strip protocol & domain

    if params:
        param_str = "_".join(f"{k}-{v}" for k, v in sorted(params.items()))
        filename = f"cache/gh_githubcache_{url_parts}_{param_str}.json"
    else:
        filename = f"cache/gh_githubcache_{url_parts}.json"
    return filename

```

This produces filenames like [`cache/gh_githubcache_users_username_repos.json`](https://github.com/interviewstreet/hiring-agent/blob/main/cache/gh_githubcache_users_username_repos.json).

### Request Short-Circuiting

Before making any HTTP request, the helper checks for cached responses (lines 106–111):

```python
cache_filename = _create_cache_filename(api_url, params)
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached GitHub data from {cache_filename}")
    cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
    return 200, cached_data

```

After a successful API call, the raw JSON response is written to disk:

```python
os.makedirs("cache", exist_ok=True)
Path(cache_filename).write_text(json.dumps(data, ensure_ascii=False, indent=2))

```

## Cache Lifecycle and Invalidation

The Hiring Agent pipeline follows a predictable lifecycle for cached data:

1. **First run (cold cache)** – The system parses the PDF, queries the GitHub API, and writes JSON files to `cache/`.
2. **Subsequent runs** – The existence check (`os.path.exists`) short-circuits heavy operations, loading previously saved JSON directly from disk.
3. **Corruption handling** – If JSON parsing fails, the code removes the invalid file and falls back to re-processing, ensuring broken caches never hide runtime errors.

## Configuration and Environment Setup

Caching is strictly **development-only** and controlled by the `DEVELOPMENT_MODE` flag defined in [`main/config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/config.py).

To enable caching during local development:

```bash
export DEVELOPMENT_MODE=1
python -m main.score path/to/resume.pdf

```

On first execution, this creates:
- `cache/resumecache_<pdf-name>.json`
- `cache/githubcache_<pdf-name>.json`

Later runs will instantly load these files instead of re-fetching data.

To manually clear the cache and force re-processing:

```python
import shutil
import pathlib

shutil.rmtree(pathlib.Path("cache"), ignore_errors=True)
print("Cache cleared – next run will re-process everything.")

```

## Summary

- **File-based storage** – All cached data lives as JSON files under the top-level `cache/` directory.
- **Development-only activation** – The `DEVELOPMENT_MODE` flag in [`main/config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/config.py) gates all caching logic; production deployments always use fresh data.
- **Two-tier caching** – [`main/score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/score.py) caches parsed resumes while [`main/github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/github.py) caches raw API responses.
- **Automatic invalidation** – Corrupted cache files trigger automatic deletion and fallback to live processing.
- **Performance impact** – Subsequent pipeline runs skip PDF parsing and API calls entirely, reducing iteration time from minutes to seconds.

## Frequently Asked Questions

### Where does the Hiring Agent pipeline store cache files?

The pipeline stores all cache files in a `cache/` directory at the project root. Resume data is saved as `resumecache_<pdf-name>.json`, while GitHub API responses follow the pattern `gh_githubcache_<url_parts>_<params>.json`.

### How do I clear the cache to force re-processing?

Delete the `cache/` directory or use Python's `shutil.rmtree()` to remove it programmatically. The next pipeline run will detect missing cache files and regenerate them by parsing PDFs and calling the GitHub API fresh.

### Why is caching disabled in production?

The `DEVELOPMENT_MODE` flag defaults to `False` in production configurations to guarantee that live deployments always work with the latest candidate data and fresh API responses. This prevents stale resumes or outdated GitHub statistics from affecting hiring decisions.

### What happens if a cache file becomes corrupted?

If `json.loads()` fails when reading a cache file, the code catches the exception, removes the corrupted file, and proceeds to re-process the data from the original source. This ensures that disk errors or manual edits to cache files never cause permanent pipeline failures.