Caching System Behavior in Development Mode: How interviewstreet/hiring-agent Optimizes Local Development

When DEVELOPMENT_MODE is set to True in config.py, the repository activates a lightweight file-based cache that speeds up résumé parsing and GitHub API calls by storing JSON responses locally, while automatically removing corrupted files and remaining completely inactive in production.

The interviewstreet/hiring-agent repository implements a development-only caching mechanism to accelerate iterative testing and local development workflows. This system leverages a boolean flag defined in config.py to toggle between persistent file-based caching during development and direct API access in production. Understanding this caching system's behavior in development mode helps developers optimize their local setup while ensuring production deployments always fetch fresh data.

How Development Mode Activates the Cache

The cache is gated by the DEVELOPMENT_MODE constant defined in config.py. When this flag evaluates to True, the application checks for existing cache files before executing expensive operations like PDF parsing or HTTP requests.

Resume Parsing Cache in score.py

In score.py, the system caches parsed résumé data to avoid reprocessing PDFs on every run.

Loading Cached Résumés

At the start of the scoring process, the code checks for an existing cache file:


# development‑mode cache check

if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached data from {cache_filename}")
    try:
        cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
        loaded_resume = JSONResume(**cached_data)
        cache_loaded = True
    except Exception as e:
        print(f"⚠️ Warning: Invalid cache file {cache_filename}: {e}")
        print("Ignoring cache and reprocessing PDF...")
        os.remove(cache_filename)

Writing New Cache Entries

After successful parsing, the system persists the structured data:

if not cache_loaded:
    # ... parsing logic ...

    if resume_data_is_valid:
        os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
        Path(cache_filename).write_text(
            json.dumps(resume_dict, ensure_ascii=False, indent=2), encoding="utf-8"
        )

GitHub API Cache in github.py

The github.py module implements similar short-circuiting for GitHub REST API calls to prevent rate limiting and reduce latency during development.

Short-Circuiting HTTP Requests

Before making network requests, the code attempts to load cached responses:

cache_filename = _create_cache_filename(api_url, params)
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached GitHub data from {cache_filename}")
    try:
        cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
        if cached_data:
            return 200, cached_data
    except Exception as e:
        print(f"⚠️ Warning: Error reading cache file {cache_filename}: {e}")
        os.remove(cache_filename)

Persisting API Responses

Successful HTTP 200 responses are written to disk for subsequent runs:

if DEVELOPMENT_MODE and status_code == 200:
    os.makedirs("cache", exist_ok=True)
    Path(cache_filename).write_text(
        json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8"
    )

Cache File Management and Invalidation

Cache files follow deterministic naming conventions based on input parameters—PDF filenames for résumé parsing and URL parameters for GitHub requests. The system handles cache corruption gracefully by catching exceptions during load operations, printing diagnostic warnings, and deleting invalid files before falling back to fresh processing.

Production Behavior

When DEVELOPMENT_MODE is set to False, all if DEVELOPMENT_MODE conditional blocks are skipped entirely. In this state, the application never checks for cache files, never writes new entries, and always processes PDFs from scratch while making live GitHub API calls.

Summary

  • Development mode (DEVELOPMENT_MODE = True in config.py) enables file-based JSON caching for both résumé parsing and GitHub API calls.
  • score.py caches parsed résumé data as resumecache_*.json files to eliminate redundant PDF processing.
  • github.py stores GitHub REST API responses as gh_githubcache_*.json files to prevent unnecessary network requests.
  • Cache files are named deterministically based on input parameters and stored in a cache/ directory created on demand via os.makedirs(..., exist_ok=True).
  • Corrupted cache files trigger automatic deletion and reprocessing, with warnings printed to the console.
  • Production mode completely bypasses the caching layer, ensuring fresh data on every execution.

Frequently Asked Questions

How do I enable or disable the caching system?

Set DEVELOPMENT_MODE to True or False in config.py. When enabled, the system caches résumé parses and GitHub API responses locally; when disabled, all cache logic is bypassed and fresh data is fetched every time.

What happens if a cache file becomes corrupted?

If json.loads() raises an exception while reading a cache file in either score.py or github.py, the system prints a warning message, deletes the corrupted file using os.remove(cache_filename), and proceeds to reprocess the PDF or reissue the API request.

Where are cache files stored and how are they named?

Cache files are stored in a cache/ directory created automatically when needed. Files are named deterministically based on the input—resumecache_*.json for résumés (based on PDF filename) and gh_githubcache_*.json for GitHub calls (based on URL and parameters)—ensuring identical inputs map to identical cache files.

Is the caching system safe to use in production?

No, the caching system is designed exclusively for development convenience. When DEVELOPMENT_MODE is False (production), the conditional blocks containing cache logic are never executed, ensuring the application always retrieves current data and never relies on potentially stale local files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →