Where Are Cached Results Stored When DEVELOPMENT_MODE Is Enabled?

When DEVELOPMENT_MODE is set to True in the interviewstreet/hiring-agent repository, cached results are stored in the top-level cache/ directory as JSON files, specifically using resumecache_*.json for resume extractions and githubcache_*.json (or gh_githubcache_*.json) for GitHub API responses.

The hiring-agent application by Interview Street uses a development mode flag to optimize iteration speed during local testing. When DEVELOPMENT_MODE is enabled, the system avoids reprocessing PDFs and re-fetching GitHub data by persisting intermediate results to disk. Understanding the exact storage location and file patterns helps developers debug processing pipelines and manually clear stale cache entries.

Cache Storage Location and File Structure

When DEVELOPMENT_MODE is active, the application creates a cache/ directory at the repository root to store serialized intermediate data. This directory is generated on-demand using os.makedirs with exist_ok=True, ensuring the application does not fail if the folder already exists.

The system maintains two distinct cache types with specific file naming patterns:

Resume Extraction Cache Files

Located at cache/resumecache_<pdf-basename>.json, these files store the JSON output of parsed PDF resumes. According to score.py, after a PDF is processed, the extracted resume dictionary is serialized and written to this location. On subsequent runs, if the file exists, the application loads the JSON directly instead of re-parsing the PDF.

GitHub Data Cache Files

GitHub API responses are cached as cache/githubcache_<pdf-basename>.json or the generic cache/gh_githubcache_*.json pattern used by the GitHub helper module. As implemented in github.py, when the application queries the GitHub API, the response payload is saved to disk. Future executions read this cached data before making new network requests, reducing API rate limit consumption.

How the Caching Mechanism Works

The caching behavior is conditional on the DEVELOPMENT_MODE flag defined in config.py. When enabled, the code checks for existing cache files using os.path.exists() before proceeding with expensive operations.

Resume Processing Logic in score.py

In score.py, the resume extraction logic creates cache files using the following approach:


# score.py – Creating the resume cache

os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
Path(cache_filename).write_text(json.dumps(resume_dict, ensure_ascii=False))

When reading cached data, the code validates file existence before loading:


# Loading cached resume data (development mode only)

if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached data from {cache_filename}")
    cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
    loaded_resume = JSONResume(**cached_data)

GitHub API Logic in github.py

Similarly, github.py implements caching for network requests:


# github.py – Creating the GitHub cache

os.makedirs("cache", exist_ok=True)
Path(cache_filename).write_text(json.dumps(github_data, ensure_ascii=False))

The retrieval logic mirrors the resume cache pattern:


# Loading cached GitHub data (development mode only)

if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached GitHub data from {cache_filename}")
    cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
    return 200, cached_data

Production Behavior vs. Development Mode

When DEVELOPMENT_MODE is set to False, the application skips all os.path.exists() checks for cache files. This ensures a fresh execution on every run, preventing stale data from influencing production scoring results. The code branches bypass both read and write operations to the cache/ directory entirely, forcing live PDF parsing and fresh GitHub API queries.

Key Source Files Controlling Cache Behavior

The caching system spans three critical files in the repository:

  • config.py: Defines the boolean DEVELOPMENT_MODE flag that toggles caching behavior across the application
  • score.py: Handles PDF processing and manages resumecache_*.json read/write operations
  • github.py: Wraps GitHub API calls and persists responses to githubcache_*.json files

Summary

  • Cache location: All cached results are stored in the repository's top-level cache/ directory when DEVELOPMENT_MODE is enabled
  • File patterns: Resume data uses resumecache_<pdf-basename>.json; GitHub data uses githubcache_<pdf-basename>.json or gh_githubcache_*.json
  • Directory creation: Both score.py and github.py use os.makedirs("cache", exist_ok=True) to ensure the directory exists before writing JSON data
  • Conditional operation: Caching only occurs when DEVELOPMENT_MODE is True; production runs bypass cache reads and writes entirely
  • Manual cleanup: Developers can delete specific JSON files from cache/ or remove the entire directory to force reprocessing of specific resumes or GitHub data

Frequently Asked Questions

What directory contains the cached results when DEVELOPMENT_MODE is enabled?

When DEVELOPMENT_MODE is set to True, the application stores all cached results in a cache/ directory at the repository root. This folder is created automatically on first write if it does not already exist, using os.makedirs with exist_ok=True.

How do I clear the cache to force reprocessing of PDFs?

Delete the specific JSON file in cache/ corresponding to the PDF basename (e.g., resumecache_candidate.pdf.json), or remove the entire cache/ directory. The application will regenerate the cache files on the next run when DEVELOPMENT_MODE is enabled, reprocessing the PDFs and refetching GitHub data as needed.

Why does the application cache GitHub API responses?

The github.py module caches API responses to avoid hitting rate limits during development and to speed up iteration cycles. When DEVELOPMENT_MODE is active, the code checks for githubcache_*.json files before making network requests, returning the cached JSON data immediately if available instead of querying the GitHub API again.

Does the cache affect production deployments?

No. When DEVELOPMENT_MODE is False (the production default), the code skips all cache existence checks and I/O operations as implemented in both score.py and github.py. This ensures every production run fetches fresh GitHub data and reprocesses PDFs without relying on potentially stale local files from the cache/ directory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →