Where Are Cached Results Stored When DEVELOPMENT_MODE Is Enabled?
When DEVELOPMENT_MODE is set to True in the interviewstreet/hiring-agent repository, cached results are stored in the top-level cache/ directory as JSON files, specifically using resumecache_*.json for resume extractions and githubcache_*.json (or gh_githubcache_*.json) for GitHub API responses.
The hiring-agent application by Interview Street uses a development mode flag to optimize iteration speed during local testing. When DEVELOPMENT_MODE is enabled, the system avoids reprocessing PDFs and re-fetching GitHub data by persisting intermediate results to disk. Understanding the exact storage location and file patterns helps developers debug processing pipelines and manually clear stale cache entries.
Cache Storage Location and File Structure
When DEVELOPMENT_MODE is active, the application creates a cache/ directory at the repository root to store serialized intermediate data. This directory is generated on-demand using os.makedirs with exist_ok=True, ensuring the application does not fail if the folder already exists.
The system maintains two distinct cache types with specific file naming patterns:
Resume Extraction Cache Files
Located at cache/resumecache_<pdf-basename>.json, these files store the JSON output of parsed PDF resumes. According to score.py, after a PDF is processed, the extracted resume dictionary is serialized and written to this location. On subsequent runs, if the file exists, the application loads the JSON directly instead of re-parsing the PDF.
GitHub Data Cache Files
GitHub API responses are cached as cache/githubcache_<pdf-basename>.json or the generic cache/gh_githubcache_*.json pattern used by the GitHub helper module. As implemented in github.py, when the application queries the GitHub API, the response payload is saved to disk. Future executions read this cached data before making new network requests, reducing API rate limit consumption.
How the Caching Mechanism Works
The caching behavior is conditional on the DEVELOPMENT_MODE flag defined in config.py. When enabled, the code checks for existing cache files using os.path.exists() before proceeding with expensive operations.
Resume Processing Logic in score.py
In score.py, the resume extraction logic creates cache files using the following approach:
# score.py – Creating the resume cache
os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
Path(cache_filename).write_text(json.dumps(resume_dict, ensure_ascii=False))
When reading cached data, the code validates file existence before loading:
# Loading cached resume data (development mode only)
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
print(f"Loading cached data from {cache_filename}")
cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
loaded_resume = JSONResume(**cached_data)
GitHub API Logic in github.py
Similarly, github.py implements caching for network requests:
# github.py – Creating the GitHub cache
os.makedirs("cache", exist_ok=True)
Path(cache_filename).write_text(json.dumps(github_data, ensure_ascii=False))
The retrieval logic mirrors the resume cache pattern:
# Loading cached GitHub data (development mode only)
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
print(f"Loading cached GitHub data from {cache_filename}")
cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
return 200, cached_data
Production Behavior vs. Development Mode
When DEVELOPMENT_MODE is set to False, the application skips all os.path.exists() checks for cache files. This ensures a fresh execution on every run, preventing stale data from influencing production scoring results. The code branches bypass both read and write operations to the cache/ directory entirely, forcing live PDF parsing and fresh GitHub API queries.
Key Source Files Controlling Cache Behavior
The caching system spans three critical files in the repository:
config.py: Defines the booleanDEVELOPMENT_MODEflag that toggles caching behavior across the applicationscore.py: Handles PDF processing and managesresumecache_*.jsonread/write operationsgithub.py: Wraps GitHub API calls and persists responses togithubcache_*.jsonfiles
Summary
- Cache location: All cached results are stored in the repository's top-level
cache/directory whenDEVELOPMENT_MODEis enabled - File patterns: Resume data uses
resumecache_<pdf-basename>.json; GitHub data usesgithubcache_<pdf-basename>.jsonorgh_githubcache_*.json - Directory creation: Both
score.pyandgithub.pyuseos.makedirs("cache", exist_ok=True)to ensure the directory exists before writing JSON data - Conditional operation: Caching only occurs when
DEVELOPMENT_MODEisTrue; production runs bypass cache reads and writes entirely - Manual cleanup: Developers can delete specific JSON files from
cache/or remove the entire directory to force reprocessing of specific resumes or GitHub data
Frequently Asked Questions
What directory contains the cached results when DEVELOPMENT_MODE is enabled?
When DEVELOPMENT_MODE is set to True, the application stores all cached results in a cache/ directory at the repository root. This folder is created automatically on first write if it does not already exist, using os.makedirs with exist_ok=True.
How do I clear the cache to force reprocessing of PDFs?
Delete the specific JSON file in cache/ corresponding to the PDF basename (e.g., resumecache_candidate.pdf.json), or remove the entire cache/ directory. The application will regenerate the cache files on the next run when DEVELOPMENT_MODE is enabled, reprocessing the PDFs and refetching GitHub data as needed.
Why does the application cache GitHub API responses?
The github.py module caches API responses to avoid hitting rate limits during development and to speed up iteration cycles. When DEVELOPMENT_MODE is active, the code checks for githubcache_*.json files before making network requests, returning the cached JSON data immediately if available instead of querying the GitHub API again.
Does the cache affect production deployments?
No. When DEVELOPMENT_MODE is False (the production default), the code skips all cache existence checks and I/O operations as implemented in both score.py and github.py. This ensures every production run fetches fresh GitHub data and reprocesses PDFs without relying on potentially stale local files from the cache/ directory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →