What Files Are Cached and Where in Development Mode in Hiring-Agent
When DEVELOPMENT_MODE is set to True, the Hiring-Agent writes three categories of JSON cache files—resumecache_*.json, githubcache_*.json, and gh_githubcache_*.json—to a cache/ directory at the repository root, enabling instant reloading of parsed PDFs and GitHub API responses without re-executing expensive operations.
The interviewstreet/hiring-agent repository implements a local file-based caching system to speed up development workflows. By persisting parsed resume data and GitHub API responses to disk, the application eliminates redundant processing when iterating on scoring logic or testing candidate evaluations against the same inputs.
Cache File Types and Locations
Three distinct JSON cache files are generated when DEVELOPMENT_MODE=True:
- Resume extraction caches: Stored as
cache/resumecache_<pdf-basename>.jsonafter PDF parsing completes inscore.py - GitHub profile caches: Saved as
cache/githubcache_<pdf-basename>.jsonwhen a GitHub URL is discovered in a resume - Generic API caches: Written as
cache/gh_githubcache_<url-parts>[_<param-hash>].jsonfor every low-level GitHub request issued bygithub.py
All files reside under the cache/ directory, which is excluded from version control via .gitignore.
How Cache Retrieval Works
The application checks for existing cache files before executing expensive operations. If a valid cache exists, the JSON is loaded directly; if the file is corrupt or missing, the operation executes and the result is serialized to disk.
Resume and Profile Caching in score.py
In score.py, the system constructs cache_filename and github_cache_filename variables based on the input PDF's basename. When DEVELOPMENT_MODE is active and os.path.exists(cache_filename) returns True, the code loads the JSON directly into a JSONResume object, bypassing the PDF parser entirely.
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
print(f"Loading cached data from {cache_filename}")
cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
loaded_resume = JSONResume(**cached_data)
cache_loaded = True
After successful processing of uncached data, the code ensures the cache/ directory exists and writes the serialized dictionary:
if not cache_loaded:
# …process PDF…
os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
Path(cache_filename).write_text(
json.dumps(resume_dict, ensure_ascii=False, indent=2),
encoding="utf-8",
)
GitHub API Caching in github.py
The github.py module implements caching through the _create_cache_filename helper, which generates deterministic filenames from request URLs and optional query parameters. The github_request function first checks for the cached file, returns the stored JSON if present, or executes the HTTP request and persists the response.
cache_filename = _create_cache_filename(api_url, params)
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
print(f"Loading cached GitHub data from {cache_filename}")
cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
return 200, cached_data
# …perform HTTP request…
os.makedirs("cache", exist_ok=True)
Path(cache_filename).write_text(json.dumps(data), encoding="utf-8")
Enabling Development Mode
To activate file caching, set the environment variable before running the application:
# In the .env file (or environment)
DEVELOPMENT_MODE=True
Summary
- The
cache/directory at the repository root stores all development artifacts resumecache_*.jsonfiles contain parsed PDF resume data generated byscore.pygithubcache_*.jsonfiles store repository lists associated with specific resumesgh_githubcache_*.jsonfiles cache raw GitHub API responses fromgithub.py- Cache files are automatically invalidated and regenerated if corrupted
- The system only reads from cache when
DEVELOPMENT_MODE=True
Frequently Asked Questions
Where are cache files stored in Hiring-Agent?
All cache files are stored in a cache/ directory at the root of the repository. This location is consistent across all three cache types and is explicitly excluded from Git via .gitignore.
How does Hiring-Agent know when to use cached data?
The code checks the DEVELOPMENT_MODE environment variable and the existence of specific JSON files using os.path.exists(). If both conditions are met, the application loads the cached JSON directly; otherwise, it executes the expensive operation and writes the result to disk for future runs.
What happens if a cache file is corrupted?
If the JSON in a cache file is invalid or cannot be parsed, the application automatically removes the corrupt file and re-executes the original operation (PDF parsing or GitHub API call), then writes fresh data to the same location.
Can I disable caching while keeping development mode enabled?
No, caching is intrinsically tied to DEVELOPMENT_MODE. To force fresh data generation, you must either set DEVELOPMENT_MODE=False or manually delete specific files from the cache/ directory before execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →