How the Development Caching Mechanism Works in Hiring-Agent
During development, the Hiring-Agent project uses a file-based JSON caching system controlled by the DEVELOPMENT_MODE flag in config.py to store intermediate results and avoid repeated expensive operations.
The interviewstreet/hiring-agent repository implements a deterministic disk-based caching layer that speeds up iterative development by persisting parsed PDF resumes and GitHub API responses as JSON files. This development caching mechanism is governed by a single boolean flag that determines whether the system reads from existing cache files or performs fresh computation.
Resume PDF Processing Cache in score.py
The resume evaluation pipeline caches parsed PDF data to eliminate redundant parsing of identical documents across multiple runs.
Cache Path Construction
In score.py, the system constructs a deterministic filename based on the input PDF name. Lines 216-219 generate paths following the pattern cache/resumecache_<pdf-basename>.json, ensuring each resume maps to a unique cache entry.
Loading and Validation
When DEVELOPMENT_MODE is enabled and the cache file exists, lines 227-231 read the JSON and instantiate a JSONResume object directly. If JSON decoding fails, lines 237-244 catch the error, log a warning, delete the corrupt file, and trigger a full PDF re-parse.
Cache Persistence
After successful parsing, lines 259-261 create the cache/ directory if needed and write the serialized resume JSON to disk. This ensures subsequent executions load instantly from the cached representation rather than reprocessing the PDF.
GitHub API Response Cache in github.py
The GitHub integration module caches API responses to avoid rate limits and network latency during repeated development runs.
Deterministic Filename Generation
The _create_cache_filename function (lines 18-25 in github.py) builds cache keys by sanitizing the API URL and hashing request parameters, producing filenames like cache/gh_githubcache_<url-parts>[_<param-hash>].json.
Cache Lookup and Error Handling
Lines 35-42 check for cached responses when DEVELOPMENT_MODE is active, returning the stored JSON immediately if present. Lines 44-50 handle invalid cache files by logging warnings, removing the corrupted data, and falling back to live API requests.
Storing Successful Responses
For HTTP 200 responses, lines 104-107 persist the response body to the cache directory. Line 111 ensures that cache write failures are logged but non-blocking, preventing filesystem errors from interrupting the workflow.
Configuration and Control
The DEVELOPMENT_MODE flag resides in config.py (line 6) and defaults to True for local development. Setting this flag to False disables the caching mechanism entirely, forcing fresh PDF parsing and live API calls for production deployments.
# config.py
DEVELOPMENT_MODE = True # Set to False for production
Practical Usage Examples
Running Resume Evaluation with Cache
from score import main
# First run parses the PDF and writes cache/resumecache_jane_doe.json
result = main("samples/jane_doe_resume.pdf")
# Subsequent runs load from cache instantly
result = main("samples/jane_doe_resume.pdf")
Fetching GitHub Data with Cache
from github import request_github_api
# First call hits the live API and caches the response
status, data = request_github_api(
api_url="https://api.github.com/users/interviewstreet",
params=None,
)
# Second call reads from cache/gh_githubcache_api_github_com_users_interviewstreet.json
status, data = request_github_api(
api_url="https://api.github.com/users/interviewstreet",
params=None,
)
Disabling Cache for Production
export DEVELOPMENT_MODE=False
python -m score samples/jane_doe_resume.pdf
Summary
- The development caching mechanism relies on the
DEVELOPMENT_MODEflag inconfig.pyto toggle between cached and live data. - Resume parsing results are stored as JSON in
cache/resumecache_<name>.jsonviascore.py. - GitHub API responses are cached using deterministic filenames generated by
_create_cache_filenameingithub.py. - Corrupted cache files trigger automatic deletion and fallback to full processing, with explicit logging at each step.
- The system creates the
cache/directory automatically and handles write errors gracefully without stopping execution.
Frequently Asked Questions
How do I enable or disable the caching mechanism?
Set DEVELOPMENT_MODE to True or False in config.py (line 6). When True, the system checks for existing cache files before performing expensive operations; when False, it skips cache lookups and writes entirely.
Where are cache files stored?
All cache files reside in the cache/ directory at the project root. Resume caches follow the pattern resumecache_<pdf-basename>.json, while GitHub caches use gh_githubcache_<url-parts>[_<param-hash>].json.
What happens if a cache file is corrupted?
Both score.py (lines 237-244) and github.py (lines 44-50) implement validation logic. If JSON parsing fails, the code logs a warning, deletes the invalid file, and proceeds with fresh processing or live API calls.
Can I use this caching mechanism in production?
While technically possible, the repository defaults DEVELOPMENT_MODE to False for production deployments. Disabling the cache ensures you always process current resume data and fresh API responses, avoiding stale data issues.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →