How the Development Mode Caching Mechanism Works in InterviewStreet's Hiring Agent
When DEVELOPMENT_MODE is set to True in config.py, the interviewstreet/hiring-agent repository activates a lightweight file-based cache that stores parsed résumé data and GitHub API responses as JSON files, automatically bypassing expensive re-processing and HTTP calls on subsequent runs while silently invalidating corrupted entries.
The development mode caching mechanism is designed to accelerate iterative development and testing by avoiding redundant computation and external API requests. According to the interviewstreet/hiring-agent source code, this system operates exclusively when the environment flag is enabled, ensuring production deployments always fetch fresh data.
Resume Parsing Cache in score.py
The first component of the development mode caching mechanism handles expensive PDF parsing operations. In score.py, parsed résumé data is serialized to JSON and stored for rapid retrieval across multiple scoring runs.
Cache Location and File Naming
Cache files follow the pattern resumecache_*.json and reside in the cache/ directory. The filename is generated deterministically based on the input PDF filename, ensuring that identical résumés map to identical cache entries.
Reading and Writing Logic
At the start of the processing pipeline (around lines 227-260 in score.py), the code checks for existing cache entries before invoking the parser:
# development‑mode cache check
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
print(f"Loading cached data from {cache_filename}")
try:
cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
loaded_resume = JSONResume(**cached_data)
cache_loaded = True
except Exception as e:
print(f"⚠️ Warning: Invalid cache file {cache_filename}: {e}")
print("Ignoring cache and reprocessing PDF...")
os.remove(cache_filename)
When parsing completes successfully and no cache was loaded, the system writes the result:
if not cache_loaded:
# ... parsing logic ...
if resume_data_is_valid:
os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
Path(cache_filename).write_text(
json.dumps(resume_dict, ensure_ascii=False, indent=2), encoding="utf-8"
)
Cache Invalidation Strategy
If json.loads() raises an exception due to malformed JSON or disk corruption, the code immediately removes the offending file using os.remove(cache_filename) and falls back to re-processing the original PDF. This defensive pattern ensures that transient disk errors or manual file edits never break the development workflow.
GitHub API Response Cache in github.py
The second cached component prevents redundant HTTP requests to the GitHub REST API. In github.py, successful API responses are persisted to disk and served on subsequent identical requests.
Request Interception and Cache Keys
Before executing any HTTP call, the code generates a deterministic cache filename using _create_cache_filename(api_url, params) (lines 35-50). If a matching file exists, the network request is skipped entirely:
cache_filename = _create_cache_filename(api_url, params)
if DEVELOPMENT_MODE and os.path.exists(cache_filename):
print(f"Loading cached GitHub data from {cache_filename}")
try:
cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
if cached_data:
return 200, cached_data
except Exception as e:
print(f"⚠️ Warning: Error reading cache file {cache_filename}: {e}")
os.remove(cache_filename)
Persisting Successful Responses
Only successful responses (status_code == 200) are cached to avoid storing error payloads. The write operation creates the cache/ directory on demand:
if DEVELOPMENT_MODE and status_code == 200:
os.makedirs("cache", exist_ok=True)
Path(cache_filename).write_text(
json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8"
)
Configuration and Activation
The entire caching layer is gated by a single boolean flag defined in config.py. By default, DEVELOPMENT_MODE = True enables the cache, while setting it to False disables all file I/O operations related to caching. This centralized configuration ensures no cache code paths execute in production environments.
Production vs. Development Behavior
When DEVELOPMENT_MODE is disabled, none of the if DEVELOPMENT_MODE … conditionals evaluate to true. Consequently:
- Résumé parsing always processes the PDF from scratch
- GitHub API calls always execute fresh HTTP requests
- No cache files are read, written, or deleted
This strict separation guarantees that production deployments reflect current data, eliminating the risk of stale cache entries affecting candidate scoring results.
Summary
- The development mode caching mechanism is controlled by the
DEVELOPMENT_MODEflag inconfig.pyand is enabled by default for local development. - Resume parsing results are cached as
resumecache_*.jsonfiles inscore.py, with automatic re-processing triggered if the JSON is corrupted. - GitHub API responses are stored as
gh_githubcache_*.jsonfiles ingithub.py, intercepting network requests only when valid cache exists. - Cache filenames are generated deterministically from input parameters, ensuring consistent cache hits across identical inputs.
- Corrupted cache files are automatically deleted and regenerated, preventing development workflow interruptions.
- In production mode, all caching logic is bypassed entirely, ensuring fresh data and API responses.
Frequently Asked Questions
How do I disable the development mode caching mechanism?
Set DEVELOPMENT_MODE = False in config.py. When this flag is disabled, all cache checks in score.py and github.py are skipped, forcing the application to parse PDFs from scratch and execute fresh GitHub API requests on every run.
What happens if a cache file becomes corrupted?
The code handles corruption defensively. In both score.py and github.py, if json.loads() raises an exception when reading a cache file, the exception is caught, a warning is printed to the console, os.remove(cache_filename) deletes the corrupted file, and the system falls back to the original data source (PDF parsing or HTTP request).
Where are the cache files stored?
Cache files are stored in the cache/ directory at the project root. The directory is created automatically on demand via os.makedirs(..., exist_ok=True) when the first cache write occurs. Resume caches follow the pattern resumecache_*.json, while GitHub API caches use gh_githubcache_*.json.
Is the cache used in production deployments?
No. The development mode caching mechanism is explicitly development-only. When DEVELOPMENT_MODE is False (the recommended setting for production), none of the cache read or write operations execute, ensuring that production instances always process fresh résumé data and current GitHub API responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →