How Development Mode Caching Works for Resumes and GitHub Data in the Hiring Agent Repository

When DEVELOPMENT_MODE is set to True in config.py, the hiring-agent repository stores parsed resume PDFs and GitHub API responses as JSON files in a local cache/ directory, bypassing expensive re-parsing and HTTP requests on subsequent runs.

The interviewstreet/hiring-agent project accelerates iterative development through an intelligent caching layer that activates exclusively in development environments. By toggling the DEVELOPMENT_MODE environment variable in config.py, developers can persist intermediate processing results for resume PDFs and GitHub API data, eliminating redundant I/O operations and network latency during debugging and feature development.

Resume PDF Caching Implementation

The resume processing pipeline in score.py implements a robust file-based caching mechanism that serializes extracted JSONResume objects to disk.

Cache File Naming Convention

The system generates deterministic cache filenames based on the input PDF's basename. In score.py (lines 215–217), the code constructs the path as follows:

cache_filename = f"cache/resumecache_{os.path.basename(pdf_path).replace('.pdf', '')}.json"

This creates files like cache/resumecache_john_doe.json for an input file named john_doe.pdf, ensuring each resume maintains its own isolated cache entry.

Loading from Cache with Validation

When DEVELOPMENT_MODE is enabled, the application checks for existing cache files before invoking the PDF extraction engine. The loading logic in score.py (lines 226–231) attempts deserialization directly into a JSONResume object:

if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached data from {cache_filename}")
    cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
    loaded_resume = JSONResume(**cached_data)

If the cache file is corrupted, contains invalid JSON, or fails schema validation, the system automatically purges the invalid file and falls back to full PDF processing (lines 237–240):

except Exception as e:
    print(f"⚠️ Warning: Invalid cache file {cache_filename}: {e}")
    os.remove(cache_filename)

Writing Extracted Data to Cache

Following a successful PDF extraction, the pipeline persists the JSONResume dictionary representation to disk only when development mode is active. The write operation in score.py (lines 259–266) ensures the cache directory exists before serialization:

if DEVELOPMENT_MODE:
    os.makedirs(os.path.dirname(cache_filename), exist_ok=True)
    Path(cache_filename).write_text(json.dumps(resume.dict()), encoding="utf-8")

GitHub Data Caching Mechanism

The GitHub integration module (github.py) implements an analogous caching strategy for API responses, preventing redundant HTTP requests during development iterations.

Deterministic Cache Key Generation

Cache filenames incorporate the request URL and query parameters to create unique identifiers for each endpoint. As implemented in github.py (lines 18–25):

filename = f"cache/gh_githubcache_{url_parts}_{param_str}.json"

This generates paths like cache/gh_githubcache_users_octocat_.json for a user profile request, ensuring distinct API calls do not collide.

Cache Retrieval and Fallback

The fetch logic prioritizes cached responses over live network calls. In github.py (lines 35–42), the system returns HTTP 200 status codes for cache hits to maintain interface consistency:

if DEVELOPMENT_MODE and os.path.exists(cache_filename):
    print(f"Loading cached GitHub data from {cache_filename}")
    cached_data = json.loads(Path(cache_filename).read_text(encoding="utf-8"))
    return 200, cached_data

Similar to the resume pipeline, corrupted GitHub cache files trigger automatic removal and fresh API requests (lines 44–49):

except Exception as e:
    print(f"⚠️ Warning: Error reading cache file {cache_filename}: {e}")
    os.remove(cache_filename)

Persisting API Responses

Successful API responses (status code 200) are serialized to JSON immediately after retrieval. The persistence logic in github.py (lines 104–107) writes only valid responses:

if DEVELOPMENT_MODE and status_code == 200:
    os.makedirs("cache", exist_ok=True)
    Path(cache_filename).write_text(json.dumps(data), encoding="utf-8")

Configuration and Environment Setup

The caching behavior is controlled by a single boolean flag in config.py. Setting DEVELOPMENT_MODE = True enables both the resume and GitHub caching layers simultaneously. When this flag is False, the application bypasses all cache read and write operations, ensuring production environments always process fresh data.

Practical Code Examples

Processing Resumes with Cache Acceleration

The following pattern demonstrates how subsequent calls to evaluate_resume leverage the cache automatically:

from score import evaluate_resume

pdf_path = "samples/jane_doe_resume.pdf"

# First execution parses the PDF and writes to cache/resumecache_jane_doe_resume.json

result1 = evaluate_resume(pdf_path)

# Second execution loads directly from cache, skipping PDF extraction

result2 = evaluate_resume(pdf_path)

Fetching GitHub Data with Local Persistence

Similarly, GitHub API calls utilize the cache layer transparently:

from github import fetch_github_data

# Initial call performs HTTP request and caches the response

status, data = fetch_github_data("https://api.github.com/users/octocat")

# Subsequent calls read from cache/gh_githubcache_users_octocat_.json

status, data = fetch_github_data("https://api.github.com/users/octocat")

Both examples require DEVELOPMENT_MODE = True in config.py to activate caching behavior.

Summary

  • Development mode caching is controlled by the DEVELOPMENT_MODE boolean in config.py and applies to both resume processing and GitHub API interactions.
  • Resume caches are stored as cache/resumecache_{filename}.json and contain serialized JSONResume objects, with automatic validation and cleanup of corrupted files.
  • GitHub caches use the naming pattern cache/gh_githubcache_{url}_{params}.json to store API responses, returning status code 200 for cache hits to maintain consistent interfaces.
  • Automatic recovery occurs when cache files are invalid—both modules delete corrupted files and fall back to fresh processing or HTTP requests.
  • On-demand directory creation ensures the cache/ folder is created automatically via os.makedirs(..., exist_ok=True) when writing the first cache entry.

Frequently Asked Questions

What happens if a cache file becomes corrupted?

Both score.py and github.py implement defensive error handling that catches deserialization exceptions, prints a warning message, deletes the invalid cache file using os.remove(), and proceeds with fresh data extraction or API requests. This ensures corrupted caches never block development workflow.

How do I completely disable caching for production deployments?

Set DEVELOPMENT_MODE = False in config.py. When this flag is disabled, the application skips all cache existence checks and write operations, forcing live PDF parsing and GitHub API calls on every execution.

Where exactly are cache files stored?

All cache files reside in a cache/ directory relative to the execution path. Resume caches follow the pattern cache/resumecache_*.json, while GitHub caches use cache/gh_githubcache_*.json. The directories are created automatically on first write via os.makedirs("cache", exist_ok=True).

Does caching bypass JSONResume schema validation?

No. When loading from cache in score.py, the system instantiates a JSONResume(**cached_data) object, which enforces the Pydantic schema validation. If the cached JSON violates the schema, the instantiation raises an exception, triggering cache deletion and re-processing of the original PDF.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →