How the Caching Mechanism Works in Development Mode in Hiring-Agent

During development, the Hiring-Agent project stores intermediate results as JSON files on disk, checking for cached data before performing expensive operations like PDF parsing or GitHub API calls when the DEVELOPMENT_MODE flag is enabled.

The interviewstreet/hiring-agent repository uses a file-based caching mechanism to accelerate iterative development workflows. When the DEVELOPMENT_MODE flag is set to True in config.py, the system persists parsed resume data and GitHub API responses to the cache/ directory. This allows developers to skip redundant processing and API calls on subsequent runs while maintaining deterministic output.

Resume Processing Cache in score.py

The resume evaluation pipeline in score.py implements a deterministic caching layer for PDF parsing results. This avoids the overhead of repeatedly extracting text and structure from the same resume files during development cycles.

Cache Path Construction and Lookup

Before parsing a PDF, the code constructs a cache filename based on the input file's basename. The pattern cache/resumecache_<pdf-basename>.json ensures each resume maps to a unique cache entry.

According to the source in score.py (lines 216-219), the path is built dynamically:

cache_filename = f"cache/resumecache_{os.path.basename(pdf_path)}.json"

When DEVELOPMENT_MODE is active, the system checks for this file's existence before invoking the PDF parser (lines 227-231). If found, the JSON is deserialized into a JSONResume object, bypassing the expensive parsing logic entirely.

Handling Corrupted Cache Files

The implementation includes robust error handling for invalid cache data. If JSON decoding fails, score.py logs a warning, deletes the corrupted file, and falls back to full PDF processing (lines 237-244). This prevents stale or malformed caches from blocking development workflows.

Writing Cache After Successful Parsing

Upon successful PDF extraction, the system creates the cache/ directory if needed and persists the structured resume data as JSON (lines 259-261). This write-back strategy ensures the next execution run loads instantly from disk rather than re-parsing the document.

GitHub API Response Cache in github.py

External API calls to GitHub are similarly cached to avoid rate limiting and network latency during development. The github.py module wraps all API requests with a deterministic caching layer.

Deterministic Filename Generation

The helper function _create_cache_filename (lines 18-25) generates cache keys by sanitizing the API URL and appending a hash of request parameters. This produces filenames like cache/gh_githubcache_<url-parts>[_<param-hash>].json, ensuring unique storage per endpoint and parameter combination.

Cache Retrieval and Fallback Logic

Before executing an HTTP request, the code checks for an existing cache file when DEVELOPMENT_MODE is enabled (lines 35-42). If present, the JSON response is loaded and returned immediately, simulating a successful API call without network overhead.

When cache files contain invalid JSON, the system logs a warning, removes the defective file, and proceeds with a live API request (lines 44-50). This graceful degradation ensures development isn't halted by cache corruption.

Persisting Successful Responses

After receiving a 200 OK response from GitHub, the response body is written to the cache directory (lines 104-107). Any exceptions during the write operation are logged but non-fatal, allowing the application to continue even if disk persistence fails (line 111).

Configuration and Activation

The caching behavior is controlled by the DEVELOPMENT_MODE boolean defined in config.py (line 6). By default, this is hard-coded to True for local development environments.


# From config.py line 6

DEVELOPMENT_MODE = True

Production deployments should override this flag to disable caching and ensure fresh data processing. This can be achieved by modifying the configuration file or injecting environment variables before application startup.

Practical Code Examples

The following examples demonstrate cache utilization in typical development scenarios.

Running the resume evaluator with automatic caching:

from score import main

# First run parses the PDF and writes cache/resumecache_jane_doe_resume.pdf.json

result = main("samples/jane_doe_resume.pdf")

# Subsequent runs load the JSONResume object from disk instantly

result_cached = main("samples/jane_doe_resume.pdf")

Fetching GitHub data with response caching:

from github import request_github_api

# Initial call contacts the live API and caches the response

status, data = request_github_api(
    api_url="https://api.github.com/users/interviewstreet",
    params=None,
)

# Second call reads from cache/gh_githubcache_api_github_com_users_interviewstreet.json

status_cached, data_cached = request_github_api(
    api_url="https://api.github.com/users/interviewstreet",
    params=None,
)

Disabling cache for production execution:


# Override the development flag to force fresh processing

export DEVELOPMENT_MODE=False
python -m score samples/jane_doe_resume.pdf

Summary

  • The development mode caching mechanism in Hiring-Agent relies on JSON file storage under the cache/ directory, activated by the DEVELOPMENT_MODE flag in config.py.
  • Resume parsing results are cached in score.py using filenames derived from PDF basenames, with automatic fallback to full parsing when caches are missing or corrupted.
  • GitHub API responses are cached in github.py using deterministic filenames generated from URL and parameter hashes, preventing redundant network requests.
  • Both implementations feature graceful degradation, deleting invalid cache files and proceeding with live processing when JSON decoding fails.
  • Cache writes occur after successful operations, ensuring subsequent development runs benefit from persisted intermediate state while maintaining data consistency.

Frequently Asked Questions

Where is the development mode cache stored in Hiring-Agent?

All cached data is stored as JSON files in the cache/ directory at the project root. Resume caches follow the pattern resumecache_<pdf-basename>.json, while GitHub API caches use the prefix gh_githubcache_ followed by URL-derived identifiers.

How does the system handle corrupted cache files?

When JSON decoding fails during cache loading, both score.py (lines 237-244) and github.py (lines 44-50) log a warning message, delete the corrupted file, and automatically fall back to live processing or API requests. This ensures that cache corruption never blocks development workflows.

Can I disable caching without modifying the source code?

While DEVELOPMENT_MODE is hard-coded to True in config.py, production deployments typically override this value via environment variables or configuration management tools. Setting this flag to False disables all cache lookups and writes, forcing fresh PDF parsing and API calls on every execution.

Does the cache mechanism affect production performance?

The caching layer is designed specifically for development mode and should be disabled in production by setting DEVELOPMENT_MODE = False. When disabled, the system bypasses all disk I/O related to caching, ensuring production runs process fresh data without filesystem overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →