How Development Mode and Caching Improve the Hiring Agent Development Workflow

Development mode uses a DEV_MODE flag to cache PDF extraction results and LLM responses, eliminating redundant computation and enabling instant iteration on scoring logic.

The Hiring Agent project from interviewstreet streamlines resume evaluation through an intelligent caching strategy. By toggling a simple configuration flag in config.py, developers can switch between production-grade processing and rapid local iteration. This development mode significantly reduces both API costs and computational overhead during the engineering workflow.

The DEV_MODE Configuration Flag

The foundation of the development workflow rests on a single global boolean defined in config.py. When DEV_MODE is set to True, the system activates caching mechanisms that persist intermediate processing results to disk.

This flag controls four critical behaviors that transform the agent from a production pipeline into a development tool:

  1. PDF extraction caching – Parsed resume content is serialized to JSON after the first run
  2. LLM response caching – Language model outputs are stored alongside extraction data
  3. Evaluation persistence – Scoring results append to resume_evaluations.csv for version-controlled tracking
  4. Fast feedback loops – Subsequent runs skip expensive computation and load from disk instead

How Caching Accelerates PDF Processing

In score.py (lines 226-269), the hiring agent implements a "run-once-store-once" pattern for resume parsing. When processing a PDF for the first time in development mode, the system extracts text using pymupdf and immediately serializes the result to cache/resumecache_<basename>.json.

On subsequent executions, the agent checks for this cached file before invoking the extractor:


# Inside score.py – snippet showing how the cache is used

if DEV_MODE and cache_path.exists():
    # Load cached extraction & LLM response

    data = json.load(open(cache_path))
else:
    # Perform full extraction and LLM call, then write cache

    data = extract_and_score(pdf_path)
    json.dump(data, open(cache_path, "w"))

This pattern eliminates repeated I/O operations and CPU-intensive text extraction, cutting pipeline execution time from minutes to milliseconds after the initial run.

Persisting LLM Responses and Evaluation Data

Development mode extends beyond PDF caching to preserve expensive LLM API calls. Each language model response is stored within the same JSON cache file, preventing redundant external API requests while you iterate on prompt templates or scoring heuristics.

Additionally, the evaluator writes structured results to resume_evaluations.csv after each run. This CSV accumulation creates a version-controlled audit trail of scoring changes, allowing engineers to compare model outputs across different code iterations without re-running the entire pipeline.

How Development Mode and Caching Enable Fast Feedback Loops

The combination of development mode and caching creates an environment optimized for rapid experimentation. Because heavy-weight steps—PDF parsing and LLM inference—execute only once, developers can modify lightweight components such as scoring logic, feature extraction algorithms, or prompt templates and see results instantly.

This architectural decision significantly reduces costs associated with paid LLM services during development. Engineers can run dozens of scoring variations against the same resume without incurring additional API charges or waiting for network latency.

Implementation Guide: Enabling Development Mode

Activate development mode by importing and setting the flag from config.py, or use environment variables for command-line execution:


# Enable development mode (typically via an environment variable or config flag)

from config import DEV_MODE
DEV_MODE = True   # or export DEV_MODE=1 before running scripts

# Running the full pipeline – first run will generate caches

from score import run_score
run_score("sample_resume.pdf")

# Subsequent runs pick up the cached data instantly

run_score("sample_resume.pdf")

For shell-based workflows:


# Command-line usage (development mode enabled via env var)

export DEV_MODE=1
python -m score path/to/resume.pdf

The cache directory populates automatically with resumecache_*.json files, while resume_evaluations.csv accumulates evaluation history for analysis.

Summary

  • Development mode is controlled by the DEV_MODE flag in config.py, toggling caching behavior across the pipeline
  • PDF extraction caching in score.py (lines 226-269) stores parsed results to cache/resumecache_<basename>.json, eliminating redundant pymupdf processing
  • LLM response caching prevents expensive API calls during iterative development of prompts and scoring logic
  • CSV persistence to resume_evaluations.csv creates version-controlled evaluation trails for regression testing
  • Fast iteration results from skipping heavy computation after the first run, reducing both time and API costs

Frequently Asked Questions

How do I enable development mode in Hiring Agent?

Set DEV_MODE = True in config.py or export the environment variable DEV_MODE=1 before running scripts. When enabled, the system in score.py automatically checks for and writes cache files to avoid redundant processing.

Where does Hiring Agent store cached extractions?

Cached data resides in the cache/ directory as JSON files named resumecache_<basename>.json, where <basename> corresponds to the original PDF filename. These files contain both the extracted resume text and the associated LLM response.

What files are generated when running in development mode?

The system generates two primary artifacts: JSON cache files in cache/resumecache_*.json containing extraction and LLM data, and a CSV file named resume_evaluations.csv that accumulates scoring results across multiple runs for version-controlled tracking.

How does caching reduce LLM API costs during development?

By storing LLM responses alongside PDF extraction data in the cache files, the system avoids repeated calls to external language-model APIs when re-processing the same resume. This allows unlimited iteration on scoring logic and prompt engineering without incurring additional API charges after the initial extraction.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →