How Company Research is Cached in the AI Job Search Framework: A 30-Day TTL Deep Dive
The AI Job Search framework stores company research results in JSON files within a company_research/ directory and checks this cache for 30 days before performing fresh web scrapes during the /apply and /interview workflows.
The MadsLorentzen/ai-job-search repository implements a lightweight file-based caching mechanism to eliminate redundant API calls and web scraping when researching employers. This system ensures that repeated queries for the same company—whether preparing a cover letter or interview materials—leverage previously gathered intelligence while maintaining strict data freshness guarantees.
Understanding the File-Based Cache Architecture
The framework treats company research caching as a data-only persistence layer, storing immutable JSON artifacts that contain no executable instructions. This design separates cached intelligence from application logic while providing rapid lookups across workflow commands.
Cache Storage Location and Naming Convention
Research outputs reside in the company_research/ directory at the repository root. Each employer receives a dedicated JSON file named after the normalized company string.
Normalization follows a strict transformation:
- Convert to lowercase
- Replace spaces with underscores
For example, "Acme Corporation" becomes company_research/acme_corporation.json. The .gitignore file explicitly excludes this directory from version control, ensuring sensitive research data never enters the Git history.
The 30-Day Time-To-Live (TTL) Policy
Every cache file includes a mandatory fetched_date field containing an ISO 8601 timestamp. According to the cache schema documented in .claude/skills/job-application-assistant/04-job-evaluation.md (lines 197-227), entries older than 30 days are considered stale and trigger fresh research passes.
This TTL balance minimizes API costs while ensuring outdated company information—such as recent funding rounds or leadership changes—does not persist indefinitely in generated application materials.
Cache Integration in Workflow Commands
The framework checks the company research cache at the entry point of both primary workflows. This early-exit pattern prevents unnecessary computation when valid cached data exists.
The /apply Command Cache Check
In .claude/commands/apply.md (lines 122-130), the /apply command first attempts to load cached research before initiating web scrapes. If the cache hit succeeds and the fetched_date is within the 30-day window, the command uses this data as the foundation for cover letter generation.
Stale or missing cache files trigger a full research pass, after which the framework persists the new findings to company_research/<normalized-name>.json.
The /interview Command Cache Check
Similarly, .claude/commands/interview.md (lines 40-45) implements identical cache-checking logic for interview preparation workflows. The command queries the same company_research/ directory and applies the same TTL validation rules, ensuring consistency across the application pipeline.
Both commands maintain the "verify before quoting" rule even on cache hits, meaning cached data undergoes validation before inclusion in final output documents.
Implementing the Cache Logic in Python
The underlying implementation uses standard library modules to handle JSON serialization and filesystem operations. The following pattern demonstrates the cache interaction as implemented in the framework:
import json
import os
import datetime
from pathlib import Path
CACHE_DIR = Path("company_research")
TTL_DAYS = 30
def _cache_path(company_name: str) -> Path:
normalized = company_name.lower().replace(" ", "_")
return CACHE_DIR / f"{normalized}.json"
def load_company_cache(company_name: str):
"""Return cached data if it exists and is fresh, otherwise None."""
cache_file = _cache_path(company_name)
if not cache_file.is_file():
return None
data = json.loads(cache_file.read_text())
fetched = datetime.datetime.fromisoformat(
data.get("fetched_date", "1970-01-01")
)
if (datetime.datetime.utcnow() - fetched).days > TTL_DAYS:
return None # stale cache
return data # fresh cache
def write_company_cache(company_name: str, research_dict: dict):
"""Persist research results to the cache with a freshness timestamp."""
cache_file = _cache_path(company_name)
research_dict["fetched_date"] = datetime.datetime.utcnow().isoformat()
cache_file.parent.mkdir(parents=True, exist_ok=True)
cache_file.write_text(json.dumps(research_dict, indent=2))
Loading and Validating Cached Data
The load_company_cache function encapsulates the lookup logic. It constructs the file path using _cache_path, verifies file existence, and performs the TTL calculation by comparing the current UTC time against the stored fetched_date.
Returning None for missing or stale entries provides a simple interface for upstream commands to determine whether fresh research is required.
Writing Fresh Research to Cache
The write_company_cache function handles persistence. It ensures the company_research/ directory exists, injects the current timestamp into the research dictionary, and writes formatted JSON to disk. This atomic write pattern prevents partial cache corruption during interrupted research operations.
Cache Schema Documentation
The complete caching specification resides in .claude/skills/job-application-assistant/04-job-evaluation.md (lines 197-227). This documentation defines:
- The storage directory path (
company_research/) - The 30-day TTL constraint
- The required
fetched_datefield format - Company name normalization rules
The feature introduction is recorded in CHANGELOG.md (lines 111-121), which documents the initial implementation of the 30-day TTL mechanism.
Summary
- The AI Job Search framework caches company research in
company_research/<normalized-name>.jsonfiles to prevent redundant web scraping - A strict 30-day TTL policy ensures data freshness while minimizing API usage
- Both
/applyand/interviewcommands check the cache before initiating fresh research passes - The cache schema requires an ISO-formatted
fetched_datefield for expiration validation - Cached data remains subject to the "verify before quoting" rule, maintaining output quality even when utilizing stored intelligence
Frequently Asked Questions
How does the framework handle company names with special characters?
The normalization logic converts company names to lowercase and replaces spaces with underscores, as implemented in the _cache_path function. Special characters outside this scope are preserved in the filename, though the framework primarily targets alphanumeric normalization for filesystem compatibility.
What happens when the 30-day TTL expires?
When load_company_cache detects a fetched_date older than 30 days, it returns None. This signals the calling command—whether /apply or /interview—to perform a fresh research pass. Upon completion, write_company_cache overwrites the stale file with updated data and a new timestamp.
Is the company research cache committed to Git?
No. The repository's .gitignore explicitly excludes the company_research/ directory. This prevents sensitive employer research data and personally identifiable information from entering version control while allowing the cache to persist locally across workflow executions.
Can the TTL be customized for specific companies?
Currently, the framework enforces a uniform 30-day TTL defined in the TTL_DAYS constant and documented in the job-evaluation skill file. All companies share this expiration window regardless of data volatility, ensuring predictable cache behavior across the MadsLorentzen/ai-job-search codebase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →