Where Is Company Research Cached in the AI Job Search Framework?

Company research data is cached in a dedicated company_research/ directory at the repository root, with each entry stored as a JSON file named after the normalized company name and a 30-day TTL.

The AI Job Search framework implements a simple file-based caching system to avoid redundant API calls when researching employers. Understanding where this cache lives and how it operates is essential for debugging stale data, clearing entries, or extending the framework. This guide examines the cache location, structure, and lifecycle based on the actual source implementation in MadsLorentzen/ai-job-search.

Cache Location and File Structure

The framework stores all cached company research in a top-level directory called company_research/.

Each cache entry follows a predictable naming convention:

The normalization process converts company names to lowercase, replaces spaces with hyphens, and removes special characters. This ensures consistent file naming regardless of how the company name was originally entered.

Cache TTL and Freshness Rules

Cached entries remain valid for 30 days from the fetch date. After this window, the cache is considered stale and triggers a fresh research fetch.

The TTL logic appears in CHANGELOG.md (lines 69-79) and is enforced by checking the fetched_date field within each JSON file:

from datetime import datetime, timedelta

def is_cache_fresh(data: dict) -> bool:
    fetched = datetime.fromisoformat(data.get("fetched_date", "1970-01-01"))
    return datetime.utcnow() - fetched <= timedelta(days=30)

This 30-day window balances data freshness with API rate limit conservation.

How Commands Read and Write the Cache

Reading from Cache

Commands like /apply and /interview first attempt to load existing research before making external calls. The apply.md command (lines 122-130) implements this check during the "Research the Company" step:

import json
from pathlib import Path
from datetime import datetime, timedelta

def load_company_cache(name: str) -> dict | None:
    cache_dir = Path("company_research")
    cache_file = cache_dir / f"{name}.json"
    if not cache_file.is_file():
        return None
    data = json.loads(cache_file.read_text())
    fetched = datetime.fromisoformat(data.get("fetched_date", "1970-01-01"))
    if datetime.utcnow() - fetched > timedelta(days=30):
        return None  # stale cache

    return data

Writing to Cache

After completing fresh research, commands persist the results using a symmetric write operation:

def write_company_cache(name: str, research: dict) -> None:
    cache_dir = Path("company_research")
    cache_dir.mkdir(exist_ok=True)
    cache_file = cache_dir / f"{name}.json"
    research["fetched_date"] = datetime.utcnow().isoformat()
    cache_file.write_text(json.dumps(research, indent=2))

The fetched_date timestamp is injected at write time to enable subsequent TTL checks.

Cache Configuration and Security

The cache behavior is formally specified in .claude/skills/job-application-assistant/04-job-evaluation.md under the "Company Research Cache" section. This specification defines:

  • Cache directory location
  • 30-day TTL requirement
  • Data-vs-instruction semantics (cached data is still subject to verification rules)

Security restrictions in tools/security_guards.py (line 93) protect the cache path from unauthorized access patterns.

Git Exclusion Pattern

Cache files are explicitly excluded from version control via .gitignore (line 105):

company_research/*.json

This prevents sensitive research data or large JSON blobs from entering commits while preserving the directory structure in the repository.

Verification and Test Coverage

The framework includes dedicated test coverage in tests/test_company_research_cache.py that validates:

  • Correct cache directory path
  • Accurate TTL calculation
  • Proper read/write semantics
  • Verification rule inheritance

Running this test suite ensures cache functionality remains intact across framework updates.

Summary

  • Cache location: company_research/ directory at repository root
  • File format: Individual JSON files named {normalized-name}.json
  • TTL: 30 days based on fetched_date field
  • Commands affected: /apply (lines 122-130) and /interview (line 40)
  • Security: Protected path in security_guards.py
  • Version control: Excluded via .gitignore

Frequently Asked Questions

How do I clear the company research cache?

Delete the specific JSON file in company_research/ or remove the entire directory. The framework will regenerate entries on the next /apply or /interview command execution.

Can I adjust the 30-day TTL?

The TTL is hardcoded across multiple files. Modify timedelta(days=30) in your local copies of the command implementations, but note this will diverge from the framework specification in 04-job-evaluation.md.

Why is my cached research not being used?

Verify the fetched_date field exists in the JSON file and is within 30 days. Empty or malformed cache files trigger automatic regeneration. Check tests/test_company_research_cache.py for debugging patterns.

Is cached data trusted without verification?

No. According to the specification in 04-job-evaluation.md, cached data is treated as data, not instructions—verification rules still apply to any claims derived from the cache.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →