Security Considerations for Using the Hiring-Agent: Protecting PII in Resume Processing Pipelines

The hiring-agent repository processes sensitive résumé PDFs and external API credentials, requiring strict environment variable management, disabled development mode in production, scoped GitHub tokens, and sanitized CSV outputs to prevent data leakage and injection attacks.

The hiring-agent is an open-source pipeline developed by InterviewStreet that ingests résumé PDFs, enriches candidate profiles with GitHub data, and leverages Large Language Models (LLMs) for automated evaluation. Because it handles personally identifiable information (PII) and third-party API secrets, understanding the security considerations for using the hiring-agent is critical before deploying it in any environment. The codebase exposes several configurable security controls through environment variables and mode flags that, if misconfigured, could lead to credential exposure or unauthorized data retention on disk.

Environment Variables and Secret Management

The application relies on sensitive configuration values that must be supplied via a .env file. According to the .env.example template, these include LLM_PROVIDER, DEFAULT_MODEL, GEMINI_API_KEY, and the optional GITHUB_TOKEN.

If these values are leaked—such as by accidentally committing the .env file to version control—an attacker could abuse your Gemini API quota, exhaust GitHub rate limits, or access private repository data. While config.py manages the DEVELOPMENT_MODE flag that controls caching behavior, it does not read the environment file directly; instead, the runtime (e.g., python-dotenv) handles this loading. This separation means config.py assumes the environment is already sanitized at startup.

Development Mode and Data Caching Risks

In config.py (lines 5-7), the DEVELOPMENT_MODE boolean determines whether the pipeline persists intermediate data to disk. When set to True, the system writes raw JSON files under cache/ and appends evaluation results to resume_evaluations.csv.

These files contain unredacted résumé text and extracted GitHub information. Leaving development mode enabled in production exposes PII on the filesystem and creates risks of accidental commits to version control. For production deployments, explicitly override this flag or ensure the environment sets DEVELOPMENT_MODE=False.


# production_config.py

from hiring_agent.config import DEVELOPMENT_MODE

# Override the flag at import time

DEVELOPMENT_MODE = False

External LLM Provider Data Transmission

The models.py module implements two provider wrappers: OllamaProvider for local inference and GeminiProvider for Google's cloud API. The prompt.py module selects the active provider based on the LLM_PROVIDER environment variable.

When configured to use Gemini (LLM_PROVIDER=gemini), the entire résumé content is transmitted to Google's API, subjecting the data to third-party data-retention policies and potential transit interception. Using OllamaProvider keeps all processing within your trusted runtime environment, eliminating network transmission risks for sensitive candidate data.

GitHub Token Scope and Rate Limiting

The github.py module retrieves the GITHUB_TOKEN from the environment to increase API rate limits beyond anonymous thresholds. However, an over-privileged token could be abused to enumerate private repositories, read sensitive user data, or trigger actions on behalf of the token owner.

You should scope tokens to public read-only access exclusively, avoiding any repo or workflow scopes. The following snippet demonstrates safe token usage in github.py, restricting calls to public endpoints only:

headers = {"Authorization": f"Bearer {GITHUB_TOKEN}"} if GITHUB_TOKEN else {}

# Only call public endpoints; avoid any POST/DELETE requests.

response = httpx.get("https://api.github.com/users/{username}", headers=headers)

PDF Parsing and Dependency Security

The pymupdf_rag.py module relies on PyMuPDF to convert PDFs to Markdown text. PDF parsers are historically vulnerable to memory-corruption attacks through maliciously crafted documents. You must pin PyMuPDF to a recent, patched version in requirements.txt and monitor CVE advisories for the library.

Additionally, the repository pulls all dependencies from PyPI, making it susceptible to supply-chain attacks if a malicious version of a dependency is published. Use hash-based verification (e.g., pip-hash) or a lockfile for production deployments to ensure dependency integrity.

Output Sanitization and Access Control

The score.py module orchestrates the final evaluation and writes results to resume_evaluations.csv. This CSV includes raw résumé fields that, if imported into spreadsheet applications without sanitization, could execute formula injection attacks via cells starting with =, +, -, or @.

When processing outputs for downstream systems, implement sanitization logic to neutralize formula injectors:

def escape_csv(value: str) -> str:
    """Prevent CSV-formula injection by prefixing a single quote."""
    if isinstance(value, str) and value and value[0] in ("=", "+", "-", "@"):
        return f"'{value}"
    return value

# When writing a row:

row = {k: escape_csv(str(v)) for k, v in original_row.items()}
writer.writerow(row)

As a command-line utility, the tool lacks built-in authentication. If exposed as a web service, implement authentication and rate limiting to prevent mass-scraping of candidate data.

Secure Configuration Checklist

Follow these steps to harden your deployment:

  1. Never commit real .env files – version control only the .env.example template.
  2. Disable development mode in production by setting DEVELOPMENT_MODE=False in config.py or via environment variables.
  3. Prefer local LLMs (Ollama) for on-premises processing to avoid sending PII to external APIs.
  4. Scope GitHub tokens to public read-only access and rotate them regularly.
  5. Pin dependencies to known-good versions with hash verification in requirements.txt.
  6. Secure the cache directory with restrictive filesystem permissions (e.g., chmod 700 cache/).
  7. Sanitize CSV outputs before importing into spreadsheet applications.

Summary

  • Environment variables containing API keys must be kept out of version control and loaded securely at runtime.
  • Development mode in config.py writes sensitive PII to disk and must be disabled for production use.
  • External LLM calls to Gemini transmit candidate data to third-party servers, while local Ollama models preserve data privacy.
  • GitHub tokens should be scoped to public read-only access to prevent repository enumeration or unauthorized actions.
  • PDF parsing via PyMuPDF requires patched dependencies to mitigate memory-corruption vulnerabilities.
  • CSV outputs from score.py require sanitization to prevent formula injection attacks in downstream applications.

Frequently Asked Questions

What happens if I leave DEVELOPMENT_MODE enabled in production?

When DEVELOPMENT_MODE remains True in config.py, the pipeline writes intermediate JSON files to cache/ and appends evaluation data to resume_evaluations.csv. These files contain raw résumé text and GitHub profile data, creating a persistent PII repository on disk that could be leaked through backups, file permissions, or accidental commits.

Is it safe to use Google Gemini with candidate résumés?

Using LLM_PROVIDER=gemini routes all résumé content through models.GeminiProvider to Google's API, transmitting PII to a third-party service subject to their data-retention policies. For maximum privacy, configure LLM_PROVIDER=ollama to process documents locally within your infrastructure, avoiding external network transmission entirely.

How should I protect my GitHub token when using the hiring-agent?

Store the GITHUB_TOKEN in your .env file only, never commit it to git, and generate a token with minimal scope—specifically, public read-only access without repo, workflow, or delete_repo permissions. The github.py module uses this token only to increase rate limits for public API calls, so elevated privileges introduce unnecessary risk.

Can CSV outputs from score.py contain malware?

While the CSV itself contains text data, maliciously crafted cell values beginning with =, +, -, or @ can execute formulas when opened in spreadsheet software, potentially leading to data exfiltration or command execution. Always sanitize outputs using a function like escape_csv() before importing resume_evaluations.csv into Excel, LibreOffice, or similar applications.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →