How to Add Custom Logging to Troubleshoot Issues in the Hiring-Agent Pipeline
You can add custom logging to the interviewstreet/hiring-agent pipeline by centralizing the configuration in a new logging_config.py module, replacing the existing logging.basicConfig call in score.py with hierarchical loggers, and injecting contextual metadata via the extra parameter to trace data flow across PDF extraction and LLM evaluation steps.
The interviewstreet/hiring-agent repository orchestrates a multi-stage pipeline that processes resumes, extracts text from PDFs, and evaluates candidates via LLM calls. To effectively debug failures in this chain, you need visibility into each stage. The current codebase already uses Python’s standard logging module—module-level loggers are instantiated with logging.getLogger(__name__) and a basic format is configured in score.py—but the setup lacks granularity for production troubleshooting.
Centralize Logging Configuration in a Dedicated Module
The existing setup in score.py configures logging via logging.basicConfig (approximately lines 21–26). To gain fine-grained control over formatting, levels, and output destinations, move this logic to a new logging_config.py file. This establishes a single source of truth for all logging behavior across the codebase.
# logging_config.py
import logging
from logging.handlers import RotatingFileHandler
from config import DEVELOPMENT_MODE
# Dynamic level based on environment
log_level = logging.DEBUG if DEVELOPMENT_MODE else logging.INFO
# Formatter that captures standard fields
formatter = logging.Formatter(
"%(asctime)s - %(name)s - %(levelname)s - %(message)s"
)
# Console output for development
console_handler = logging.StreamHandler()
console_handler.setFormatter(formatter)
# Persistent file output with rotation (5 MB per file, 3 backups)
file_handler = RotatingFileHandler(
"logs/hiring_agent.log", maxBytes=5_000_000, backupCount=3
)
file_handler.setFormatter(formatter)
# Apply globally
logging.basicConfig(
level=log_level,
handlers=[console_handler, file_handler],
)
def get_logger(name: str) -> logging.Logger:
"""Convenience factory for hierarchical logger names."""
return logging.getLogger(name)
After creating this module, remove the logging.basicConfig call from score.py and import the centralized configuration instead.
Implement Hierarchical Loggers for Pipeline Components
Rather than using __name__ inconsistently, adopt a hierarchical naming convention such as pipeline.<component>. This structure allows you to selectively enable DEBUG mode for specific subsystems (e.g., pipeline.pdf) while keeping others at INFO.
In score.py, replace the existing logger initialization:
# score.py
from logging_config import get_logger
logger = get_logger("pipeline.score")
# Previous basicConfig call removed; now handled centrally
Apply the same pattern to other core modules:
# pdf.py
from logging_config import get_logger
logger = get_logger("pipeline.pdf")
# evaluator.py
from logging_config import get_logger
logger = get_logger("pipeline.evaluator")
This hierarchy enables runtime filtering—for example, setting logging.getLogger("pipeline.pdf").setLevel(logging.DEBUG) to isolate text extraction issues without cluttering logs from the scoring module.
Inject Contextual Data with Extra Fields
To trace data through the pipeline, pass contextual identifiers using the extra parameter. This captures metadata like resume file names, GitHub usernames, or character counts without polluting the main log message.
In pdf.py, instrument the extraction function:
def extract_text_from_pdf(self, pdf_path: str) -> Optional[str]:
logger.debug("Starting PDF extraction", extra={"resume": pdf_path})
try:
# extraction logic
...
except Exception:
logger.exception("Failed to read PDF", extra={"resume": pdf_path})
return None
In evaluator.py, log the LLM request context:
def evaluate_resume(self, resume_text: str) -> EvaluationData:
logger.debug("Sending evaluation request", extra={"chars": len(resume_text)})
try:
# LLM call logic
...
except Exception:
logger.exception("LLM evaluation failed")
raise
To surface these fields in your output, update the formatter in logging_config.py to include %(resume)s or %(chars)s as appropriate, or use a custom logging.Formatter subclass that handles dynamic keys.
Configure Rotating File Handlers for Persistent Debugging
Intermittent failures in the hiring-agent pipeline require historical log analysis. The RotatingFileHandler configured in logging_config.py ensures logs persist across runs and automatically rotate when files exceed 5 MB, preventing disk exhaustion while retaining the last 3 backups for forensics.
Ensure the logs/ directory exists before runtime, or add directory creation logic to logging_config.py:
import os
os.makedirs("logs", exist_ok=True)
Wire Logging to Configuration Flags
The config.py module defines a DEVELOPMENT_MODE flag that currently toggles caching behavior. Extend this to control log verbosity by importing it into logging_config.py (as shown in the first code block) to set log_level dynamically. This ensures verbose DEBUG output is automatically enabled in development environments while production remains at INFO or WARNING levels.
Summary
- Centralize configuration by moving
logging.basicConfigfromscore.pyto a newlogging_config.pymodule that defines formatters, handlers, and global levels. - Use hierarchical logger names (e.g.,
pipeline.pdf,pipeline.evaluator) to enable targeted debugging of specific pipeline stages. - Pass contextual metadata via the
extraparameter inlogger.debug()andlogger.exception()calls to correlate log entries with specific resumes or evaluation steps. - Persist logs using
RotatingFileHandlerto capture intermittent failures across multiple runs without manual intervention. - Leverage
DEVELOPMENT_MODEfromconfig.pyto automatically toggle betweenDEBUGandINFOlevels based on the environment.
Frequently Asked Questions
How do I view the extra contextual fields in my log output?
By default, Python’s standard formatter ignores keys passed via the extra parameter. To display them, modify the format string in logging_config.py to explicitly include the keys you use most frequently, such as %(resume)s or %(chars)s. Alternatively, implement a custom logging.Formatter subclass that iterates over the extra dictionary and appends it to the message.
Why should I replace logging.getLogger(__name__) with hierarchical names like pipeline.pdf?
While __name__ works, it couples your logger structure to the file system layout. Using explicit hierarchical names like pipeline.pdf decouples the logging namespace from the module structure, making it easier to reorganize code files later. It also allows you to configure logging levels for entire subtrees (e.g., pipeline.*) via a single configuration line.
What is the difference between logger.error() and logger.exception()?
Use logger.exception() exclusively inside exception handlers (typically within an except block). This method automatically captures the full exception traceback and stack trace in the log output, which is essential for debugging LLM evaluation failures in evaluator.py. Use logger.error() for error conditions that do not involve an active exception.
How can I prevent log files from growing indefinitely in production?
The RotatingFileHandler configured in logging_config.py automatically rotates files when they reach the maxBytes threshold (5 MB in the example) and maintains only backupCount historical files (3 backups). Older files are automatically deleted, ensuring disk usage remains bounded even during high-volume processing runs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →