How to Log and Monitor TextFlow Pipeline Execution

TextFlow provides a lightweight, configurable logging system via setup_logger() in src/logger.py that writes timestamped files and colorized console output for every pipeline stage (VQA, Textualizer, Reasoner, Evaluation).

The junyiye/textflow repository includes a production-ready mechanism to log and monitor TextFlow pipeline execution across all stages. The system consists of three core components: a JSON configuration file that defines output directories and verbosity levels, a centralized logger factory that manages file and console handlers, and standardized integration points within each pipeline script. This architecture ensures every run generates isolated, queryable records suitable for both real-time monitoring and post-hoc debugging.

Configure the Logging Environment

All logging behavior is controlled through config.json at the repository root. The logging stanza specifies the root directory for log files and the default verbosity threshold.

{
  "logging": {
    "log_dir": "logs",
    "log_level": "info"
  }
}
  • log_dir: Defines the base folder where pipeline scripts create dataset-specific subdirectories (e.g., logs/flowvqa/).
  • log_level: Sets the global threshold (DEBUG, INFO, WARNING, ERROR). The default is INFO, meaning DEBUG messages are suppressed unless you change this value to debug.

Initialize the Logger Factory

The src/logger.py module exports setup_logger(log_file), which configures a root logger with dual handlers: one for persistent file storage and one for colorized terminal output. The factory prevents duplicate handlers by checking root_logger.hasHandlers() before attachment.

def setup_logger(log_file):
    log_level = config["logging"]["log_level"].upper()
    # ensure directory exists

    log_dir = os.path.dirname(log_file)
    if not os.path.exists(log_dir):
        os.makedirs(log_dir)

    # file handler (plain text)

    file_handler = logging.FileHandler(log_file)
    file_handler.setLevel(getattr(logging, log_level))
    file_formatter = logging.Formatter("%(asctime)s - %(name)s - %(levelname)s - %(message)s")
    file_handler.setFormatter(file_formatter)

    # root logger – only add handlers once

    root_logger = logging.getLogger()
    if not root_logger.hasHandlers():
        root_logger.setLevel(getattr(logging, log_level))
        root_logger.addHandler(file_handler)

    # console handler with colour (uses ColorfulFormatter)

    console_handler = logging.StreamHandler()
    console_handler.setLevel(getattr(logging, log_level))
    console_formatter = ColorfulFormatter("%(asctime)s - %(name)s - %(levelname)s - %(message)s")
    console_handler.setFormatter(console_formatter)
    root_logger.addHandler(console_handler)

Key implementation details according to the source code:

  • File Handler: Writes every record to the path supplied by the caller using the format %(asctime)s - %(name)s - %(levelname)s - %(message)s.
  • Console Handler: Uses ColorfulFormatter to emit cyan (DEBUG), green (INFO), yellow (WARNING), and red (ERROR) messages for live monitoring.
  • Singleton Guard: The hasHandlers() check ensures that repeated calls to setup_logger within the same process do not create duplicate log entries.

Integrate Logging into Pipeline Stages

Each entry-point script—src/vqa.py, src/textualizer.py, src/reasoner.py, and src/evaluation.py—follows an identical four-step initialization pattern. This guarantees that every pipeline stage generates an isolated log file named with the dataset, model, and timestamp.


# 1️⃣ Build a unique log file name (timestamp + model)

timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
log_file = os.path.join(
    config["logging"]["log_dir"], dataset,
    f"vqa_{model_name}_{timestamp}.log"
)

# 2️⃣ Initialise the logging system

setup_logger(log_file)

# 3️⃣ Get a module‑level logger

logger = logging.getLogger(__name__)

# 4️⃣ Emit progress information

logger.info("Starting the Vision Question Answering program...")
for arg, value in vars(args).items():
    logger.info(f"{arg}: {value}")
logger.info(f"Logs saved to {os.path.abspath(log_file)}")

To add logging to a new custom pipeline, replicate this pattern:


# my_new_pipeline.py

from datetime import datetime
import os
import logging
from config import config
from logger import setup_logger

def run():
    timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
    log_file = os.path.join(
        config["logging"]["log_dir"], "my_pipeline",
        f"run_{timestamp}.log"
    )
    setup_logger(log_file)                     # initialise logger

    logger = logging.getLogger(__name__)        # module logger

    logger.info("Custom pipeline started")
    # ... your pipeline logic ...

    logger.debug("Intermediate variable x=%s", x)
    logger.warning("Potential issue detected")
    logger.error("Fatal error – aborting")
    logger.info("Custom pipeline finished")

Because each script constructs its own log_file path, you can audit the VQA, Textualizer, Reasoner, and Evaluation stages independently or aggregate them later for end-to-end tracing.

Real-Time Monitoring Techniques

Once a pipeline is running, you can monitor progress through multiple channels:

  • Live Console View: Run any pipeline script directly in your terminal. The ColorfulFormatter provides immediate visual feedback via color-coded severity levels.
  • File Tailing: For detached or background jobs, follow the log stream using standard Unix tools:

# start the VQA job (non‑blocking)

python -m src.vqa --dataset flowvqa --model_name Qwen2-VL-7B &

# tail the generated log (adjust the timestamp)

tail -f logs/flowvqa/vqa_Qwen2-VL-7B_20240305_152030.log
  • Post-Run Analysis: Log files contain structured timestamps, argument dictionaries, data-path resolutions, and final output locations. You can grep for specific severity levels to build automated alerts:
grep -E "ERROR|WARNING" logs/flowvqa/*.log > alerts.txt

Summary

  • config.json controls the global log_dir and log_level for all TextFlow components.
  • src/logger.py provides the setup_logger() factory, which wires a file handler and a colorized console handler to the root logger while preventing duplicate entries via hasHandlers().
  • Pipeline scripts (src/vqa.py, src/textualizer.py, src/reasoner.py, src/evaluation.py) initialize logging with unique, timestamped file names at startup using the pattern f"{pipeline}_{model}_{timestamp}.log".
  • Monitoring options include live colorized console output, tail -f on specific log files, and automated parsing of severity levels for CI/CD integration.

Frequently Asked Questions

Where does TextFlow store pipeline log files?

By default, TextFlow writes logs to the directory specified in config.json under logging.log_dir (default: logs/). Each pipeline run creates a subdirectory named after the dataset (e.g., logs/flowvqa/) containing files formatted as {pipeline}_{model}_{timestamp}.log.

How do I change the logging verbosity in TextFlow?

Edit config.json and set logging.log_level to debug, info, warning, or error. The change takes effect immediately on the next run; setup_logger() reads this value at initialization and applies it to both the file and console handlers.

Can I add logging to a custom pipeline module?

Yes. Import setup_logger from src/logger.py, construct a unique log file path using datetime.now().strftime("%Y%m%d_%H%M%S"), call setup_logger(log_file), then retrieve a module logger via logging.getLogger(__name__). All subsequent logger.info(), logger.debug(), or logger.error() calls will write to both your log file and the console.

Why are my log entries appearing twice in TextFlow?

Duplicate entries occur if setup_logger() is called multiple times without the internal guard. The src/logger.py implementation checks root_logger.hasHandlers() to avoid re-attaching handlers, but if you manually add handlers elsewhere in your code, you may circumvent this protection. Ensure you only call setup_logger() once per process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →