# How to Log and Monitor TextFlow Pipeline Execution

> Learn to log and monitor TextFlow pipeline execution using its built-in logger. Track VQA, Textualizer, Reasoner, and Evaluation stages with timestamped files and colorized console output.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: how-to-guide
- Published: 2026-03-05

---

**TextFlow provides a lightweight, configurable logging system via `setup_logger()` in [`src/logger.py`](https://github.com/junyiye/textflow/blob/main/src/logger.py) that writes timestamped files and colorized console output for every pipeline stage (VQA, Textualizer, Reasoner, Evaluation).**

The `junyiye/textflow` repository includes a production-ready mechanism to log and monitor TextFlow pipeline execution across all stages. The system consists of three core components: a JSON configuration file that defines output directories and verbosity levels, a centralized logger factory that manages file and console handlers, and standardized integration points within each pipeline script. This architecture ensures every run generates isolated, queryable records suitable for both real-time monitoring and post-hoc debugging.

## Configure the Logging Environment

All logging behavior is controlled through [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) at the repository root. The `logging` stanza specifies the root directory for log files and the default verbosity threshold.

```json
{
  "logging": {
    "log_dir": "logs",
    "log_level": "info"
  }
}

```

- **`log_dir`**: Defines the base folder where pipeline scripts create dataset-specific subdirectories (e.g., `logs/flowvqa/`).  
- **`log_level`**: Sets the global threshold (`DEBUG`, `INFO`, `WARNING`, `ERROR`). The default is **INFO**, meaning DEBUG messages are suppressed unless you change this value to `debug`.

## Initialize the Logger Factory

The [`src/logger.py`](https://github.com/junyiye/textflow/blob/main/src/logger.py) module exports `setup_logger(log_file)`, which configures a **root logger** with dual handlers: one for persistent file storage and one for colorized terminal output. The factory prevents duplicate handlers by checking `root_logger.hasHandlers()` before attachment.

```python
def setup_logger(log_file):
    log_level = config["logging"]["log_level"].upper()
    # ensure directory exists

    log_dir = os.path.dirname(log_file)
    if not os.path.exists(log_dir):
        os.makedirs(log_dir)

    # file handler (plain text)

    file_handler = logging.FileHandler(log_file)
    file_handler.setLevel(getattr(logging, log_level))
    file_formatter = logging.Formatter("%(asctime)s - %(name)s - %(levelname)s - %(message)s")
    file_handler.setFormatter(file_formatter)

    # root logger – only add handlers once

    root_logger = logging.getLogger()
    if not root_logger.hasHandlers():
        root_logger.setLevel(getattr(logging, log_level))
        root_logger.addHandler(file_handler)

    # console handler with colour (uses ColorfulFormatter)

    console_handler = logging.StreamHandler()
    console_handler.setLevel(getattr(logging, log_level))
    console_formatter = ColorfulFormatter("%(asctime)s - %(name)s - %(levelname)s - %(message)s")
    console_handler.setFormatter(console_formatter)
    root_logger.addHandler(console_handler)

```

Key implementation details according to the source code:

- **File Handler**: Writes every record to the path supplied by the caller using the format `%(asctime)s - %(name)s - %(levelname)s - %(message)s`.
- **Console Handler**: Uses `ColorfulFormatter` to emit cyan (DEBUG), green (INFO), yellow (WARNING), and red (ERROR) messages for live monitoring.
- **Singleton Guard**: The `hasHandlers()` check ensures that repeated calls to `setup_logger` within the same process do not create duplicate log entries.

## Integrate Logging into Pipeline Stages

Each entry-point script—[`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py), [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), and [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py)—follows an identical four-step initialization pattern. This guarantees that every pipeline stage generates an isolated log file named with the dataset, model, and timestamp.

```python

# 1️⃣ Build a unique log file name (timestamp + model)

timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
log_file = os.path.join(
    config["logging"]["log_dir"], dataset,
    f"vqa_{model_name}_{timestamp}.log"
)

# 2️⃣ Initialise the logging system

setup_logger(log_file)

# 3️⃣ Get a module‑level logger

logger = logging.getLogger(__name__)

# 4️⃣ Emit progress information

logger.info("Starting the Vision Question Answering program...")
for arg, value in vars(args).items():
    logger.info(f"{arg}: {value}")
logger.info(f"Logs saved to {os.path.abspath(log_file)}")

```

To add logging to a new custom pipeline, replicate this pattern:

```python

# my_new_pipeline.py

from datetime import datetime
import os
import logging
from config import config
from logger import setup_logger

def run():
    timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
    log_file = os.path.join(
        config["logging"]["log_dir"], "my_pipeline",
        f"run_{timestamp}.log"
    )
    setup_logger(log_file)                     # initialise logger

    logger = logging.getLogger(__name__)        # module logger

    logger.info("Custom pipeline started")
    # ... your pipeline logic ...

    logger.debug("Intermediate variable x=%s", x)
    logger.warning("Potential issue detected")
    logger.error("Fatal error – aborting")
    logger.info("Custom pipeline finished")

```

Because each script constructs its own `log_file` path, you can audit the **VQA**, **Textualizer**, **Reasoner**, and **Evaluation** stages independently or aggregate them later for end-to-end tracing.

## Real-Time Monitoring Techniques

Once a pipeline is running, you can monitor progress through multiple channels:

- **Live Console View**: Run any pipeline script directly in your terminal. The `ColorfulFormatter` provides immediate visual feedback via color-coded severity levels.
- **File Tailing**: For detached or background jobs, follow the log stream using standard Unix tools:

```bash

# start the VQA job (non‑blocking)

python -m src.vqa --dataset flowvqa --model_name Qwen2-VL-7B &

# tail the generated log (adjust the timestamp)

tail -f logs/flowvqa/vqa_Qwen2-VL-7B_20240305_152030.log

```

- **Post-Run Analysis**: Log files contain structured timestamps, argument dictionaries, data-path resolutions, and final output locations. You can grep for specific severity levels to build automated alerts:

```bash
grep -E "ERROR|WARNING" logs/flowvqa/*.log > alerts.txt

```

## Summary

- **[`config.json`](https://github.com/junyiye/textflow/blob/main/config.json)** controls the global `log_dir` and `log_level` for all TextFlow components.
- **[`src/logger.py`](https://github.com/junyiye/textflow/blob/main/src/logger.py)** provides the `setup_logger()` factory, which wires a file handler and a colorized console handler to the root logger while preventing duplicate entries via `hasHandlers()`.
- **Pipeline scripts** ([`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py), [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py)) initialize logging with unique, timestamped file names at startup using the pattern `f"{pipeline}_{model}_{timestamp}.log"`.
- **Monitoring options** include live colorized console output, `tail -f` on specific log files, and automated parsing of severity levels for CI/CD integration.

## Frequently Asked Questions

### Where does TextFlow store pipeline log files?

By default, TextFlow writes logs to the directory specified in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) under `logging.log_dir` (default: `logs/`). Each pipeline run creates a subdirectory named after the dataset (e.g., `logs/flowvqa/`) containing files formatted as `{pipeline}_{model}_{timestamp}.log`.

### How do I change the logging verbosity in TextFlow?

Edit [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) and set `logging.log_level` to `debug`, `info`, `warning`, or `error`. The change takes effect immediately on the next run; `setup_logger()` reads this value at initialization and applies it to both the file and console handlers.

### Can I add logging to a custom pipeline module?

Yes. Import `setup_logger` from [`src/logger.py`](https://github.com/junyiye/textflow/blob/main/src/logger.py), construct a unique log file path using `datetime.now().strftime("%Y%m%d_%H%M%S")`, call `setup_logger(log_file)`, then retrieve a module logger via `logging.getLogger(__name__)`. All subsequent `logger.info()`, `logger.debug()`, or `logger.error()` calls will write to both your log file and the console.

### Why are my log entries appearing twice in TextFlow?

Duplicate entries occur if `setup_logger()` is called multiple times without the internal guard. The [`src/logger.py`](https://github.com/junyiye/textflow/blob/main/src/logger.py) implementation checks `root_logger.hasHandlers()` to avoid re-attaching handlers, but if you manually add handlers elsewhere in your code, you may circumvent this protection. Ensure you only call `setup_logger()` once per process.