# System Health Monitoring and Error Recovery in Embedded Radar Systems: Inside the AERIS-10 Open Source Radar

> Explore system health monitoring and error recovery in embedded radar with AERIS-10. Discover its three-layer architecture for unattended field operation.

- Repository: [NawfalMotii79/PLFM_RADAR](https://github.com/NawfalMotii79/PLFM_RADAR)
- Tags: internals
- Published: 2026-08-20

---

**The AERIS-10 embedded radar platform implements a three-layer health monitoring and error recovery architecture across its **FPGA**, **STM32 microcontroller**, and **Python GUI**, using Qt signals for real-time fault propagation and automatic hardware resets to maintain unattended field operation.**

System health monitoring and error recovery in embedded radar systems demand tight coordination between **FPGA** fabric, **STM32 microcontroller** firmware, and **Python GUI** dashboards. The PLFM_RADAR repository by NawfalMotii79 demonstrates how the AERIS-10 platform achieves this reliability by combining Python-based worker threads, low-level FTDI protocol wrappers, and cross-layer reset logic. The PyQt application entry point in [`9_Firmware/9_3_GUI/v7/GUI_V7_PyQt.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/9_Firmware/9_3_GUI/v7/GUI_V7_PyQt.py) wires the workers to the dashboard at startup, creating a unified pipeline that detects, logs, and recovers from faults without manual intervention.

## System Health Monitoring and Error Recovery Architecture

### Worker-Level Error Counters and Qt Signals

In [`9_Firmware/9_3_GUI/v7/workers.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/9_Firmware/9_3_GUI/v7/workers.py), each data-acquisition worker such as `RadarDataWorker` and `GPSDataWorker` tracks internal failures through **error counting** via `self._error_count`. Whenever an exception is caught—commonly `ValueError` or `struct.error`—the worker increments the counter, emits an `errorOccurred(str)` **Qt signal**, and writes a timestamped entry via `logger.error`.

This dual reporting mechanism ensures that both the UI and persistent logs receive the fault data simultaneously. The `stats` signal broadcasts a dictionary containing `"errors"` and `"frames"` keys every second, giving the dashboard a real-time health feed.

```python
class RadarDataWorker(QThread):
    errorOccurred = pyqtSignal(str)      # <-- emitted on error

    stats = pyqtSignal(dict)             # <-- periodic status updates

    def run(self):
        while not self.isInterruptionRequested():
            try:
                data = self._read_radar_frame()
                self._process(data)
            except (ValueError, struct.error) as e:
                self._error_count += 1
                self.errorOccurred.emit(str(e))
                logger.error(f"RadarDataWorker error: {e}")

            # Emit health stats every second

            self.stats.emit({"errors": self._error_count,
                            "frames": self._frames_processed})

```

### Centralized Dashboard Handling

The main dashboard defined in [`9_Firmware/9_3_GUI/v7/dashboard.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/9_Firmware/9_3_GUI/v7/dashboard.py) connects each worker’s `errorOccurred` signal to a centralized slot named `_on_worker_error`. This slot provides **centralized error handling** that updates the status bar, writes to the unified log, and can trigger immediate recovery actions such as restarting the worker.

```python
class RadarDashboard(QWidget):
    def __init__(self):
        ...
        self._radar_worker.errorOccurred.connect(self._on_worker_error)

    def _on_worker_error(self, msg: str):
        logger.error(f"Worker error: {msg}")
        # Immediate UI feedback

        self.statusBar().showMessage(f"Error: {msg}", 5000)
        # Automatic recovery if error count grows

        if self._radar_worker._error_count > 10:
            self._restart_acquisition()

```

## Hardware-Level Fault Isolation in Embedded Radar Drivers

### Graceful Degradation for FTDI USB Communication

Low-level I/O in [`9_Firmware/9_3_GUI/v7/radar_protocol.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/9_Firmware/9_3_GUI/v7/radar_protocol.py) implements **hardware-level fault isolation** by wrapping every FT2232H USB transaction in a `try/except` block. Rather than crashing the acquisition thread, the driver catches the exception, logs an informative message such as `log.error("FT2232H open failed: …")`, and returns `None` to the caller, enabling **graceful degradation**. This design lets higher-level code decide between an immediate retry, a pipeline abort, or a hardware reset.

```python
def read_bytes(self, length: int) -> Optional[bytes]:
    try:
        raw = self._ftdi.readbytes(length)
        return bytes(raw)
    except Exception as e:
        log.error(f"FT2232H read error: {e}")
        return None      # Caller can decide to retry or abort

```

## Automatic Recovery Loops and Cross-Layer Resets

### Threshold-Based Hardware Restart Logic

Beyond immediate slot reactions, the GUI’s main loop leverages **automatic recovery loops** that periodically poll the `stats` signal emitted by each worker. When the reported error count exceeds a configurable threshold, the dashboard executes a more extensive **cross-layer reset**: it stops the current radar acquisition, resets the STM32 microcontroller via [`9_Firmware/9_3_GUI/v7/hardware.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/9_Firmware/9_3_GUI/v7/hardware.py), and restarts the acquisition pipeline.

This strategy bridges the Python GUI and the embedded STM32, ensuring that transient communication glitches do not require operator presence in the field.

## Unified Logging for Remote Diagnostics

All platform components rely on **unified logging** through the standard Python `logging` module, configured to emit timestamps, severity levels, and module names in every record. This consistency allows operators to ship log files to remote diagnostics servers or parse them locally with standard Unix tools such as `grep` and `awk`. When errors propagate from [`radar_protocol.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/radar_protocol.py) through [`workers.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/workers.py) to [`dashboard.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/dashboard.py), the shared logger preserves a complete audit trail that makes post-mortem analysis straightforward.

## Summary

- **Worker-level counters:** `RadarDataWorker` and `GPSDataWorker` in [`workers.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/workers.py) maintain `self._error_count` and emit `errorOccurred` and `stats` signals for real-time visibility.
- **Centralized dashboard handling:** [`dashboard.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/dashboard.py) connects worker signals to `_on_worker_error`, providing UI feedback and triggering automatic recovery when thresholds are breached.
- **Hardware fault isolation:** [`radar_protocol.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/radar_protocol.py) wraps FTDI USB calls in `try/except` blocks and returns `None` on failure, preventing thread crashes.
- **Cross-layer resets:** The dashboard can reset the STM32 via [`hardware.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/hardware.py) and restart acquisition, enabling unattended recovery loops.
- **Unified Python logging:** Standard `logging` across all modules supports remote diagnostics and grep-friendly post-mortem analysis.

## Frequently Asked Questions

### How does the AERIS-10 radar detect errors in real time?

Each acquisition worker increments an internal `self._error_count` variable whenever it catches an exception and immediately emits an `errorOccurred(str)` Qt signal. The dashboard receives this signal within milliseconds, updating the UI and writing a timestamped log entry so operators see faults as they happen.

### What happens when the FTDI USB communication fails?

The low-level driver in [`radar_protocol.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/radar_protocol.py) traps the exception, logs a message such as `FT2232H read error`, and returns `None` instead of raw bytes. The calling worker can then choose to retry the read, increment its error counter, or escalate to a full hardware reset based on the current error rate.

### Can the system recover automatically without human intervention?

Yes. The dashboard monitors the `stats` signal for error count thresholds. If a worker exceeds the configured limit—typically ten consecutive errors—the dashboard stops acquisition, resets the STM32 microcontroller through [`hardware.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/hardware.py), and restarts the pipeline, allowing the radar to resume unattended field operation.

### Which files implement the health monitoring and recovery stack?

The four core files are [`9_Firmware/9_3_GUI/v7/workers.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/9_Firmware/9_3_GUI/v7/workers.py) for worker threads and error counting, [`9_Firmware/9_3_GUI/v7/dashboard.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/9_Firmware/9_3_GUI/v7/dashboard.py) for signal handling and UI alerts, [`9_Firmware/9_3_GUI/v7/radar_protocol.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/9_Firmware/9_3_GUI/v7/radar_protocol.py) for FTDI USB fault isolation, and [`9_Firmware/9_3_GUI/v7/hardware.py`](https://github.com/NawfalMotii79/PLFM_RADAR/blob/main/9_Firmware/9_3_GUI/v7/hardware.py) for STM32 reset and re-initialization APIs.