System Health Monitoring and Error Recovery in Embedded Radar Systems: Inside the AERIS-10 Open Source Radar

The AERIS-10 embedded radar platform implements a three-layer health monitoring and error recovery architecture across its FPGA, STM32 microcontroller, and Python GUI, using Qt signals for real-time fault propagation and automatic hardware resets to maintain unattended field operation.

System health monitoring and error recovery in embedded radar systems demand tight coordination between FPGA fabric, STM32 microcontroller firmware, and Python GUI dashboards. The PLFM_RADAR repository by NawfalMotii79 demonstrates how the AERIS-10 platform achieves this reliability by combining Python-based worker threads, low-level FTDI protocol wrappers, and cross-layer reset logic. The PyQt application entry point in 9_Firmware/9_3_GUI/v7/GUI_V7_PyQt.py wires the workers to the dashboard at startup, creating a unified pipeline that detects, logs, and recovers from faults without manual intervention.

System Health Monitoring and Error Recovery Architecture

Worker-Level Error Counters and Qt Signals

In 9_Firmware/9_3_GUI/v7/workers.py, each data-acquisition worker such as RadarDataWorker and GPSDataWorker tracks internal failures through error counting via self._error_count. Whenever an exception is caught—commonly ValueError or struct.error—the worker increments the counter, emits an errorOccurred(str) Qt signal, and writes a timestamped entry via logger.error.

This dual reporting mechanism ensures that both the UI and persistent logs receive the fault data simultaneously. The stats signal broadcasts a dictionary containing "errors" and "frames" keys every second, giving the dashboard a real-time health feed.

class RadarDataWorker(QThread):
    errorOccurred = pyqtSignal(str)      # <-- emitted on error

    stats = pyqtSignal(dict)             # <-- periodic status updates

    def run(self):
        while not self.isInterruptionRequested():
            try:
                data = self._read_radar_frame()
                self._process(data)
            except (ValueError, struct.error) as e:
                self._error_count += 1
                self.errorOccurred.emit(str(e))
                logger.error(f"RadarDataWorker error: {e}")

            # Emit health stats every second

            self.stats.emit({"errors": self._error_count,
                            "frames": self._frames_processed})

Centralized Dashboard Handling

The main dashboard defined in 9_Firmware/9_3_GUI/v7/dashboard.py connects each worker’s errorOccurred signal to a centralized slot named _on_worker_error. This slot provides centralized error handling that updates the status bar, writes to the unified log, and can trigger immediate recovery actions such as restarting the worker.

class RadarDashboard(QWidget):
    def __init__(self):
        ...
        self._radar_worker.errorOccurred.connect(self._on_worker_error)

    def _on_worker_error(self, msg: str):
        logger.error(f"Worker error: {msg}")
        # Immediate UI feedback

        self.statusBar().showMessage(f"Error: {msg}", 5000)
        # Automatic recovery if error count grows

        if self._radar_worker._error_count > 10:
            self._restart_acquisition()

Hardware-Level Fault Isolation in Embedded Radar Drivers

Graceful Degradation for FTDI USB Communication

Low-level I/O in 9_Firmware/9_3_GUI/v7/radar_protocol.py implements hardware-level fault isolation by wrapping every FT2232H USB transaction in a try/except block. Rather than crashing the acquisition thread, the driver catches the exception, logs an informative message such as log.error("FT2232H open failed: …"), and returns None to the caller, enabling graceful degradation. This design lets higher-level code decide between an immediate retry, a pipeline abort, or a hardware reset.

def read_bytes(self, length: int) -> Optional[bytes]:
    try:
        raw = self._ftdi.readbytes(length)
        return bytes(raw)
    except Exception as e:
        log.error(f"FT2232H read error: {e}")
        return None      # Caller can decide to retry or abort

Automatic Recovery Loops and Cross-Layer Resets

Threshold-Based Hardware Restart Logic

Beyond immediate slot reactions, the GUI’s main loop leverages automatic recovery loops that periodically poll the stats signal emitted by each worker. When the reported error count exceeds a configurable threshold, the dashboard executes a more extensive cross-layer reset: it stops the current radar acquisition, resets the STM32 microcontroller via 9_Firmware/9_3_GUI/v7/hardware.py, and restarts the acquisition pipeline.

This strategy bridges the Python GUI and the embedded STM32, ensuring that transient communication glitches do not require operator presence in the field.

Unified Logging for Remote Diagnostics

All platform components rely on unified logging through the standard Python logging module, configured to emit timestamps, severity levels, and module names in every record. This consistency allows operators to ship log files to remote diagnostics servers or parse them locally with standard Unix tools such as grep and awk. When errors propagate from radar_protocol.py through workers.py to dashboard.py, the shared logger preserves a complete audit trail that makes post-mortem analysis straightforward.

Summary

  • Worker-level counters: RadarDataWorker and GPSDataWorker in workers.py maintain self._error_count and emit errorOccurred and stats signals for real-time visibility.
  • Centralized dashboard handling: dashboard.py connects worker signals to _on_worker_error, providing UI feedback and triggering automatic recovery when thresholds are breached.
  • Hardware fault isolation: radar_protocol.py wraps FTDI USB calls in try/except blocks and returns None on failure, preventing thread crashes.
  • Cross-layer resets: The dashboard can reset the STM32 via hardware.py and restart acquisition, enabling unattended recovery loops.
  • Unified Python logging: Standard logging across all modules supports remote diagnostics and grep-friendly post-mortem analysis.

Frequently Asked Questions

How does the AERIS-10 radar detect errors in real time?

Each acquisition worker increments an internal self._error_count variable whenever it catches an exception and immediately emits an errorOccurred(str) Qt signal. The dashboard receives this signal within milliseconds, updating the UI and writing a timestamped log entry so operators see faults as they happen.

What happens when the FTDI USB communication fails?

The low-level driver in radar_protocol.py traps the exception, logs a message such as FT2232H read error, and returns None instead of raw bytes. The calling worker can then choose to retry the read, increment its error counter, or escalate to a full hardware reset based on the current error rate.

Can the system recover automatically without human intervention?

Yes. The dashboard monitors the stats signal for error count thresholds. If a worker exceeds the configured limit—typically ten consecutive errors—the dashboard stops acquisition, resets the STM32 microcontroller through hardware.py, and restarts the pipeline, allowing the radar to resume unattended field operation.

Which files implement the health monitoring and recovery stack?

The four core files are 9_Firmware/9_3_GUI/v7/workers.py for worker threads and error counting, 9_Firmware/9_3_GUI/v7/dashboard.py for signal handling and UI alerts, 9_Firmware/9_3_GUI/v7/radar_protocol.py for FTDI USB fault isolation, and 9_Firmware/9_3_GUI/v7/hardware.py for STM32 reset and re-initialization APIs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →