System Health Monitoring and Error Recovery in Embedded Radar Systems: Inside the AERIS-10 Open Source Radar
The AERIS-10 embedded radar platform implements a three-layer health monitoring and error recovery architecture across its FPGA, STM32 microcontroller, and Python GUI, using Qt signals for real-time fault propagation and automatic hardware resets to maintain unattended field operation.
System health monitoring and error recovery in embedded radar systems demand tight coordination between FPGA fabric, STM32 microcontroller firmware, and Python GUI dashboards. The PLFM_RADAR repository by NawfalMotii79 demonstrates how the AERIS-10 platform achieves this reliability by combining Python-based worker threads, low-level FTDI protocol wrappers, and cross-layer reset logic. The PyQt application entry point in 9_Firmware/9_3_GUI/v7/GUI_V7_PyQt.py wires the workers to the dashboard at startup, creating a unified pipeline that detects, logs, and recovers from faults without manual intervention.
System Health Monitoring and Error Recovery Architecture
Worker-Level Error Counters and Qt Signals
In 9_Firmware/9_3_GUI/v7/workers.py, each data-acquisition worker such as RadarDataWorker and GPSDataWorker tracks internal failures through error counting via self._error_count. Whenever an exception is caught—commonly ValueError or struct.error—the worker increments the counter, emits an errorOccurred(str) Qt signal, and writes a timestamped entry via logger.error.
This dual reporting mechanism ensures that both the UI and persistent logs receive the fault data simultaneously. The stats signal broadcasts a dictionary containing "errors" and "frames" keys every second, giving the dashboard a real-time health feed.
class RadarDataWorker(QThread):
errorOccurred = pyqtSignal(str) # <-- emitted on error
stats = pyqtSignal(dict) # <-- periodic status updates
def run(self):
while not self.isInterruptionRequested():
try:
data = self._read_radar_frame()
self._process(data)
except (ValueError, struct.error) as e:
self._error_count += 1
self.errorOccurred.emit(str(e))
logger.error(f"RadarDataWorker error: {e}")
# Emit health stats every second
self.stats.emit({"errors": self._error_count,
"frames": self._frames_processed})
Centralized Dashboard Handling
The main dashboard defined in 9_Firmware/9_3_GUI/v7/dashboard.py connects each worker’s errorOccurred signal to a centralized slot named _on_worker_error. This slot provides centralized error handling that updates the status bar, writes to the unified log, and can trigger immediate recovery actions such as restarting the worker.
class RadarDashboard(QWidget):
def __init__(self):
...
self._radar_worker.errorOccurred.connect(self._on_worker_error)
def _on_worker_error(self, msg: str):
logger.error(f"Worker error: {msg}")
# Immediate UI feedback
self.statusBar().showMessage(f"Error: {msg}", 5000)
# Automatic recovery if error count grows
if self._radar_worker._error_count > 10:
self._restart_acquisition()
Hardware-Level Fault Isolation in Embedded Radar Drivers
Graceful Degradation for FTDI USB Communication
Low-level I/O in 9_Firmware/9_3_GUI/v7/radar_protocol.py implements hardware-level fault isolation by wrapping every FT2232H USB transaction in a try/except block. Rather than crashing the acquisition thread, the driver catches the exception, logs an informative message such as log.error("FT2232H open failed: …"), and returns None to the caller, enabling graceful degradation. This design lets higher-level code decide between an immediate retry, a pipeline abort, or a hardware reset.
def read_bytes(self, length: int) -> Optional[bytes]:
try:
raw = self._ftdi.readbytes(length)
return bytes(raw)
except Exception as e:
log.error(f"FT2232H read error: {e}")
return None # Caller can decide to retry or abort
Automatic Recovery Loops and Cross-Layer Resets
Threshold-Based Hardware Restart Logic
Beyond immediate slot reactions, the GUI’s main loop leverages automatic recovery loops that periodically poll the stats signal emitted by each worker. When the reported error count exceeds a configurable threshold, the dashboard executes a more extensive cross-layer reset: it stops the current radar acquisition, resets the STM32 microcontroller via 9_Firmware/9_3_GUI/v7/hardware.py, and restarts the acquisition pipeline.
This strategy bridges the Python GUI and the embedded STM32, ensuring that transient communication glitches do not require operator presence in the field.
Unified Logging for Remote Diagnostics
All platform components rely on unified logging through the standard Python logging module, configured to emit timestamps, severity levels, and module names in every record. This consistency allows operators to ship log files to remote diagnostics servers or parse them locally with standard Unix tools such as grep and awk. When errors propagate from radar_protocol.py through workers.py to dashboard.py, the shared logger preserves a complete audit trail that makes post-mortem analysis straightforward.
Summary
- Worker-level counters:
RadarDataWorkerandGPSDataWorkerinworkers.pymaintainself._error_countand emiterrorOccurredandstatssignals for real-time visibility. - Centralized dashboard handling:
dashboard.pyconnects worker signals to_on_worker_error, providing UI feedback and triggering automatic recovery when thresholds are breached. - Hardware fault isolation:
radar_protocol.pywraps FTDI USB calls intry/exceptblocks and returnsNoneon failure, preventing thread crashes. - Cross-layer resets: The dashboard can reset the STM32 via
hardware.pyand restart acquisition, enabling unattended recovery loops. - Unified Python logging: Standard
loggingacross all modules supports remote diagnostics and grep-friendly post-mortem analysis.
Frequently Asked Questions
How does the AERIS-10 radar detect errors in real time?
Each acquisition worker increments an internal self._error_count variable whenever it catches an exception and immediately emits an errorOccurred(str) Qt signal. The dashboard receives this signal within milliseconds, updating the UI and writing a timestamped log entry so operators see faults as they happen.
What happens when the FTDI USB communication fails?
The low-level driver in radar_protocol.py traps the exception, logs a message such as FT2232H read error, and returns None instead of raw bytes. The calling worker can then choose to retry the read, increment its error counter, or escalate to a full hardware reset based on the current error rate.
Can the system recover automatically without human intervention?
Yes. The dashboard monitors the stats signal for error count thresholds. If a worker exceeds the configured limit—typically ten consecutive errors—the dashboard stops acquisition, resets the STM32 microcontroller through hardware.py, and restarts the pipeline, allowing the radar to resume unattended field operation.
Which files implement the health monitoring and recovery stack?
The four core files are 9_Firmware/9_3_GUI/v7/workers.py for worker threads and error counting, 9_Firmware/9_3_GUI/v7/dashboard.py for signal handling and UI alerts, 9_Firmware/9_3_GUI/v7/radar_protocol.py for FTDI USB fault isolation, and 9_Firmware/9_3_GUI/v7/hardware.py for STM32 reset and re-initialization APIs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →