FPGA Software Replica for Offline Radar Data Processing and Replay: A Complete Guide

The AERIS-10 radar platform provides a bit-accurate Python replica (SoftwareFPGA) that mirrors the complete FPGA signal processing chain, enabling offline radar data replay and algorithm development without hardware.

This guide explores the SoftwareFPGA module from the open-source AERIS-10 phased-array radar project. You'll learn how to use this FPGA software replica for offline radar data processing, from loading raw I/Q samples through generating detection-ready RadarFrame objects identical to hardware output.

What Is the AERIS-10 Radar Platform?

AERIS-10 is a complete 10.5 GHz phased-array radar system combining physical hardware with a comprehensive software stack. The repository at NawfalMotii79/PLFM_RADAR contains everything from PCB schematics to FPGA firmware—and critically, a software replica of the FPGA processing chain that runs entirely in Python.

The platform supports two antenna configurations:

  • Nexus: 8 × 16 patch array, approximately 3 km range
  • Extended: 32 × 16 slotted waveguide array, approximately 20 km range

The hardware FPGA (Xilinx XC7A50T) executes a fixed pipeline: range FFT, range-bin decimation, MTI cancellation, Doppler FFT, DC notch filtering, and CFAR detection. The SoftwareFPGA class in 9_Firmware/9_3_GUI/v7/software_fpga.py replicates this pipeline exactly, bit-for-bit.

Hardware Signal Processing Pipeline

Understanding the hardware chain clarifies what the software replica implements. The FPGA processes each chirp through these deterministic stages:

Stage Function Hardware Location
ADC Capture 16-bit I/Q sampling Main board ADC
Range FFT 1024-point FFT per chirp XC7A50T FPGA fabric
Range Decimation 1024 → 64 bins FPGA
MTI Canceller Static clutter suppression FPGA
Doppler FFT 16-point FFT across chirps FPGA
DC Notch Zero DC bin removal FPGA
Detection Threshold or CFAR FPGA
Frame Assembly Pack into RadarFrame FPGA → USB

The SoftwareFPGA class mirrors this exact sequence using reference implementations from golden_reference.py.

The SoftwareFPGA API: Bit-Accurate Offline Processing

The SoftwareFPGA class provides a register-mirror architecture that matches the FPGA's control interface. This design ensures configuration parameters (CFAR enable, guard cells, training cells) produce identical results in software and hardware.

Core API Structure

from software_fpga import SoftwareFPGA, quantize_raw_iq

# Initialize the replica with default register values

fpga = SoftwareFPGA()

# Configure detection parameters

fpga.set_cfar_enable(True)      # Enable CFAR vs. simple threshold

fpga.set_cfar_guard(2)           # Guard cells around CUT

fpga.set_cfar_train(8)           # Training cells for noise estimation

# Process quantized I/Q data

frame = fpga.process_chirps(iq_i, iq_q, frame_number=0, timestamp=0.0)

Register Mirror Implementation

The constructor populates default register values matching the RTL reset state:


# From software_fpga.py - initialization mimics FPGA reset

self._registers = {
    'cfar_enable': 0,
    'cfar_guard_cells': 2,
    'cfar_training_cells': 8,
    'cfar_mode': 0,  # CA (Cell Averaging)

    'mti_enable': 1,
    'dc_notch_enable': 1,
    # ... additional register mirrors

}

This register-mirror pattern enables offline validation of FPGA configurations before deployment.

Processing Stages in Detail

Each SoftwareFPGA processing stage calls a reference implementation from golden_reference.py. These functions are the golden standard against which the VHDL/Verilog implementation is verified.

1. Range FFT (run_range_fft)


# Called internally by process_chirps

from golden_reference import run_range_fft

# 1024-point complex FFT with FPGA-compatible twiddle factors

range_fft_result = run_range_fft(quantized_samples, twiddle_file='tw_range_1024.txt')

The function loads twiddle-factor files when available to ensure numerical matching with the hardware FFT implementation.

2. Range-Bin Decimation (run_range_bin_decimator)

Reduces 1024 range bins to 64 through configurable decimation filtering, matching the FPGA's resource-optimized filter structure.

3. MTI Canceller (run_mti_canceller)

from golden_reference import run_mti_canceller

# Two-pulse canceller: y[n] = x[n] - x[n-1]

mti_output = run_mti_canceller(decimated_data)

Removes stationary clutter by subtracting consecutive chirps. Critical for detecting moving targets against static backgrounds.

4. Doppler FFT (run_doppler_fft)

16-point FFT across the chirp dimension (slow-time) to resolve velocity. The small FFT size reflects the 16-chirp coherent processing interval.

5. DC Notch (run_dc_notch)

Zeros the zero-Doppler bin to suppress leakage from imperfect MTI cancellation.

6. Detection

Two modes available:

  • Simple threshold (run_detection): Fixed magnitude threshold
  • CFAR (run_cfar_ca): Cell-Averaging CFAR with configurable guard and training cells

CFAR mode selection uses the _CFAR_MODE_MAP dictionary supporting CA, GO (Greatest Of), and SO (Smallest Of) variants.

RadarFrame Data Structure

The process_chirps method returns a RadarFrame dataclass defined in radar_protocol.py:

from dataclasses import dataclass, field
import numpy as np

NUM_RANGE_BINS = 64
NUM_DOPPLER_BINS = 32

@dataclass
class RadarFrame:
    """One complete radar frame (64 range × 32 Doppler)."""
    timestamp: float = 0.0
    range_doppler_i: np.ndarray = field(
        default_factory=lambda: np.zeros((NUM_RANGE_BINS, NUM_DOPPLER_BINS), dtype=np.int16)
    )
    range_doppler_q: np.ndarray = field(
        default_factory=lambda: np.zeros((NUM_RANGE_BINS, NUM_DOPPLER_BINS), dtype=np.int16)
    )
    magnitude: np.ndarray = field(
        default_factory=lambda: np.zeros((NUM_RANGE_BINS, NUM_DOPPLER_BINS), dtype=np.float64)
    )
    detections: np.ndarray = field(
        default_factory=lambda: np.zeros((NUM_RANGE_BINS, NUM_DOPPLER_BINS), dtype=np.uint8)
    )
    range_profile: np.ndarray = field(
        default_factory=lambda: np.zeros(NUM_RANGE_BINS, dtype=np.float64)
    )
    detection_count: int = 0
    frame_number: int = 0

This structure matches exactly what the FPGA transmits over USB, enabling seamless switching between hardware and software processing.

Practical Code Examples

Loading and Processing SDR-Captured Data

import numpy as np
from software_fpga import SoftwareFPGA, quantize_raw_iq

# Load complex baseband from SDR or previous capture

# Shape: (num_chirps, samples_per_chirp)

raw_iq = np.fromfile('sdr_capture.bin', dtype=np.complex64).reshape(-1, 1024)

# Convert to FPGA 16-bit format (saturation arithmetic)

iq_i, iq_q = quantize_raw_iq(raw_iq)

# Initialize and configure software FPGA

fpga = SoftwareFPGA()
fpga.set_cfar_enable(True)
fpga.set_cfar_guard(2)
fpga.set_cfar_train(8)

# Process entire capture

frame = fpga.process_chirps(
    iq_i, iq_q,
    frame_number=1,
    timestamp=123.456
)

print(f"Detections: {frame.detection_count}")
print(f"Peak range bin: {frame.magnitude.argmax() // 32}")
print(f"Peak Doppler bin: {frame.magnitude.argmax() % 32}")

Batch Processing Multiple Captures

from pathlib import Path

fpga = SoftwareFPGA()
fpga.set_cfar_enable(True)

results = []
for capture_path in Path('captures/').glob('*.bin'):
    raw_iq = np.fromfile(capture_path, dtype=np.complex64).reshape(-1, 1024)
    iq_i, iq_q = quantize_raw_iq(raw_iq)
    
    frame = fpga.process_chirps(iq_i, iq_q)
    results.append({
        'file': capture_path.name,
        'detections': frame.detection_count,
        'max_magnitude': frame.magnitude.max()
    })

Visualizing Range-Doppler Maps

import matplotlib.pyplot as plt

def plot_range_doppler(frame: RadarFrame, title: str = "Range-Doppler Map"):
    """Visualize processed radar frame with proper axes."""
    fig, axes = plt.subplots(1, 2, figsize=(12, 5))
    
    # Magnitude in dB

    mag_db = 20 * np.log10(frame.magnitude + 1e-6)
    im1 = axes[0].imshow(
        mag_db,
        aspect='auto',
        extent=[-16, 15, 0, 64],
        origin='lower',
        cmap='viridis'
    )
    axes[0].set_title('Magnitude (dB)')
    axes[0].set_xlabel('Doppler Bin')
    axes[0].set_ylabel('Range Bin')
    plt.colorbar(im1, ax=axes[0])
    
    # Detection mask

    im2 = axes[1].imshow(
        frame.detections,
        aspect='auto',
        extent=[-16, 15, 0, 64],
        origin='lower',
        cmap='Reds'
    )
    axes[1].set_title(f'Detections (count: {frame.detection_count})')
    axes[1].set_xlabel('Doppler Bin')
    axes[1].set_ylabel('Range Bin')
    
    plt.suptitle(title)
    plt.tight_layout()
    return fig

# Usage

fig = plot_range_doppler(frame, "AERIS-10 Offline Processing")
plt.show()

Integration with GUI Replay Mode

The PyQt-based GUI (GUI_V7_PyQt.py) uses SoftwareFPGA when hardware is unavailable. The worker thread (workers.py) instantiates the replica transparently:


# Simplified from workers.py

class ProcessingWorker(QObject):
    update_plot = pyqtSignal(RadarFrame)
    
    def __init__(self, use_hardware: bool = False):
        super().__init__()
        self.use_hardware = use_hardware
        self.fpga = None
    
    def initialize(self):
        if self.use_hardware:
            # Initialize USB connection to FPGA

            pass
        else:
            # Use software replica for offline replay

            self.fpga = SoftwareFPGA()
            self.fpga.set_cfar_enable(self.config.get('cfar_enable', True))
    
    def process_file(self, filepath: str):
        raw_iq = np.fromfile(filepath, dtype=np.complex64).reshape(-1, 1024)
        iq_i, iq_q = quantize_raw_iq(raw_iq)
        
        frame = self.fpga.process_chirps(iq_i, iq_q)
        self.update_plot.emit(frame)

This architecture ensures identical processing whether connected to hardware or replaying saved data.

Key Source Files and Locations

File Path Purpose
software_fpga.py 9_Firmware/9_3_GUI/v7/software_fpga.py Main SoftwareFPGA class implementation
golden_reference.py 9_Firmware/9_2_FPGA/tb/cosim/real_data/golden_reference.py Reference signal processing functions
radar_protocol.py 9_Firmware/9_3_GUI/radar_protocol.py RadarFrame and StatusResponse dataclasses
GUI_V7_PyQt.py 9_Firmware/9_3_GUI/v7/GUI_V7_PyQt.py PyQt GUI with replay mode
workers.py 9_Firmware/9_3_GUI/v7/workers.py Background processing threads

Performance Considerations

The SoftwareFPGA replica prioritizes bit-accuracy over speed. For large datasets:

  • NumPy operations provide reasonable performance for development
  • Batch processing enables overnight analysis of capture archives
  • FFTW twiddle factors (when available) accelerate FFT computation

For production deployments, the hardware FPGA implementation maintains real-time performance at the cost of flexibility.

Summary

  • SoftwareFPGA provides a bit-accurate Python replica of the AERIS-10 FPGA processing chain
  • Register-mirror architecture ensures configuration compatibility with hardware
  • Reference implementations in golden_reference.py validate the VHDL/Verilog design
  • RadarFrame dataclass enables seamless data exchange between software and hardware paths
  • GUI integration supports transparent switching between live hardware and offline replay

Frequently Asked Questions

How does SoftwareFPGA guarantee bit-accurate results?

SoftwareFPGA uses fixed-point arithmetic, saturation quantization, and identical twiddle factors to the hardware FFT implementation. Each processing stage calls a reference function from golden_reference.py that is co-simulated against the RTL during FPGA verification. The constructor initializes register values matching the FPGA's reset state.

Can I use SoftwareFPGA with my own radar hardware?

Yes, provided you adapt the input quantization (quantize_raw_iq) to match your ADC bit-width and sample format. The processing chain parameters (FFT sizes, decimation ratios, detection thresholds) are configurable through the register interface. You may need to modify NUM_RANGE_BINS and NUM_DOPPLER_BINS in radar_protocol.py for different geometries.

What is the performance difference between SoftwareFPGA and hardware?

The hardware XC7A50T FPGA processes chirps in real-time at the radar PRF. The Python replica processes data post-capture, typically at 10-100× slower than real-time depending on CPU. For algorithm development, the software replica enables rapid iteration; for deployment, the bitstream synthesis ensures identical behavior.

How do I validate that my software processing matches hardware output?

Capture raw I/Q data from the hardware's debug interface, process through SoftwareFPGA, and compare against the FPGA's USB output. The repository includes co-simulation infrastructure in 9_Firmware/9_2_FPGA/tb/cosim/ for automated verification. Bit-exact matching confirms algorithm correctness before RTL modification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →