# Range FFT and Doppler FFT Implementation in FPGA Signal Processing: A Complete Technical Guide

> Learn how to implement Range FFT and Doppler FFT in FPGA signal processing using a radix-2 DIT engine. Explore a complete technical guide for PLFM RADAR firmware.

- Repository: [NawfalMotii79/PLFM_RADAR](https://github.com/NawfalMotii79/PLFM_RADAR)
- Tags: deep-dive
- Published: 2026-08-20

---

**The PLFM RADAR FPGA firmware implements Range FFT and Doppler FFT using a fully-parameterized radix-2 DIT engine with 1024-point transforms for range processing and dual 16-point staggered-PRF transforms for Doppler velocity extraction.**

This article examines the complete Verilog implementation of **Range FFT and Doppler FFT in FPGA signal processing** as found in the open-source **PLFM RADAR** repository by NawfalMotii79. Both modules are synthesizable, vendor-IP-free designs optimized for Xilinx DSP48E1 blocks and block RAMs, providing a production-ready reference for radar firmware developers.

## Range FFT Engine Architecture

The **Range FFT** module (`fft_engine.v`) converts raw ADC samples into range bins through a 1024-point radix-2 Decimation-In-Time (DIT) transform. The design prioritizes resource efficiency and deterministic throughput over latency, making it ideal for real-time radar pipelines.

### Core Parameters and Pipeline Structure

The engine accepts compile-time parameters that govern every aspect of the transform:

| Parameter | Default | Function |
|-----------|---------|----------|
| `N` | 1024 | FFT size (power of two) |
| `LOG2N` | 10 | log₂(N), drives counter widths |
| `DATA_W` | 16 | Input/output data width |
| `TWIDDLE_W` | 16 | Twiddle factor precision |
| `TWIDDLE_FILE` | — | Init file for cosine ROM |

The **four-stage butterfly pipeline** (READ → TW → MULT2 → WRITE) processes one butterfly per clock cycle, executing `N/2` butterflies across `LOG2N` stages. This yields a fixed throughput of `N × LOG2N` cycles per FFT, ignoring the initial load phase.

The data path uses dedicated pipeline registers defined in `fft_engine.v` lines 21-33:

```verilog
// Butterfly pipeline registers
reg signed [INTERNAL_W-1:0] rd_a_re, rd_a_im;   // BRAM port A data
reg signed [INTERNAL_W-1:0] rd_b_re, rd_b_im;   // BRAM port B data
reg signed [PROD_W:0] bf_prod_re, bf_prod_im;   // DSP48E1 multiply result

```

### Quarter-Wave Twiddle ROM Optimization

The **twiddle factor storage** exploits cosine symmetry to reduce ROM usage by 75%. Only the first quadrant of cosine values are stored; sine values are generated through index manipulation:

```verilog
(* rom_style = "block" *) reg signed [TWIDDLE_W-1:0] cos_rom [0:TW_QUARTER-1];
initial begin
    $readmemh(TWIDDLE_FILE, cos_rom);
end

always @(*) begin : tw_lookup
    if (k == 0)          tw_cos_lookup = cos_rom[0];
    else if (k == FFT_N_QTR) tw_cos_lookup = {TWIDDLE_W{1'b0}};
    else if (k < FFT_N_QTR) begin
        tw_cos_lookup = cos_rom[k[TW_ADDR_W-1:0]];
        tw_sin_lookup = cos_rom[FFT_N_QTR[LOG2N-1:0] - k];
    end else begin
        tw_cos_lookup = -cos_rom[FFT_N_HALF[LOG2N-1:0] - k];
        tw_sin_lookup =  cos_rom[k - FFT_N_QTR[LOG2N-1:0]];
    end
end

```

This technique, documented in lines 101-124 of `fft_engine.v`, eliminates the need for separate sine storage while maintaining full spectral accuracy.

### Bit-Reversed Addressing and FSM Control

Input data are written in **bit-reversed order** during the load phase, eliminating the traditional post-FFT reordering pass. The address generator uses a simple bit-reversal network on the load counter.

The state machine (lines 338-380) provides clean external handshaking:

```verilog
always @(posedge clk or negedge reset_n) begin
    if (!reset_n) begin
        state <= ST_IDLE;
    end else begin
        case (state)
            ST_IDLE:   if (start) state <= ST_LOAD;
            ST_LOAD:   if (load_count == FFT_N_M1) state <= ST_BF_READ;
            ST_BF_READ:    state <= ST_BF_TW;
            ST_BF_TW:      state <= ST_BF_MULT2;
            ST_BF_MULT2:   state <= ST_BF_WRITE;
            ST_BF_WRITE:   if (bfly_count == FFT_N_HALF_M1) … // stage complete
            ST_OUTPUT:   … // streaming results
            ST_DONE:     done <= 1'b1; state <= ST_IDLE;
        endcase
    end
    end
end

```

The **`inverse`** input selects forward or inverse operation. For IFFT mode, the output scaling by `1/N` is implemented through a right-shift of `LOG2N` bits on the final result.

## Doppler FFT Processor with Staggered-PRF Support

The **Doppler processor** (`doppler_processor.v`) handles velocity estimation through a dual-sub-frame architecture that resolves the range-Doppler ambiguity inherent in pulsed radar systems.

### Staggered-PRF Architecture Rationale

A single 32-point FFT on a staggered-PRF frame would produce **spectral contamination** because the non-uniform sampling violates the Nyquist criterion. Instead, the PLFM RADAR implementation splits each 32-chirp frame into two uniform 16-chirp sub-frames:

- **Sub-frame 0**: 16 chirps at **long PRI**
- **Sub-frame 1**: 16 chirps at **short PRI**

Each sub-frame undergoes independent 16-point FFT processing. The differing PRI lengths cause identical target velocities to map to **different Doppler bin indices** between sub-frames, enabling post-processing ambiguity resolution.

### Memory Organization and Windowing

The processor maintains two block RAMs (`doppler_i_mem`, `doppler_q_mem`) sized for all range bins:

```verilog
// Memory address computation (lines 151-156)
assign mem_write_addr = (write_chirp_index * RANGE_BINS) + write_range_bin;
assign mem_read_addr  = (read_doppler_index * RANGE_BINS) + read_range_bin;

```

With **`RANGE_BINS = 64`**, each sub-frame stores 16 × 64 = 1024 complex samples.

A **Hamming window** (Q15 format) is applied to every sample before FFT loading to suppress spectral leakage:

```verilog
// Window coefficients (lines 73-101)
window_coeff[0]  = 16'h0A3D;  // 0.0800 * 32767 = 2621
window_coeff[1]  = 16'h0E5C;  // 0.1116 * 32767 = 3676
...
window_coeff[15] = 16'h0A3D;

```

### Sub-Frame FSM and FFT Interface

The main state machine tracks which sub-frame is active and sequences the FFT engine:

```verilog
localparam S_IDLE = 3'b000, S_ACCUMULATE = 3'b001, S_LOAD_FFT = 3'b010,
           S_FFT_WAIT = 3'b011, S_OUTPUT = 3'b100;

always @(posedge clk or negedge reset_n) begin
    if (!reset_n) state <= S_IDLE;
    else case (state)
        S_IDLE:       if (frame_start_pulse) state <= S_ACCUMULATE;
        S_ACCUMULATE: … // window, write to RAM
        S_LOAD_FFT:   … // feed fft_engine
        S_FFT_WAIT:   if (fft_ready) state <= S_OUTPUT;
        S_OUTPUT:     … // stream with doppler_bin tagging
    endcase
end

```

The **output tagging system** encodes both velocity bin and PRI ambiguity information in a 5-bit field:

```verilog
assign doppler_bin = {current_sub_frame, fft_bin[3:0]};

```

This allows downstream logic to correlate detections across sub-frames and compute unambiguous velocity through lookup tables or arithmetic comparison.

## Integration Code Examples

### Instantiating the Range FFT Engine

The following wrapper demonstrates proper instantiation for 1024-point range processing:

```verilog
module range_fft_wrapper (
    input  wire        clk,
    input  wire        reset_n,
    input  wire        start,
    input  wire signed [15:0] din_re,
    input  wire signed [15:0] din_im,
    input  wire        din_valid,
    output wire signed [15:0] dout_re,
    output wire signed [15:0] dout_im,
    output wire        dout_valid,
    output wire        busy,
    output wire        done
);
    fft_engine #(
        .N            (1024),
        .LOG2N        (10),
        .DATA_W       (16),
        .TWIDDLE_W    (16),
        .TWIDDLE_FILE ("fft_twiddle_1024.mem")
    ) u_fft (
        .clk        (clk),
        .reset_n    (reset_n),
        .start      (start),
        .inverse    (1'b0),          // forward FFT for range processing
        .din_re     (din_re),
        .din_im     (din_im),
        .din_valid  (din_valid),
        .dout_re    (dout_re),
        .dout_im    (dout_im),
        .dout_valid (dout_valid),
        .busy       (busy),
        .done       (done)
    );
endmodule

```

For IFFT operation (matched filter synthesis), toggle `.inverse(1'b1)` and the engine automatically applies `1/N` scaling.

### Doppler Processor Integration

```verilog
module doppler_top (
    input  wire         clk,
    input  wire         reset_n,
    input  wire [31:0]  range_data,   // {I[31:16], Q[15:0]}
    input  wire         data_valid,
    input  wire         new_chirp_frame,
    output wire [31:0]  doppler_output,
    output wire         doppler_valid,
    output wire [4:0]   doppler_bin,
    output wire [5:0]   range_bin,
    output wire         sub_frame,
    output wire         processing_active,
    output wire         frame_complete
);
    doppler_processor_optimized #(
        .DOPPLER_FFT_SIZE    (16),
        .RANGE_BINS          (64),
        .CHIRPS_PER_FRAME    (32),
        .CHIRPS_PER_SUBFRAME (16),
        .WINDOW_TYPE         (0)   // Hamming
    ) u_doppler (
        .clk                (clk),
        .reset_n            (reset_n),
        .range_data         (range_data),
        .data_valid         (data_valid),
        .new_chirp_frame    (new_chirp_frame),
        .doppler_output     (doppler_output),
        .doppler_valid      (doppler_valid),
        .doppler_bin        (doppler_bin),
        .range_bin          (range_bin),
        .sub_frame          (sub_frame),
        .processing_active  (processing_active),
        .frame_complete     (frame_complete)
    );
endmodule

```

### Verified Testbench Pattern

The reference testbench in `tb_fft_engine.v` demonstrates stimulus generation and completion detection:

```verilog
module tb_fft_engine;
    reg clk = 0;
    always #5 clk = ~clk;  // 100 MHz

    reg reset_n = 0, start = 0;
    reg signed [15:0] din_re = 0, din_im = 0;
    reg din_valid = 0;

    fft_engine #(.N(1024), .LOG2N(10)) dut (…);

    initial begin
        #20 reset_n = 1;
        #10 start = 1;
        repeat (1024) begin
            @(posedge clk);
            din_valid = 1;
            din_re = $random;
            din_im = $random;
        end
        din_valid = 0; start = 0;
        wait (done);
        $display("FFT completed: %0d + j%0d", dout_re, dout_im);
    end
endmodule

```

## Complete File Reference

| File | Path | Purpose |
|------|------|---------|
| `fft_engine.v` | `/9_Firmware/9_2_FPGA/fft_engine.v` | Generic radix-2 DIT FFT/IFFT engine |
| `doppler_processor.v` | `/9_Firmware/9_2_FPGA/doppler_processor.v` | Staggered-PRF Doppler processor |
| `matched_filter_processing_chain.v` | `/9_Firmware/9_2_FPGA/matched_filter_processing_chain.v` | Range profile generation |
| `range_bin_decimator.v` | `/9_Firmware/9_2_FPGA/range_bin_decimator.v` | 1024→64 bin reduction |
| `tb_fft_engine.v` | `/9_Firmware/tests/cross_layer/tb_fft_engine.v` | FFT engine verification |
| `tb_doppler_cosim.v` | `/9_Firmware/9_2_FPGA/tb/tb_doppler_cosim.v` | Doppler cosimulation |

## Summary

- The **Range FFT engine** (`fft_engine.v`) provides a 1024-point radix-2 DIT implementation with quarter-wave twiddle ROM optimization, achieving 75% memory reduction through cosine symmetry exploitation.

- **Bit-reversed loading** eliminates output reordering, while the four-stage pipeline sustains one butterfly per clock with deterministic `N × LOG2N` cycle latency.

- The **Doppler processor** implements **staggered-PRF processing** through dual 16-point FFTs, applying Hamming windowing and sub-frame tagging (`{sub_frame, bin[3:0]}`) to resolve velocity ambiguities.

- Both modules expose **simple handshaking** (`start`/`busy`/`done`) and require **no vendor IP**, mapping cleanly to Xilinx BRAM and DSP48E1 resources.

## Frequently Asked Questions

### What FFT algorithm does the PLFM RADAR implementation use?

The implementation uses a **radix-2 Decimation-In-Time (DIT) Cooley-Tukey algorithm** with iterative butterfly processing. This choice balances hardware efficiency against algorithmic complexity, avoiding the routing congestion of higher-radix designs while maintaining sufficient throughput for real-time radar processing.

### How does the staggered-PRF Doppler processor resolve velocity ambiguity?

By splitting each 32-chirp frame into two 16-chirp sub-frames with different PRI values, the processor generates **two distinct Doppler mappings** for the same physical velocity. The 5-bit `doppler_bin` output tags each result with its sub-frame origin (`{sub_frame, bin[3:0]}`), allowing downstream logic to correlate detections and compute unambiguous velocity through table lookup or arithmetic comparison of the bin indices.

### Can the FFT engine be used for sizes other than 1024?

Yes. The `N` and `LOG2N` parameters enable any power-of-two transform size. The twiddle ROM must be regenerated for the target size using the quarter-wave symmetry (storing only N/4 values), and the `TWIDDLE_FILE` init path updated accordingly. The pipelined architecture scales linearly in BRAM and DSP48E1 usage with transform size.

### What is the purpose of the `range_bin_decimator.v` module?

This module reduces the 1024-point range profile output from the matched filter to **64 range bins** before Doppler processing. This decimation lowers the memory bandwidth and storage requirements for the Doppler stage by 16×, trading range resolution for system resource efficiency in accordance with the radar's operational requirements.