Range FFT and Doppler FFT Implementation in FPGA Signal Processing: A Complete Technical Guide
The PLFM RADAR FPGA firmware implements Range FFT and Doppler FFT using a fully-parameterized radix-2 DIT engine with 1024-point transforms for range processing and dual 16-point staggered-PRF transforms for Doppler velocity extraction.
This article examines the complete Verilog implementation of Range FFT and Doppler FFT in FPGA signal processing as found in the open-source PLFM RADAR repository by NawfalMotii79. Both modules are synthesizable, vendor-IP-free designs optimized for Xilinx DSP48E1 blocks and block RAMs, providing a production-ready reference for radar firmware developers.
Range FFT Engine Architecture
The Range FFT module (fft_engine.v) converts raw ADC samples into range bins through a 1024-point radix-2 Decimation-In-Time (DIT) transform. The design prioritizes resource efficiency and deterministic throughput over latency, making it ideal for real-time radar pipelines.
Core Parameters and Pipeline Structure
The engine accepts compile-time parameters that govern every aspect of the transform:
| Parameter | Default | Function |
|---|---|---|
N |
1024 | FFT size (power of two) |
LOG2N |
10 | log₂(N), drives counter widths |
DATA_W |
16 | Input/output data width |
TWIDDLE_W |
16 | Twiddle factor precision |
TWIDDLE_FILE |
— | Init file for cosine ROM |
The four-stage butterfly pipeline (READ → TW → MULT2 → WRITE) processes one butterfly per clock cycle, executing N/2 butterflies across LOG2N stages. This yields a fixed throughput of N × LOG2N cycles per FFT, ignoring the initial load phase.
The data path uses dedicated pipeline registers defined in fft_engine.v lines 21-33:
// Butterfly pipeline registers
reg signed [INTERNAL_W-1:0] rd_a_re, rd_a_im; // BRAM port A data
reg signed [INTERNAL_W-1:0] rd_b_re, rd_b_im; // BRAM port B data
reg signed [PROD_W:0] bf_prod_re, bf_prod_im; // DSP48E1 multiply result
Quarter-Wave Twiddle ROM Optimization
The twiddle factor storage exploits cosine symmetry to reduce ROM usage by 75%. Only the first quadrant of cosine values are stored; sine values are generated through index manipulation:
(* rom_style = "block" *) reg signed [TWIDDLE_W-1:0] cos_rom [0:TW_QUARTER-1];
initial begin
$readmemh(TWIDDLE_FILE, cos_rom);
end
always @(*) begin : tw_lookup
if (k == 0) tw_cos_lookup = cos_rom[0];
else if (k == FFT_N_QTR) tw_cos_lookup = {TWIDDLE_W{1'b0}};
else if (k < FFT_N_QTR) begin
tw_cos_lookup = cos_rom[k[TW_ADDR_W-1:0]];
tw_sin_lookup = cos_rom[FFT_N_QTR[LOG2N-1:0] - k];
end else begin
tw_cos_lookup = -cos_rom[FFT_N_HALF[LOG2N-1:0] - k];
tw_sin_lookup = cos_rom[k - FFT_N_QTR[LOG2N-1:0]];
end
end
This technique, documented in lines 101-124 of fft_engine.v, eliminates the need for separate sine storage while maintaining full spectral accuracy.
Bit-Reversed Addressing and FSM Control
Input data are written in bit-reversed order during the load phase, eliminating the traditional post-FFT reordering pass. The address generator uses a simple bit-reversal network on the load counter.
The state machine (lines 338-380) provides clean external handshaking:
always @(posedge clk or negedge reset_n) begin
if (!reset_n) begin
state <= ST_IDLE;
end else begin
case (state)
ST_IDLE: if (start) state <= ST_LOAD;
ST_LOAD: if (load_count == FFT_N_M1) state <= ST_BF_READ;
ST_BF_READ: state <= ST_BF_TW;
ST_BF_TW: state <= ST_BF_MULT2;
ST_BF_MULT2: state <= ST_BF_WRITE;
ST_BF_WRITE: if (bfly_count == FFT_N_HALF_M1) … // stage complete
ST_OUTPUT: … // streaming results
ST_DONE: done <= 1'b1; state <= ST_IDLE;
endcase
end
end
end
The inverse input selects forward or inverse operation. For IFFT mode, the output scaling by 1/N is implemented through a right-shift of LOG2N bits on the final result.
Doppler FFT Processor with Staggered-PRF Support
The Doppler processor (doppler_processor.v) handles velocity estimation through a dual-sub-frame architecture that resolves the range-Doppler ambiguity inherent in pulsed radar systems.
Staggered-PRF Architecture Rationale
A single 32-point FFT on a staggered-PRF frame would produce spectral contamination because the non-uniform sampling violates the Nyquist criterion. Instead, the PLFM RADAR implementation splits each 32-chirp frame into two uniform 16-chirp sub-frames:
- Sub-frame 0: 16 chirps at long PRI
- Sub-frame 1: 16 chirps at short PRI
Each sub-frame undergoes independent 16-point FFT processing. The differing PRI lengths cause identical target velocities to map to different Doppler bin indices between sub-frames, enabling post-processing ambiguity resolution.
Memory Organization and Windowing
The processor maintains two block RAMs (doppler_i_mem, doppler_q_mem) sized for all range bins:
// Memory address computation (lines 151-156)
assign mem_write_addr = (write_chirp_index * RANGE_BINS) + write_range_bin;
assign mem_read_addr = (read_doppler_index * RANGE_BINS) + read_range_bin;
With RANGE_BINS = 64, each sub-frame stores 16 × 64 = 1024 complex samples.
A Hamming window (Q15 format) is applied to every sample before FFT loading to suppress spectral leakage:
// Window coefficients (lines 73-101)
window_coeff[0] = 16'h0A3D; // 0.0800 * 32767 = 2621
window_coeff[1] = 16'h0E5C; // 0.1116 * 32767 = 3676
...
window_coeff[15] = 16'h0A3D;
Sub-Frame FSM and FFT Interface
The main state machine tracks which sub-frame is active and sequences the FFT engine:
localparam S_IDLE = 3'b000, S_ACCUMULATE = 3'b001, S_LOAD_FFT = 3'b010,
S_FFT_WAIT = 3'b011, S_OUTPUT = 3'b100;
always @(posedge clk or negedge reset_n) begin
if (!reset_n) state <= S_IDLE;
else case (state)
S_IDLE: if (frame_start_pulse) state <= S_ACCUMULATE;
S_ACCUMULATE: … // window, write to RAM
S_LOAD_FFT: … // feed fft_engine
S_FFT_WAIT: if (fft_ready) state <= S_OUTPUT;
S_OUTPUT: … // stream with doppler_bin tagging
endcase
end
The output tagging system encodes both velocity bin and PRI ambiguity information in a 5-bit field:
assign doppler_bin = {current_sub_frame, fft_bin[3:0]};
This allows downstream logic to correlate detections across sub-frames and compute unambiguous velocity through lookup tables or arithmetic comparison.
Integration Code Examples
Instantiating the Range FFT Engine
The following wrapper demonstrates proper instantiation for 1024-point range processing:
module range_fft_wrapper (
input wire clk,
input wire reset_n,
input wire start,
input wire signed [15:0] din_re,
input wire signed [15:0] din_im,
input wire din_valid,
output wire signed [15:0] dout_re,
output wire signed [15:0] dout_im,
output wire dout_valid,
output wire busy,
output wire done
);
fft_engine #(
.N (1024),
.LOG2N (10),
.DATA_W (16),
.TWIDDLE_W (16),
.TWIDDLE_FILE ("fft_twiddle_1024.mem")
) u_fft (
.clk (clk),
.reset_n (reset_n),
.start (start),
.inverse (1'b0), // forward FFT for range processing
.din_re (din_re),
.din_im (din_im),
.din_valid (din_valid),
.dout_re (dout_re),
.dout_im (dout_im),
.dout_valid (dout_valid),
.busy (busy),
.done (done)
);
endmodule
For IFFT operation (matched filter synthesis), toggle .inverse(1'b1) and the engine automatically applies 1/N scaling.
Doppler Processor Integration
module doppler_top (
input wire clk,
input wire reset_n,
input wire [31:0] range_data, // {I[31:16], Q[15:0]}
input wire data_valid,
input wire new_chirp_frame,
output wire [31:0] doppler_output,
output wire doppler_valid,
output wire [4:0] doppler_bin,
output wire [5:0] range_bin,
output wire sub_frame,
output wire processing_active,
output wire frame_complete
);
doppler_processor_optimized #(
.DOPPLER_FFT_SIZE (16),
.RANGE_BINS (64),
.CHIRPS_PER_FRAME (32),
.CHIRPS_PER_SUBFRAME (16),
.WINDOW_TYPE (0) // Hamming
) u_doppler (
.clk (clk),
.reset_n (reset_n),
.range_data (range_data),
.data_valid (data_valid),
.new_chirp_frame (new_chirp_frame),
.doppler_output (doppler_output),
.doppler_valid (doppler_valid),
.doppler_bin (doppler_bin),
.range_bin (range_bin),
.sub_frame (sub_frame),
.processing_active (processing_active),
.frame_complete (frame_complete)
);
endmodule
Verified Testbench Pattern
The reference testbench in tb_fft_engine.v demonstrates stimulus generation and completion detection:
module tb_fft_engine;
reg clk = 0;
always #5 clk = ~clk; // 100 MHz
reg reset_n = 0, start = 0;
reg signed [15:0] din_re = 0, din_im = 0;
reg din_valid = 0;
fft_engine #(.N(1024), .LOG2N(10)) dut (…);
initial begin
#20 reset_n = 1;
#10 start = 1;
repeat (1024) begin
@(posedge clk);
din_valid = 1;
din_re = $random;
din_im = $random;
end
din_valid = 0; start = 0;
wait (done);
$display("FFT completed: %0d + j%0d", dout_re, dout_im);
end
endmodule
Complete File Reference
| File | Path | Purpose |
|---|---|---|
fft_engine.v |
/9_Firmware/9_2_FPGA/fft_engine.v |
Generic radix-2 DIT FFT/IFFT engine |
doppler_processor.v |
/9_Firmware/9_2_FPGA/doppler_processor.v |
Staggered-PRF Doppler processor |
matched_filter_processing_chain.v |
/9_Firmware/9_2_FPGA/matched_filter_processing_chain.v |
Range profile generation |
range_bin_decimator.v |
/9_Firmware/9_2_FPGA/range_bin_decimator.v |
1024→64 bin reduction |
tb_fft_engine.v |
/9_Firmware/tests/cross_layer/tb_fft_engine.v |
FFT engine verification |
tb_doppler_cosim.v |
/9_Firmware/9_2_FPGA/tb/tb_doppler_cosim.v |
Doppler cosimulation |
Summary
-
The Range FFT engine (
fft_engine.v) provides a 1024-point radix-2 DIT implementation with quarter-wave twiddle ROM optimization, achieving 75% memory reduction through cosine symmetry exploitation. -
Bit-reversed loading eliminates output reordering, while the four-stage pipeline sustains one butterfly per clock with deterministic
N × LOG2Ncycle latency. -
The Doppler processor implements staggered-PRF processing through dual 16-point FFTs, applying Hamming windowing and sub-frame tagging (
{sub_frame, bin[3:0]}) to resolve velocity ambiguities. -
Both modules expose simple handshaking (
start/busy/done) and require no vendor IP, mapping cleanly to Xilinx BRAM and DSP48E1 resources.
Frequently Asked Questions
What FFT algorithm does the PLFM RADAR implementation use?
The implementation uses a radix-2 Decimation-In-Time (DIT) Cooley-Tukey algorithm with iterative butterfly processing. This choice balances hardware efficiency against algorithmic complexity, avoiding the routing congestion of higher-radix designs while maintaining sufficient throughput for real-time radar processing.
How does the staggered-PRF Doppler processor resolve velocity ambiguity?
By splitting each 32-chirp frame into two 16-chirp sub-frames with different PRI values, the processor generates two distinct Doppler mappings for the same physical velocity. The 5-bit doppler_bin output tags each result with its sub-frame origin ({sub_frame, bin[3:0]}), allowing downstream logic to correlate detections and compute unambiguous velocity through table lookup or arithmetic comparison of the bin indices.
Can the FFT engine be used for sizes other than 1024?
Yes. The N and LOG2N parameters enable any power-of-two transform size. The twiddle ROM must be regenerated for the target size using the quarter-wave symmetry (storing only N/4 values), and the TWIDDLE_FILE init path updated accordingly. The pipelined architecture scales linearly in BRAM and DSP48E1 usage with transform size.
What is the purpose of the range_bin_decimator.v module?
This module reduces the 1024-point range profile output from the matched filter to 64 range bins before Doppler processing. This decimation lowers the memory bandwidth and storage requirements for the Doppler stage by 16×, trading range resolution for system resource efficiency in accordance with the radar's operational requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →