This page describes how a number of open source FFTs perform relative to the BxBFFT and to each other in AMD/Xilinx Ultrascale+ FPGAs. Comparisons in other AMD/Xilinx FPGAs follow the same general patterns.
To summarize the measurements below, the BxBFFT is at the top in the two most important categories: Fmax and power consumption, as well as most other categories. BxBFFTs also have significant advantages in features, capabilities, options, testing, documentation, examples, and support.
BxB has studied the performance of other FFTs for a particular reason, which is to provide the best possible BxBFFT. An understanding of where these FFTs excel and where they are limited has led to improvements in the BxBFFT. In addition, it is hard to claim that the BxBFFT has exceptional performance if the performance of other FFTs hasn't been fully studied. The plots below show the BxBFFT's exceptional performance.
The open source FFTs chosen for this comparison meet two critera. First, they must be high-speed FFTs that operate on multiple input samples each clock. There are many additional FFTs designed for slower speeds, but they are in a different class and are not comparable. Second, RTL must be obtainable for the FFT, either VHDL or Verilog. So far, only one FFT, the CASPER FFT, has been excluded because RTL was too difficult to obtain. However, one point of comparison with the CASPER FFT was found, and it is included in a section below.
Here's a list of the open source FFTs being compared:
Each of these FFTs is implemented with Vivado to determine achievable Fmax, estimated power consumption, and resources used. The AMD/Xilinx XFFT is also included on the plots for comparison, although there is a separate detailed comparison of the AMD/Xilinx XFFT versus the BxBFFT here. The results are plotted below.
For each plot, the X-Axis of the plot is divided into 6 sections, with different numbers of parallel-processed Points Per Clock (PPC). These vary from PPC2 to PPC64. In each section, FFT size varies from 128 to 262144. This allows a single plot to show a wide overview of FFT performance across a range of FFT Sizes and levels of parallel processing..
| Excellent Performers: | BxBFFT | ||
| Good Performers: | htfft | SGen FFT | Spiral FFT |
| Average Performers: | Xilinx XFFT | ZipCPU FFT | |
| Poor Performers: | Astron FFT | CAStron FFT |
This plots show setup-limited Fmax when each FFT is compiled with nothing else in the FPGA. "Setup-limited" means that the plot doesn't take into account speed limits internal to the DSPs, BRAMs, or URAMs. Achievable speeds in a real design will be lower, both because of these internal limits and because resource contention with other parts of a larger design brings down speeds.
Even when a design has far higher Fmax than is needed, the setup margin has advantages. It gives headroom to absorb timing degradation caused by resource contention from other IP in the FPGA. As a result, designs with higher setup-limited Fmax will close timing more easily than other designs, and will close timing where designs with lower Fmax cannot.
| Excellent Performers: | BxBFFT | SGen FFT | |
| Good Performers: | Spiral FFT | ||
| Average Performers: | CAStron FFT | ZipCPU FFT | |
| Poor Performers: | Astron FFT | htfft | Xilinx XFFT |
This shows that the worst FFTs often use 1.5 times the power (or more) than the best FFTs.
| Excellent Performers: | BxBFFT | ||
| Good Performers: | SGen FFT | ||
| Average Performers: | htfft | Spiral FFT | ZipCPU FFT |
| Poor Performers: | Astron FFT | CAStron FFT | Xilinx XFFT |
This shows that the worst FFTs often use 1.5 times the LUTs (or more) than the best FFTs.
| Excellent Performers: | SGen FFT | Spiral FFT | |
| Good Performers: | None | ||
| Average Performers: | Astron FFT | BxBFFT | htfft |
| Poor Performers: | CAStron FFT | Xilinx XFFT | ZipCPU FFT |
REG performance isn't as important as other resources, since there are twice the available REGs as there are LUTs.
| Excellent Performers: | BxBFFT | |||
| Good Performers: | Xilinx XFFT | SGen FFT | Spiral FFT | ZipCPU FFT |
| Average Performers: | None | |||
| Poor Performers: | Astron FFT | CAStron FFT | htfft |
DSP usage is often not incredibly important, since many FPGAs have large numbers of DSPs. However, for certain problems or certain FPGAs they can be critical.
| Excellent Performers: | CAStron FFT | Xilinx XFFT | |
| Good Performers: | BxBFFT | ||
| Average Performers: | SGen FFT | Spiral FFT | ZipCPU FFT |
| Poor Performers: | Astron FFT | htfft |
Sometimes FFTs with high LUT usage but low BRAM usage have moved some data storage from BRAMs into distributed LUT RAM. This may be the case with the CAStron FFT and the Xilinx XFFT.
The BxBFFT supports a much wider range of features than any other FFT. For example, the BxBFFT supports real-to-complex FFTs, non-power-of-2 FFTs, amplitude management controls, pipelining controls, memory placement controls, and controls for automatic generation of twiddles.
The CASPER FFT currently has no RTL implementation, instead being instantiated through other means such as Matlab Simulink. For this reason, it was too difficult to include it in the extensive plots above. However, at least one comparison point has been published in the literature for CASPER FFT resource usage. This comes from a paper "A Digital Correlator Upgrade for the Arcminute MicroKelvin Imager" by Jack Hickish, Nima Razavi-Ghods, et al. The resources are for a Dual-Band FFT that has a customized layout allowing resource-reduction over a standard CASPER FFT. "Dual Band" indicates that this FFT block takes 2 FFTs simultaneously, sharing resources between the two. The FFT is a real-to-complex FFT with 4096 real input points and 2048 complex output points. It operates at a rate of 16 real input samples per clock and 8 complex output samples per clock. Input is 8 bit, output is 18 bit.
The BxBFFT doesn't have a Dual-Band mode, so this FFT would be replaced by two BxBFFTs. The BxBFFT doesn't operate in Virtex-6, so for comparison the BxBFFT is implemented in a Kintex Ultrascale FPGA, giving the BxBFFT a modest architecture advantage. The processing rate is matched, and also the real-to-complex FFT operation. A full 18-bit BxBFFT is implemented.
The comparison gives the following results:
| FFT | LUTs | REGs | DSPs | BRAM36s |
|---|---|---|---|---|
| CASPER FFT (Dual) | 38206 | 100748 | 736 | 112 |
| 2X BxBFFT | 23346 | 55696 | 196 | 87 |
The difference between these two results is quite significant, and more than can be easily explained by the difference between a Virtex-6 and a Kintex Ultrascale. Thus the BxBFFT appears to significantly outperform the CASPER FFT.
The BxBFFT is faster than all other studied FFTs, and one of the best on power. The BxBFFT is tops in most types of resource usage. The BxBFFT has better controls than other FFTs, and many more capabilities and options. The BxBFFT has extensive documentation, examples, verification, and support.
With the BxBFFT, you do get what you pay for. But if you can't afford anything, there are some other options, and these plots may help you choose.