BxB Logo BxBFFT vs Dillon Engineering FFTs

Summary of the Comparison

The BxBFFT far surpasses the Dillon Engineering (DE) FFTs in resource utilization. The one piece of data that DE has made publicly available shows that the DE FFT_Pipe uses more than 4 times the LUTs and more than 3 times the REGs of a comparable BxBFFT.

The BxBFFT's achievable clock speed is also higher, with Vivado measurements showing significantly higher FMax than is claimed by DE. Because of this higher FMax, BxBFFT designs will close timing when DE FFTs do not.

Power consumption of a BxBFFT will also be significantly lower than a DE FFT, in part because of the lower resource count.

The resource advantage can also be leveraged to build BxBFFTs that achieve significanly higher levels of parallelism than what is possible with DE FFTs. Higher parallelism gives the BxBFFT higher achievable throughput and lower achievable latency than the DE FFTs can provide.

In addition, the BxBFFT appears to have a much wider range of features supported out-of-the-box. For example, the BxBFFT has options for real-to-complex FFTs, pipelining controls, memory type controls, and controls for automatic generation of twiddles. Dillon Engineering doesn't advertise these. It's possible DE can supply some of these, but presumably as a cost-added consulting job.

The DE FFT's advantages appear to be that DE has more experience moving FFT designs to ASICs, and that DE has more capability for ultra-long FFTs with a customized intermediate staging memory. However, BxB is also quite willing to take on these customization tasks, which should produce better results than DE would ordinarily provide, as seen from the BxBFFT performance data below.

Comparison Data

Evaluation of Dillon Engineering (DE) FFTs must rest on the data that DE has publicly provided. Most public DE data is antique, dating to 2008 on the Virtex 5. However one piece of information is from their recent re-release of its FFT_Pipe FFT, found here.

This DE FFT_Pipe information is missing important details, such as the bit width that was used and whether a bit-reverse was included. The comparison below assumes what should be the worst case: a BxBFFT with a bit reverse and with a large bit width of 27 bits. (Sizing for an 18-bit BxBFFT is also included to show how the numbers scale.)

The table below gives this comparison of the DE FFT_Pipe FFT to a matching BxBFFT:

1024-point FFT, processing one complex point per clock
Resource DE FFT_Pipe BxBFFT 18-bit BxBFFT 27-bit
LUTs ~9800 1933 2210
Flip-Flops ~15000 4610 6225
BRAM (18Kb) 20 17 27
DSP48 20 12 12
Setup FMax MHz ~500 775.8 775.8
FMax MHz ~500 637.3 637.3

The BxBFFT clearly blows Dillon Engineering's FFT_Pipe out of the water. But it's not the only FFT to do so. Most other FFT implementations are significantly better on LUTs and FLOPs than these numbers for the DE FFT_Pipe FFT.

One might wonder whether the DE FFT_Pipe is an anomaly. Perhaps DE provided data from a case where they do especially poorly. Perhaps DE is really giving data from a floating point case. Perhaps DE has enabled their AXI-Lite FFT control module with these numbers. Perhaps DE's other FFTs are better.

Any of these are possible, since DE provides insufficient public data to rule them out. All but the last would be very poor marketing. The last would be poor engineering -- for development, debugging, testing, and maintainability it makes sense to have significant common code between multiple FFT implementations, rather than designing and maintaining them all in isolation.

In addition, it may be that the DE's FFT_Pipe may be incorporated into some of the other DE FFT offerings. So its inefficiency would reflect directly on them. For example, to obtain an FFT processing P points in parallel, DE may use P FFT_Pipes, followed by a twiddle stage, a P-point parallel radix stage, and then a bit reverse. This may be how DE designs their FFTs with higher parallelism.

BxB recommends taking DE marketing claims with a grain of salt. For example, on this page Dillon Engineering makes the apparently false claims that their FFT cores are "up to 3x faster than the nearest competitor", with "20-40% less FPGA resources required". No data is supplied to support these claims.

Based on the one piece of data that DE has provided, Dillon Engineering FFTs use more resources than most competitors, not less.

Dillon Engineering's "3x faster" claim seems especially outrageous. In the extensive testing BxB has done on multiple FFTs, there are occasionally cases where the FMax of the fastest FFT (almost always the BxBFFT) is 3x faster than the FMax of the worst FFT. However, this almost never happens vs the nearest competitor. DE's heavy FFT_Pipe resource usage also make it unlikely that DE could exceed the parallelism achievable with the best FFTs, and especially not the BxBFFT.

Conclusions

It is hard to make 100% certain conclusions from a comparison where Dillon Engineering has provided so little data. However, one must assume that this data puts DE's best foot forward.

If the data that Dillon Engineering has provided shows the best performance of DE's FFT_Pipe, it is extremely far behind the BxBFFT. Thus DE FFTs can be expected to show poor performance in LUTs, REGs, power usage, Fmax, and in meeting timing. Parallel-processing built from FFT_Pipe or using similar code can be expected to have trouble fitting in the FPGA at much lower levels of parallelism than the BxBFFT.

In addition, the BxBFFT appears to support a much wider selection of options and user controls to meet algorithmic and implementation needs.

Dillon Engineering marketing appears to be prone to hyperbole and a lack of hard data. Before trusting that DE has a good solution, BxB recommends that the DE solution should be compared to other implementations. Resource, timing, and power data for BxBFFTs can be supplied almost instantly to assist with such comparisons.

Links

Bit by Bit Signal Processing Main Page
BxBFFT Product Main Page
BxBChan Product Main Page