FPGA Families and Choosing the Right FPGA
1. The FPGA Is the Heart of the Project
What is an FPGA, briefly?
An FPGA (Field-Programmable Gate Array) is a logic array that can be configured in the field, after manufacturing. It does not run a program like a processor; based on the hardware description you write, it becomes the circuit itself. Inside it are LUTs, flip-flops, DSP blocks, block RAM and the programmable interconnect that ties them together. For the fundamentals, see the first article in the series.
Why is choosing the right FPGA a critical decision?
The first technical question at the start of a project is often the wrong one: “Which FPGA should we use?” The right question is: which corner do my requirements push me into? Because even the cleanest RTL can’t be rescued on the wrong device. A wrong choice comes back in three typical ways:
- Cost: an oversized device inflates the budget both through unit price and paid tool licenses.
- Lack of performance: without enough transceivers, DSPs or the right hard block, the design never meets its target.
- Power and heat: the wrong segment drains the battery or forces cooling onto the board.
All of this is decided in the first decision of the design, before a single line of code is written.
2. The Major Players Driving the Market
Before choosing a device, understand the vendor’s philosophy; because picking a vendor is often an ecosystem decision, not a technical one.
AMD
Market leader. The broadest product range and the largest ecosystem: Vivado for synthesis/implementation, Vitis for embedded and HLS, a ready IP catalog and a large community. It spans from the low-cost Spartan to the AI Engines in Versal under one roof. This breadth is why the “default choice” reflex so often goes to AMD.
Altera
The other giant. Its toolchain is Quartus; its center of gravity is the data center and high-performance computing. At the high end the Stratix and current Agilex 7/9 families offer strong transceivers and memory bandwidth, while Agilex 5 and Agilex 3 provide a current line in the mid and low/edge segments. At RunX we are an Altera design partner, which means direct access to current tool releases, the roadmap and hands-on engineering support.
Microchip
Its focus is not the performance race but reliability, low power, radiation tolerance and security. Its toolchain is Libero SoC. A significant part of its devices are flash-based which means instant boot, low static power and higher immunity to single-event upsets (SEU).
Note (Libero): Microchip’s flow differs a little from the others, and being prepared for it up front saves time. Synthesis runs through Synplify Pro and the flow can diverge from other tools at the EDIF/netlist stages; on first setup, the license server and netlist-flow settings are what take the most time. On the other hand, on a flash-based PolarFire/IGLOO2 it is a big convenience to program the design straight to the chip and see instant boot with no external config memory; the fact that static power comes out markedly lower in the power analysis changes the whole calculus for battery and space applications.
Lattice Semiconductor
The address for ultra-low-power, small-form-factor and cost-focused solutions. Tiny devices like the iCE40 are small and cheap enough to replace a microcontroller; open-source toolchains (Yosys/nextpnr) exist for iCE40/ECP5. On the current Nexus platform, CertusPro-NX offers a low-power option in the mid segment, while the Avant family reaches higher logic and interface needs with hardened PCIe and 25 Gb/s SerDes. At RunX we are a Lattice Worldwide Design Partner, with early access and direct support across these families.
3. An Overview of the Main FPGA Families
The way to look at this without getting confused is to split each vendor’s line into a few segments.
Entry level / low cost
AMD Spartan/Artix, Altera Cyclone, Lattice iCE40/MachXO. Small devices ranging from a few thousand to a few hundred thousand logic cells, with few or no transceivers. Concrete examples:
- AMD Artix-7 (e.g. XC7A200T): ~215,000 logic cells, 740 DSP slices, ~13 Mb block RAM and GTP transceivers up to 6.6 Gb/s. Spartan-7 has no gigabit transceivers and tops out around ~102,000 cells for pure logic and glue work.
- Altera Cyclone IV/V / Cyclone 10: Roughly the ~150,000–300,000 LE band; up to 12.5 Gb/s transceivers on Cyclone 10 GX. On the edge side, the current counterpart is Agilex 3.
- **Lattice iCE40 / Certus-NX: **ICE40 at ~1,000–8,000 LUTs, milliwatt power and tiny packages; on the current Nexus side, Certus-NX is a more capable, still low-power alternative.
The right place for sensor interfacing, glue logic and simple protocol bridges.
Mid segment
AMD Kintex, Altera Arria, Microchip PolarFire. Hundreds of thousands of LUTs, hundreds of DSPs, mid-speed transceivers and hard PCIe/DDR controllers:
- AMD Kintex-7 (e.g. XC7K480T): ~478,000 logic cells, ~1,920 DSP slices, GTX transceivers up to 12.5 Gb/s. In the UltraScale generation (Kintex UltraScale) this rises to 16 Gb/s+ lanes and cell counts in the millions.
- **Altera Arria 10 / Agilex 5: Arria 10 up to ~1.15 million LEs, ~3,000 DSP blocks, 17.4 Gb/s (up to ~25.8 Gb/s on GT) transceivers. In the current line, Agilex 5 ** with a hard-processor option and PCIe Gen4/5 is a more efficient mid-segment alternative (common on SoMs).
- Microchip PolarFire (e.g. MPF500T): ~481,000 logic elements, ~1,480 math blocks, transceivers up to 12.7 Gb/s all of it flash-based and low power.
- Lattice CertusPro-NX (Nexus): Tens of thousands of logic cells, SerDes up to 10 Gb/s and embedded memory; for a low-power mid segment.
Most serious signal processing and video pipelines are comfortable in this band.
High end / high performance
AMD Virtex, Altera Stratix/Agilex. Millions of LUTs, thousands of DSPs and very high-speed serial links:
- AMD Virtex UltraScale+ (e.g. VU13P): ~3.8 million logic cells, ~12,288 DSP slices, GTY transceivers at 32.75 Gb/s; on-package 8–16 GB HBM2 memory on the HBM variants.
- Altera Agilex 7 / 9: Millions of LEs and up to ~11,500 DSP blocks; transceivers up to 116 Gb/s on** Agilex 7.** The previous-generation Stratix 10 is still used at the high end (up to 58 Gb/s PAM4).
- Lattice Avant: Lattice’s largest family; reaches the mid-to-high segment with hardened PCIe, 25 Gb/s SerDes and low power.
In return it demands high power, a challenging board design and, often, a paid tool license.
SoC (System on Chip) solutions
Those with a hardware processor on board: AMD Zynq / Zynq UltraScale+, Altera Cyclone V SoC / Agilex SoC, Microchip SmartFusion2 / PolarFire SoC.
- AMD Zynq-7000 (e.g. Z-7020): dual-core Arm Cortex-A9 + ~85,000-cell PL. Zynq UltraScale+ MPSoC: quad Cortex-A53 + dual R5 real-time cores, up to ~1.1 million cells of PL and transceivers up to 32.75 Gb/s.
- **Altera Agilex SoC / Cyclone V SoC: **An Arm-based hard processor system + fabric (including Agilex 5 SoC in the current line).
- Microchip PolarFire SoC (e.g. MPFS250T): a quad-core RISC-V (U54) plus a monitor core, ~254,000 logic elements, flash-based.
It places a real processor next to the fabric; the Linux control software lives on the processor, the deterministic datapath in the fabric.
Moving up a segment isn’t just about getting “more LUTs”; power, board complexity, tool licensing and engineering effort come with it.
4. Criteria for Choosing the Right FPGA for Your Project
This is the most important section. Choose the device from requirements with a checklist, not from the catalog. Order matters: transceivers and hard blocks eliminate a device before capacity even enters the picture.
Logic element & cell capacity
How big is my design, how much area do I need? Estimate it, then add headroom on top. A good rule: pick a device you’ll stay within 70–80% of. Filling an FPGA to 95% makes it nearly impossible for the place-and-route tools to close timing; the remaining 20–30% headroom is not a luxury, it’s a precondition for timing closure.
I/O (input/output) needs
How many pins do I need, and in which electrical standards (LVDS, MIPI, voltage levels)? Do I need high-speed serial links (Gigabit Transceivers PCIe, JESD204B, SDI, etc.)? Even if total capacity looks sufficient, the device is eliminated if there aren’t enough transceivers at the right speed or enough pins in the right bank. So check this first.
Hardware resources (BRAM & DSP)
Will I do image processing, filtering or heavy math? Multiply-accumulate-heavy work lives in the DSP blocks; FFTs, FIR filters and matrix multiplies are measured here. Buffering and pipeline bandwidth are set by block RAM, and that is often the resource that runs out before logic does.
Power consumption and thermal management
Will the system run on battery, can I fit a heatsink? The static vs dynamic power distinction is critical here: flash-based devices (e.g. Microchip) draw markedly lower static power, which changes the math for battery and space applications. High-end SRAM-based devices, on the other hand, generate serious heat at full load and may require active cooling on the board.
Physical size and package
What is my PCB design capability? Package choice directly affects board cost and manufacturability: BGA packages offer high density but are hard both to solder and to route (many layers, via-in-pad); leaded packages like QFP are much easier but limited in pin count and speed. For a small team or a fast prototype, the package is often as decisive as the device.
Supply chain and product lifetime
In industrial and defense production, a device’s shelf life matters as much as its specs. What is negligible for a lab prototype can decide the program for a product going into serial production. Vendors publish longevity (longevity / PDN) programs; some families come with a supply commitment of 10–15 years and more. Look at: the family’s long-term support commitment, the availability of industrial and military temperature grades, current lead times, and, where possible, a second source. A design portable across vendors is cheap insurance against supply-chain disruptions and should be planned at the start of the program.
Ecosystem, IP cores and time-to-market
There are weeks sometimes months of difference between writing a feature from scratch and integrating a ready, verified IP core. A mature ecosystem a broad IP catalog, reference designs, evaluation boards and a strong community directly shortens time-to-market. By the same logic, having functions like PCIe/DDR/Ethernet as hard blocks removes soft-IP effort entirely. In industrial production, time is money directly; reusable and vendor-portable IP both speeds up the first delivery and pays off again in later products.
Capacity bottleneck (future flexibility / headroom)
Beyond the 20–30% headroom you leave for timing, leave headroom for the product’s future too. A product in serial production evolves over time: new features, field updates, changing standards. A device that fits perfectly today can become a bottleneck in the next revision. A practical method: choose a family that hosts several die sizes in the same package. That way, when you later need a bigger chip, instead of redesigning the board you can drop a pin-compatible larger device into the same footprint. This is the cheapest way to avoid a board respin.
5. Development Tools (Toolchain) and License Costs
The software ecosystem is often as important as the hardware itself; because when you choose the device, you also choose the toolchain.
- AMD Vivado / Vitis: The free edition covers a broad range of devices, but the largest UltraScale+/Versal devices often require a paid license. The ecosystem and documentation are the richest.
- Altera Quartus Prime: The Lite edition is free and covers the Cyclone/MAX families; Stratix, Agilex (7/9) and large Arria devices require the paid edition. Current Agilex 5/3 support is in the newer Quartus releases.
- Microchip Libero SoC: Supports a broad range including PolarFire with a free license; the flow is a bit different (Synplify Pro synthesis) but practical on the programming and power-analysis side for flash-based devices.
- Lattice Radiant / Diamond: Radiant covers the current Nexus families (Certus-NX/CertusPro-NX) and Avant, while Diamond serves ECP5/MachXO; an open-source chain (Yosys/nextpnr) is also possible for iCE40/ECP5.
Practical upshot: on a small program, a paid tool license can cost more than the device itself. Put the license cost in the budget when you choose.
6. FPGA Recommendations by Example Scenario
Students and hobby projects
If the goal is to learn and iterate quickly, a cheap development board is the right move: an AMD Artix-7, Altera Cyclone IV or Lattice-based (iCE40/ECP5/Certus-NX) board. These are both economical and run on free (and, on iCE40, open-source) toolchains; the barrier to entry is low.
AI and image processing
You need plenty of DSP, enough memory and usually a processor. For real-time image processing, for example, an SoC like Zynq / Zynq UltraScale+ or a high-DSP device is a natural candidate: control and interfacing on the processor, deterministic processing in the fabric. A concrete sense of scale: 4K30 uncompressed video means a few Gb/s of raw data on the wire (8-bit 4:2:0 ≈ 3.7 Gb/s; ~5 Gb/s at 10-bit 4:2:2), so a dual-stream video pipeline needs enough transceivers and DSP but can often be solved in the mid segment without going high end. Likewise, a system that has to process high-speed ADC data (over JESD204B) while running Linux at the same time points to an SoC family.
Space/aerospace or low power
If the priority is reliability, low static power and radiation tolerance, Microchip PolarFire or IGLOO2 is preferred. Their flash-based structure needs no external config memory, boots instantly and is more resilient to SEUs; blocks like hard PCIe also remove soft-IP effort. Here capacity is secondary; the config memory type and quality grade are the primary axes.
7. Conclusion and Summary
There is no perfect FPGA; there is only the “best fit for the need.” The right device is not the most powerful one but the smallest that meets the requirement with headroom because every excess comes back as power, cost and complexity. In short: count the requirements, eliminate first on the mandatory blocks and transceivers, budget capacity with 20–30% headroom, then apply the power, package, supply and tool/license constraints.
One final piece of advice: before committing to a chip, always get a development board from the relevant family and start prototyping. Seeing the toolchain, the boot flow and the real resource usage early is far cheaper than running into a surprise in the middle of the design.
At RunX Technology we make this decision every day across the AMD, Altera, Lattice and Microchip ecosystems; from radar signal processing to video pipelines and DO-254 verification.Our design partnerships with Lattice and Altera give us direct access to vendor support, current roadmaps and early silicon/eval boards in both ecosystems so we can weigh device selection not just against today’s datasheet but against the vendor’s long-term commitments and next-generation plans. From ultra-low-power Lattice families to Altera’s high-performance, high-speed-transceiver devices, this proximity becomes a practical edge beyond the spec sheet.
If you have a requirements list in hand, we can find the right device together.
Talk to us: info@run-x.com | www.run-x.com