AMD says its FPGA is ready to emulate your biggest chips
- Reference: 1687870896
- News link: https://www.theregister.co.uk/2023/06/27/amd_versal_fpga_emulation/
- Source link:
However, AMD's – formerly Xilinx's – latest Versal FPGAs unveiled Tuesday can do a bit better than simulate a 30-year-old microprocessor. The parts are designed to emulate, test, and debug chips before they've even been built.
Taping a chip out for manufacturing is an incredibly expensive prospect, and even more so if you discover a defect after the fact. Using these FPGAs, chip designers can "create a digital twin or a digital version of their upcoming ASIC or SOC well ahead of silicon tape out," Rob Bauer, senior product line manager for AMD's Versal family, told The Register . "They can verify, they can begin software development much earlier in the design cycle, etc."
[1]
According to Bauer, this is only going to get more difficult for chipmakers as the transition to advanced packaging techniques like 2.5D and 3D chiplet architectures. "If you're a chip designer, no longer are you doing verification and software development for a single die, you're doing it for a multi-die chiplet-based device," he explained.
[2]
[3]
This is where AMD is positioning its Versal Premium VP1902. Measuring roughly 77x77mm, the massive chip boasts 18.5 million logic cells – twice that of the outgoing [4]VU19P – as well as dedicated Arm cores for control-plane operations, and onboard networking to assist with debugging.
The idea here is that by including general compute and networking functionality, less of the FPGA's logic is used up by I/O, debugging or control plane, and more of it for emulating the ASIC or SoC.
[5]AMD seeks luck of the Irish with $135M investment for adaptive computing R&D
[6]Is this the year 100GE NICs go mainstream? If you're into AI, it might be
[7]While Intel XPUs are delayed, here's some more FPGAs to tide you over
[8]Chips in space: Reprogrammable AMD AI SoC cleared for liftoff
In addition to doubling the gate density, AMD says the part also offers twice the bandwidth, which translates into a higher effective cloud rate when emulating silicon. Meanwhile, the chip features a new chiplet architecture that places four FPGA tiles in quadrants, which Bauer says helps to reduce latency and congestion as data moves through the chips.
While all of this might sound impressive, anyone who has spent any time playing with emulation will know it tends to be highly inefficient, slow, and expensive compared to running on native hardware, and the situation is no different here.
[9]
Emulating modern SoCs with billions of transistors is a pretty resource intensive process to begin with. Depending on the size and complexity of the chip, Bauer says dozens or even hundreds of FPGAs spanning multiple racks may be required, and even then clock speeds are severely limited compared to what you'd find in hard silicon.
According to AMD, while just 24 devices are required to emulate a billion logic gates, it can be scaled out to support up to 60 billion gates at clock speeds in excess of 50MHz.
Bauer notes that the effective clock rate does depend on the number of FPGAs involved. "For example, if you had a piece of IP that can live in a single VP1902, you're gonna see much higher performance," he said.
[10]
While AMD's latest FPGA is largely aimed at chipmakers, the company says the chips are also well suited to companies doing firmware development and testing, IP block and subsystem prototyping, peripheral validation, and other test use cases.
As for compatibility, we're told the new chip will take advantage of the same underlying Vivado ML software development suite as the company's previous FPGAs. AMD says it's also working in collaboration with leading EDA vendors, like Cadence, Siemens and Synopsys, to add support for the chip's more advanced features.
AMD's VP1902 is slated to start sampling to customers in Q3 with general availability beginning in early 2024. ®
Get our [11]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZJsHnesgD62FgKj@g2K9lgAAAlA&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJsHnesgD62FgKj@g2K9lgAAAlA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJsHnesgD62FgKj@g2K9lgAAAlA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://www.theregister.com/2019/08/26/hot_chips_roundup/
[5] https://www.theregister.com/2023/06/21/amd_135m_ireland/
[6] https://www.theregister.com/2023/03/10/100g_nics_mainstream/
[7] https://www.theregister.com/2023/03/08/intel_fpga_agilex/
[8] https://www.theregister.com/2022/11/15/amd_versal_space_chip/
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJsHnesgD62FgKj@g2K9lgAAAlA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJsHnesgD62FgKj@g2K9lgAAAlA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[11] https://whitepapers.theregister.com/
Re: Rarely Useful
I'm with you on the painful dev tools. There has to be a better way (or many better ways)
I don't think the RAM limitations are inherently worse than for CPU/GPU, are they? If customers want masses of RAM, I'm sure that Xilinx will oblige, probably not on-die, but chiplets of cache in the same package or nice fast DDRx interfaces - nothing that they haven't done before, right down to the cheapie zynq devices.
If the workload is such that a CPU is the right answer, an FPGA won't be, that's not going to change.
Re: Rarely Useful
The thing is that if you start needing off die ram, overall performance starts relying on the memory bus performance. Because GPUs and CPUs are all about memory performance and have a lot of it (Intel are up to, what, 200GBps memory bandwidth per chip? NVidia GPUs go beyond that I think), they win.
FPGAs are torn between providing acres of programmable logic and using the silicon real estate for things like DDR5 memory interfaces. CPUs and GPUs don't make the same trade, because they store their programs in off-chip memory. Generally speaking, this means FPGAs are pressured to skimp on things like DDR4 interfaces.
FPGAs have always been slow but potentially heavily parallel devices, and become just slow devices as soon as something forces them to be more (or fully) sequential (like a lack of memory interfaces).
Meanwhile, CPUs and GPUs have become much more parallel with specific instructions. SSE / AVX in Intel AMD CPUs are pretty good. Altivec in PowerPC / Cell was epic for the day.
To compete at all, FPGAs have had to include hard cores for certain DSP routines like FMA, and similarly now for AI applications, because doing the same thing in programmable logic cells is very slow indeed. I think that rather spoils the point of the programmable logic because all it is doing is pushing data in and out of hard cores. In comparison, all that CPUs and GPUs are doing is pushing data in and out of their own SIMD vector units / cores. The difference is that programming the logic is hard and it runs at, say, 400MHz, whilst the software is easy and can easily clock along at 4GHz. FPGAs needs a lot more hard DSP core hardware to be competitive. Thus they become large and expensive chips.
Also, FPGAs tend not to be keen on floating point (maybe they've got better?), and AVX / Altivec and GPUs love chewing through floating point arithmetic.
Re: Rarely Useful
I'm no FPGA developer, so could be very wrong, but it seems to me FPGAs are good for short product runs that require custom silicon, whether that silicon is a new design, perhaps in testing, or an emulation of an old design (for retro computing, or controlling old machinery).
It does seem as though once the number of units required gets quite high, you'd be better off looking at making actual chips.
Sometimes necessary
Had just read an article where FPGAs were needed to do impossible math: [1]Ninth Dedekind number discovered: Scientists solve long-known problem in mathematics . Kinda funny that the supercomputer was 'merely' the host for the FPGAs.
[1] https://phys.org/news/2023-06-ninth-dedekind-scientists-long-known-problem.html
Rarely Useful
Interesting to see the same old use cases being rolled out, again. Quite a lot of them are not very mass-market, much more niche-market.
FPGAs remain ******* hard to get working effectively. There's a lot of talk of direct model synthesis, but that's never very good. Someone doing some proper coding gets better results, and the number of people who actual work in VHDL or Verilog and are good at it is surpassing few. You've really got to need to use an FPGA, to justify using one. Software on a CPU is a lot easier.
They're also expensive, hot, and prone to having insufficient on-chip resources. If you thought Apple Sillicon had expensive on-chip RAM, wait until you're paying Xilinx / Altera prices for it. They do not have that much on-chip storage - 100sMByte tops - and as soon as that becomes insufficient then you're better off with a CPU and GPU. The thing I think is interesting is that with these monster FPGAs, they're firmly in the land of large and hot; i.e. a large GPU or CPU is volumetrically and thermally competitive. The problem is that for computational applications, the GPU or CPU can easily out perform an FPGA if the problem data set requires data storage in off-FPGA RAM (GPUs especially have huge memory bandwidths in comparison), and are a whole lot cheaper to buy.
That's the benefit of there being a huge consumer market for CPUs and GPUs - the cost of the chips gets to be quite low. With very little in the way of a market (in comparison) for FPGAs, they're always an expensive solution to a problem. There is a very small range of problems where it's worth it.