News: 1677151806

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Unless things change, first zettaflop systems will need nuclear power, AMD's Su says

(2023/02/23)


Within the next 10 years, the world's most powerful supercomputers won't just simulate nuclear reactions, they may well run on them. That is, if we don't take drastic steps to improve the efficiency of our compute architectures, AMD CEO Lisa Su said during her keynote at the International Solid-State Circuits Conference this week.

The root of the problem is that while companies like AMD and Intel have managed to roughly double the performance of their CPUs and GPUs every 2.4 years, and companies like HPE, Atos, and Lenovo have achieved similar gains roughly every 1.2 years at the system level, Su says power efficiency is lagging behind.

Citing the performance and efficiency figures gleaned from the top supercomputers, AMD says gigaflops per watt is doubling roughly every 2.2 years, about half the pace the systems are growing.

[1]

Assuming this trend continues unchanged, AMD estimates that we'll achieve a zettaflop-class supercomputer in about 10-years give or take. For reference, the US powered on its first exascale supercomputer, Oak Ridge National Laboratory's [2]Frontier system , last year. A supercomputer capable of a zettaflop of FP64 performance would be 1,000x more powerful.

[3]

[4]

To AMD's credit, its estimate for when we'll cross the zettaflop barrier is at least a little more conservative than Intel's rather hyperbolic claims that it'd cross that threshold by [5]2027 . What's more, the AMD CEO says such a machine won't exactly be practical unless compute architectures get drastically more efficient and soon.

If things continue on their current trajectory, AMD estimates that a zettaflop-class supercomputer would need somewhere in the neighborhood of 500 megawatts of power. "That's probably too much," Su admits. "That's on the scale of what a nuclear power plant would be."

[6]

"This flattening of efficiency becomes the largest challenge that we have to solve, both from a technology standpoint as well as from a sustainability standpoint," she said. "Our challenge is to figure out how over the next decade we think about compute efficiency as the number one priority."

Correcting course

Part of the problem facing chipmakers is the means they've traditionally relied on to achieve generational efficiency gains are becoming less effective.

Echoing Nvidia's leather jacket aficionado and CEO Jensen Huang, Su admits Moore's Law is slowing down. "It's getting much, much harder to get density performance as well as efficiency" out of smaller process tech.

"As we get into the advanced nodes, we still see improvements, but those improvements are at a much slower pace," she added, referencing efforts to shrink process tech much beyond 5nm or even 3nm.

But while improvements in process tech are slowing down, Su argues there are still opportunities to be had, and, perhaps unsurprisingly, most of them center around AMD's chiplet-centric worldview. "The package is the new motherboard," she said.

[7]

Over the past few years, several chipmakers have embraced this philosophy. In addition to AMD, which arguably popularized the approach with its Epyc datacenter chips and later brought the tech to its Instinct GPUs, chipmakers — including Intel, Apple, and Amazon — are now employing multi-die architectures to combat bottlenecks and accelerate workloads.

Chiplets, argues the AMD boss, will allow chipmakers to address three of the low hanging fruits when it comes to compute efficiency: compute energy, communications energy, and memory energy.

Modular chiplet or tile architectures have numerous advantages. For instance, they can allow chipmakers to use optimal process tech for each component. AMD uses some of TSMC's densest process tech for its CPUs and GPU dies, but often employs larger nodes for things like I/O and analog signaling which don't scale as efficiently.

[8]Intel slashes shareholder dividend by two-thirds as cash crunch bites

[9]Taking notes from AWS, Google prepares custom Arm server chips of its own

[10]US Department of Energy solicits AMD's help with nuke sims

[11]Intel, AMD just created a headache for datacenters

Chiplets also help reduce the amount of power required for communications between the components since the compute, memory, and I/O can be packaged in closer proximity. And when stacked vertically, as AMD has done with SRAM on its X-series Epycs and Intel is doing with HBM on its Ponte Vecchio GPUs, the gains are even greater, the chipmakers claim.

AMD expects advanced 3D packaging techniques will yield 50x more efficient communications compared to conventional off-package memory and I/O.

This is no doubt why AMD, Intel, and Nvidia have started integrating CPUs, GPUs, and AI accelerators into their next-gen silicon. For example, AMD's upcoming MI300 will integrate its Zen 4 CPU cores with its CDNA3 GPUs and a boatload of HBM memory. Intel's Falcon shores platform will follow a similar trajectory. Meanwhile, Nvidia's Grace Hopper superchips, while not integrated to the same degree, still co-package an Arm CPU with 512GB of LPDDR5 with a Hopper GPU die and 80GB of HBM.

AMD isn't stopping at CPUs, GPUs, or memory either. The company has thrown its support behind the Universal Chiplet Interconnect Express (UCIe) consortium, which is trying to establish standards for chiplet-to-chiplet communication, so a chiplet from one vendor can be packaged alongside one from another.

AMD is also is actively working to integrate IP from its Xilinx and Pensando acquisitions into new products. During Su's keynote, she highlighted the potential for co-packaged optical networking, stacked DRAM, and even in-memory compute as potential opportunities to further improve power efficiency.

Is it time to give AI a crack at HPC?

But while there's opportunity to improve the architecture, Su also suggests that it may be time to reevaluate the way we go about conducting HPC workloads, which have traditionally relied on high-precision computational simulation using massive datasets.

Instead, the AMD CEO makes the case that it may be time to make heavier use of AI and machine learning in HPC. And she's not alone in thinking this. Nvidia and Intel have both been pushing the advantages of lower precision compute, particularly for machine learning where trading a few decimal places of accuracy can mean the difference between days and hours for training.

Nvidia has arguably been the most [12]egregious , claiming systems capable of multiple "AI exaflops." What they conveniently leave out, or bury in the fine print, is the fact they're talking about FP16, FP8, or Int8 performance, not the FP64 calculations typically used in most HPC workloads.

"Just taking a look at the relative performance over the last 10 years, as much as we've improved in traditional metrics around SpecInt Rate or flops, the AI flops have improved much faster," the AMD chief said. "They've improved much faster because we've had all these mixed precision capabilities."

One of the first applications of AI/ML for HPC could be for what Su refers to as AI surrogate physics models. The general principle is that practitioners employ traditional HPC in a much more targeted way and use machine learning to help narrow the field and reduce the computational power required overall.

Several DoE labs are already [13]exploring the use of AI/ML to improve everything from climate models and drug discovery to simulated nuclear weapons testing and maintenance.

"It's early. There is a lot of work to be done on the algorithms here, and there's a lot of work to be done in how to partition the problems," Su said. ®

Get our [14]Tech Resources



[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Y-ebsRv@MGhAFEComRX--wAAAFM&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[2] https://www.theregister.com/2022/05/30/us_frontier_supercomputer_ousts_japans/

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y-ebsRv@MGhAFEComRX--wAAAFM&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y-ebsRv@MGhAFEComRX--wAAAFM&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[5] https://www.nextplatform.com/2021/10/27/intel-aims-for-zettaflops-by-2027-pushes-aurora-above-2-exaflops/

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y-ebsRv@MGhAFEComRX--wAAAFM&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y-ebsRv@MGhAFEComRX--wAAAFM&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[8] https://www.theregister.com/2023/02/22/intel_quarterly_dividend/

[9] https://www.theregister.com/2023/02/14/google_prepares_its_own_custom/

[10] https://www.theregister.com/2023/01/31/doe_amd_nuclear_simulation/

[11] https://www.theregister.com/2023/01/19/intel_amd_uptime_cooling/

[12] https://www.theregister.com/2022/07/06/fp8_wont_make_ai_perf/

[13] https://www.theregister.com/2022/12/23/doe_fusion_ai/

[14] https://whitepapers.theregister.com/



Intel-lect

elsergiovolador

The problem with late stage capitalism is that corporations are owned by investment funds.

Now imagine that investments fund owns share in Intel or AMD and also own shares in energy companies.

How are they going to resolve such conflict of interest?

If we push Intel or AMD to lower the power consumption, that will reduce profits of energy companies we are invested in!

So they wheel out the CEO to do a bit of green washing and year later they'll release the same CPU just with a different name slightly different clock maybe a few extra cores.

And oh they are also invested in companies making motherboards.

So new motherboads please!

Et voila! PROFIT

Re: late stage capitalism

Catkin

Is a great phrase to use to tell someone you don't understand capitalism without saying you don't understand capitalism.

For example, how do you prevent the competition from being similarly energy hungry and how do you ensure that it's your energy running your inefficient supercomputer?

Re: late stage capitalism

elsergiovolador

I am writing about late stage capitalism where competition does not exist.

If everything is owned by a handful of investment funds, they'll make sure each company they own don't step on other toes.

Competition and where actual capitalism is concerned lives with SMEs, but corrupt governments ensure they are more and more marginalised.

Big corporations don't want people to start their own independent businesses, first because they get deprived of talent and second, independent business may challenge the funds investment strategy if it grows big enough.

Re: late stage capitalism

Catkin

Do tell about the investment funds that own both Intel and AMD

Re: late stage capitalism

IGotOut

Vanguard, BlackRock and State Street for starters.

You're welcome.

Re: late stage capitalism

Catkin

[citation needed], doesn't look like any of those hold major stakes in both.

Re: late stage capitalism

Charlie Clark

The argument of monopsidic investors in several branches is fairly convincing. It's alleged to be the case in US airline companies (and railroads) and does provide a reasonable explanation for the perceived lack of competition. But, across branches as the poster suggests is going too far IMO.

Re: late stage capitalism

Lis

@Catkin

er, you didn't mention anything about major stakes in your original post did you? Re-read your post.

Re: late stage capitalism

Catkin

Initially, I said own but I thought that was too narrow and mentioned major because those are the ones that they disclose up front (i.e. not precluding minor but less well documented). Any evidence at all of cross investment would be a nice starting point.

Re: late stage capitalism

elsergiovolador

Ever heard of the term duopoly?

Re: late stage capitalism

Catkin

Sorry, we're talking about "late stage capitalism" here where, I'm told, competing corporations don't exist.

Re: late stage capitalism

Roland6

Which is what you have with a duopoly…

To have ‘competition’ you need at least two companies to chose from. However, if you don’t buy enough from one, it will cease to trade result8ng in a monopoly; which probably end stage capitalism.

Old solution

Flocke Kroes

Back when I was a PFY the power problem was anticipated and the basis of a final solution dreamed up. If it was patented the patents should have expired years ago.

Currently when you want to output a 1, you charge a wire up to a low voltage. When you want a 0 you discharge the wire (effectively a tiny capacitor) back to 0. That is where most of the power gets wasted. To fix it, replace every wire with two superconducting wires at opposite voltages. To switch to the opposite state, connect the wires with a transistor. One wire discharges into the other. Inductance keeps the current flowing until the voltages swap then you switch the transistor to its non-conducting state.

Re: Old solution

Neil Barnes

And of course, since the practical superconductors work at liquid nitrogen temperatures or lower, no heating problem!

Er, wait...

Ok, rethink on a more practical note. Perhaps it's just the right time for combined power, heat, and computing facilities?

Re: Old solution

Charlie Clark

Work on optical components looks the most promising and has already started.

This pitch sounds very much like a plea for government cash.

Re: Old solution

Flocke Kroes

A superconducting laptop would be impractical: switch on and wait an hour to get down to temperature. Superconductors in a data centre are much less impractical.

Back in the far distant past, logic gates were made out of bipolar junction transistors. Long before scale would have led to melted chips with BJTs IBM researched CMOS. Many years ago CPUs used aluminium to connect transistors. When it became clear that would become a limitation AMD researched the switch to copper. FinFETs were conceived a long time ago but decades later the manufacturing difficulties were worth the pain to reduce leakage currents (Intel call them 3D transistors because changing the name works around prior art invalidating patents).

I do not know if we are several years or decades from superconducting connections. Carbon nanotubes might delay the need for a while if someone can work out how to put a few billion in the right places. The power problem has turned up again and again as the number of transistors increased. One day the only solution will left will be to zero out resistance then super conductors will be assimilated into the collective.

JohnTill123

"This flattening of efficiency becomes the largest challenge that we have to solve, both from a technology standpoint as well as from a sustainability standpoint," she said. "Our challenge is to figure out how over the next decade we think about compute efficiency as the number one priority."

Way back in the days when Xerox made mainframes ( https://en.wikipedia.org/wiki/SDS_Sigma_series ), programmers used to work to optimize their code. This was required to fit into memory and to finish in a useful time.

The modern predilection to throw together masses of "libraries" into a bloated, inefficient, and slow binary and calling it "enterprise-class programming" is sad. It is a primary cause of wasted time and electricity. Go back to writing tight, efficient code and that will improve "compute efficiency" far more than just by throwing hardware at it.

Anonymous Coward

As compute power has increased, the "intelligence" applied to modelling often shows little signs of evolution. When a simple solution is good enough; one does not need a solution with 50,000 extra steps. In my own line of work, empirical models are well established and tested for thermal properties for many conditions.

In niche conditions that don't fit the empirical models I will sometimes roll out a finite-element analysis monster to do the job instead; at enormous difference in cost and time of solution.

In practise, one could still use the empirical model with a "fudge safety factor" with absolutely no concerns. The FEA's many decimal places decidedly and pretty rainbow graphs are misleading indicators. Analysis of the uncertainties around both models will confirm that both methods are "good enough". Doing the FEA to confirm Empirical-with-fudge agree is a useful exercise. Doing the FEA over and over again, is not.

The models being dumped onto supercomputers now; not least of which climate change models, offer negligible additional insight that the underlying evidence already offers. Extrapolating out from evidence that tells you nothing that was not already known - only that a range of possible forecast outcomes exists. I won't be drawn into a debate on the validity of the climate change modelling (it's most assuredly valid!) but I absolutely WILL question the utility of throwing power at a problem whose solution cannot be created by playing with ever-fatter supercomputers - the solutions lie in dragging the world off coal and oil in spite of the kicking and screaming.

This is little different to the nuclear weapons simulations that are a favourite of the supercomputing complex. You can throw infinite compute at the problem, but to REALLY know if your models are accurate; relying on inputs from 30 year old tests becomes increasingly uncertain.

I'm sure this post will wind up the usual commenters...

Anonymous Coward

Can't they just add a turbo button to current supercomputers?

Catkin

But there's a risk that the turbo button slows it down (it's not certain, manufacturers generally used an on position/light to indicate a restricted clock speed but this was far from universal).

Add the turbo button...

MiguelC

...but make sure it's [1]far away from the Reset button !

[1] https://en.wikipedia.org/wiki/File:Casebuttons.jpg

Re: Add the turbo button...

Anonymous Coward

Best not to try and turn it off after being turned on.

For any millennials not in on the joke

Flocke Kroes

Back in the [1]stupid ages computers had a button that reduced the CPU speed to make games slow enough to play. Putting it between the power and reset buttons is yet more evidence that the we worked hard to deserve to be ridiculed mercilessly by future generations.

[1] https://theinfosphere.org/Stupid_Ages

Wirth's Law

Fruit and Nutcase

Wirth's law is an adage on computer performance which states that software is getting slower more rapidly than hardware is becoming faster.

[1]https://en.wikipedia.org/wiki/Wirth's_law

Some Variations:

"What Intel giveth, Microsoft taketh away"

Gates's law "The speed of software halves every 18 months"

May's law "Software efficiency halves every 18 months, compensating Moore's law"

[1] https://en.wikipedia.org/wiki/Wirth's_law

Drastic steps ?

Pascal Monett

I'm sorry, since the first 8086 graced our desktops, computing power has been (de)multiplied by more than you can count.

The 8086 had [1]29,000 transistors in a single core. Today's CPUs have so many billions of transistors they're not even counted any more (okay, 6+ billion), and [2]cores by the dozen .

How much more drastic can you get ?

We've gone from an architecture that did one thing at a time in the single available core, to an architecture that delegates computing to the best platform, stores terabytes of data faster than I can drink my whisky, does real-time raytracing on a 3840 x 2160 pixel screen and all that managed by a chef d'orchestre who can juggle 24 threads at the same time and calculate the millionth digit of Pi before I've finished my morning coffee.

The only drastic step left is R2-D2.

Call me when he rolls off the production line.

[1] https://www.extremetech.com/computing/270926-happy-40th-anniversary-to-the-original-8086-and-the-x86-architecture

[2] https://www.intel.com/content/www/us/en/products/docs/processors/core/13th-gen-core-desktop-brief.html

Bus error -- please leave by the rear door.