News: 1708358526

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

What's going on with Eos, Nvidia's incredible shrinking supercomputer?

(2024/02/19)


Analysis Nvidia can't seem to make up its mind just how big its Eos supercomputer is.

In a blog post this month [1]re-revealing the ninth most powerful supercomputer from last fall's TOP500 ranking, the GPU slinger said the system was built using 576 DGX H100 systems totaling 4,608 GPUs. That's about what we expected from the system back when it was first [2]announced .

While impressive in its own right, that's less than half the number of GPUs Nvidia claimed the system had back in November when it took to the net to talk up its performance in a variety of MLPerf AI training benchmarks.

[3]

Back then, the big iron [4]boasted a complement of 10,752 H100 GPUs, which would have spanned 1,344 DGX systems. With nearly 4 petaFLOPS of sparse FP8 performance per GPU, that super would have been capable of 42.5 exaFLOPS of peak AI compute. Compare that to the 18.4 AI exaFLOPS Nvidia says the system is capable of outputting today, and Eos appears to have lost some muscle tone.

[5]

[6]

Oh, and if you're not familiar with the term AI exaFLOPS, it's a metric commonly used to describe floating point performance at lower precision than you'd typically see in double-precision HPC benchmarks, like LINPACK. In this case, Nvidia is arriving at these figures using sparse 8-bit floating point math, but another vendor, like Cerebras, might calculate AI FLOPS using FP16 or BF16.

So where did the other 60 percent of the system go? We put this question to Nvidia and were told "the supercomputer used for MLPerf LLM training with 10,752 H100 GPUs is a different system built with the same DGX SuperPOD architecture."

[7]

"The system ranked number nine on the 2023 TOP500 list is the 4,608 GPU Eos system featured in today's blog post and video," the spokesperson added.

Except that doesn't appear to be true either. Eos's TOP500 [8]score of 121 petaFLOPS of FP64 out of an estimated peak of 188.65 is too low. The latter should be somewhere between the 275 petaFLOPS [9]originally claimed and the 308 petaFLOPS of FP64 that Nvidia's spec sheet says 4,608 H100s should actually net you.

So while Nvidia hasn't admitted exactly how many GPUs were used, based on these performance figures we can estimate the November run was made using somewhere between 2,816 and 3,161 GPUs.

[10]

Nvidia's decision to put forward a smaller version of the system on last fall's TOP500 ranking, when it had already demonstrated a much larger Eos cluster, strikes us as odd.

With more than ten thousand H100s on board, the larger Eos config would have boasted 720 petaFLOPS of peak double-precision performance. Granted, real-world performance would have been a fair bit lower.

[11]Oxide reimagines private cloud as... a 3,000-pound blade server?

[12]Someone had to say it: Scientists propose AI apocalypse kill switches

[13]Microsoft says it'll throw €3.2B at AI ops in Germany

[14]Upstart retrofits an Nvidia GH200 server into a €47,500 workstation

We asked Nvidia for clarification on these discrepancies and were told that the timeline didn't permit for a TOP500 run on the larger system. Why? They didn't say. "Our teams are racing towards GTC and are not able to provide more details on last year's TOP500 submission at this time," a spokesperson told The Register .

Having said that, Nvidia wouldn't be the only one that couldn't get a full run of their machine done in time. Argonne National Laboratory's Aurora supercomputer, the flagship for Intel's Xeon and GPU Max families, was only able to manage a [15]partial run too. This suggests we may catch a glimpse of an even more powerful Eos system on this spring's TOP500.

If we had to guess, Nvidia may have run into trouble with stability on the full cluster. The LINPACK benchmark is, as you might have guessed, quite the stress test for any system, let alone one assembled from hundreds of GPU nodes.

In any case, Eos's shapeshifting does highlight one of the conveniences associated with Nvidia's modular DGX SuperPOD architecture. It can be scaled out and broken into chunks depending on what it's needed for.

Each SuperPod is made up of what Nvidia calls scalable units (SU) containing 32 DGX H100 nodes containing eight GPUs connected via its 400Gb/s Quantum-2 InfiniBand network. Additional SUs can be added to scale the system to support larger workloads. Officially, Nvidia supports up to four SUs per pod, but the company notes that larger configurations are possible, which is clearly the case with Eos.

As for how big Eos really is, it appears the answer to that depends entirely on how big Nvidia wants it to be at any given moment. ®

Want more analysis? Getting your hands on a H100 GPU is probably the most difficult thing in the world right now – [16]even for Nvidia itself .

Get our [17]Tech Resources



[1] https://blogs.nvidia.com/blog/eos/

[2] https://www.nextplatform.com/2022/03/23/nvidia-will-be-a-prime-contractor-for-big-ai-supercomputers/

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZdOJNWW47fMNOW@9pnRObwAAABE&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://blogs.nvidia.com/blog/scaling-ai-training-mlperf/

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZdOJNWW47fMNOW@9pnRObwAAABE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZdOJNWW47fMNOW@9pnRObwAAABE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZdOJNWW47fMNOW@9pnRObwAAABE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[8] https://www.top500.org/system/180239

[9] https://nvidianews.nvidia.com/news/nvidia-announces-dgx-h100-systems-worlds-most-advanced-enterprise-ai-infrastructure

[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZdOJNWW47fMNOW@9pnRObwAAABE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[11] https://www.theregister.com/2024/02/16/oxide_3000lb_blade_server/

[12] https://www.theregister.com/2024/02/16/boffins_propose_regulating_ai_hardware/

[13] https://www.theregister.com/2024/02/16/microsoft_germany_ai/

[14] https://www.theregister.com/2024/02/14/german_gh200_workstation/

[15] https://www.theregister.com/2023/11/13/aurora_top500_no2/

[16] https://www.nextplatform.com/2024/02/15/half-eosd-even-nvidia-cant-get-enough-h100s-for-its-supercomputer/

[17] https://whitepapers.theregister.com/



Cos nobody gives a flying fsck

Yet Another Anonymous coward

In the good old days Cray would announce their latest bug box and everyone below it on the list would buy one and everyone would laugh at the foreigners for coming in at number 99.

Now nobody cares if your shed full of GPUs is bigger than some other government labs shed full of GPUs.

Re: Cos nobody gives a flying fsck

b0llchit

Of course you care about the sheds full of power sucking devices.

You only get to boast your numbers and performance if it is scaled by the amount of "free" power you are able to extract combined with the amount of cash registers you are allowed to plunder at nanosecond intervals.

Good news is just life's way of keeping you off balance.