AMD, Nvidia, HPE tapped to triple the speed of US weather super with $35m upgrade
- Reference: 1611790603
- News link: https://www.theregister.co.uk/2021/01/27/us_cheyenne_supercomputer/
- Source link:
That system running today, code-named Cheyenne, is four years old, and now its government-funded lab wants a bigger and faster super to forecast natural disasters unfolding on Earth. Uncle Sam shopped around and awarded the contract, worth more than $35m, to Hewlett Packard Enterprise.
[1]
The upgraded supercomputer has yet to be given an official name, and is expected to be up and running in 2022, when it will replace Cheyenne, we're told. Kids in the lab's home state of Wyoming will be asked to propose a name for the thing.
The computer will be based on HPE’s Cray EX (previously Shasta) supercomputer blueprints, and is expected to reach a theoretical maximum performance of 19.87 petaFLOPs – that’s much faster than the 5.34 petaFLOPs Cheyenne [2]today offers .
[3]
“That is almost 3.5 times the speed of scientific computing performed by the Cheyenne supercomputer, and the equivalent of every man, woman, and child on the planet solving one equation every second for a month,” [4]according to NCAR on Wednesday. “Once operational, the HPE-powered system is expected to rank among the top 25 or so fastest supercomputers in the world.”
Once operational, the HPE-powered system is expected to rank among the top 25 or so fastest supercomputers in the world
The new machine will sport 2,570 compute nodes: 2,488 of those will contain AMD’s 7nm third-gen Epyc Milan processors – due to officially launch in March – and 82 remaining nodes will be a mixture of Milan chips and Nvidia’s 7nm A100 GPUs. It’ll have a total RAM capacity of a whopping 692TB. The connectivity between the nodes is powered by HPE’s [5]Slingshot interconnect architecture, boasting a bandwidth of 200Gbps per direction per network switch port.
By using GPUs to accelerate workloads, NCAR will be moving away from its [6]previous CPU-only approach. Although Cheyenne has more 4,032 computation nodes, they’re filled with Intel’s Xeon Broadwell workhorses, and this system design is less efficient, in terms of power and silicon required, than more modern supercomputers that have a mix of CPUs and GPUs to crunch through math operations at high speed. Specialist hardware wins out over generic processors.
Nvidia signs up for an Italian Job: Building for Europe the 'world's fastest AI supercomputer' by 2022 [7]READ MORE
“In terms of hardware, the new system will be hugely helpful to AI and machine learning,” David Hosansky, a spokesperson for NCAR, told The Register .
“The big benefit of GPUs is performing large numbers of computations simultaneously on one chip, resulting in much less power usage and hardware for the same number of parallel operations. GPUs have less on-board memory than CPUs, but the ones being used in this new system are top of the line in terms of both memory and number of cores. That will allow our scientists to load more data and train larger machine learning models than they could before.”
NCAR scientists will use machine-learning algorithms to simulate models of freak weather events, including hurricanes, hail storms, wildfires, and solar storms. These models ingest huge amounts of weather data to output forecasts, and help scientists understand the impacts of climate change.
For example, AI is used to estimate the amount of moisture in vegetation and that data is fed into another model that combines real data from satellite observations to map out regions at risk of wildfires. Weather changes quickly, and these maps are regenerated every day.
"We are now working on using GOES-R satellite data and machine learning to produce hourly fuel moisture content maps over [the US]," NCAR scientist Branko Kosovic, director of NCAR's Weather Systems and Assessment Program, told El Reg . "Machine learning could be used also to develop more accurate and frequent high resolution maps of fuel characteristics. Furthermore, machine learning models could be developed to improve rate of fire spread parameterizations by combining observations and theoretical considerations and developments."
[8]
For more conversations with NCAR staff, check out our sister site, [9]The Next Platform . ®
Get our [10]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_datacentre/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YBJE6@sXBP2FJxB5-B-fTgAAAAw&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://www2.cisl.ucar.edu/resources/computational-systems/cheyenne
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_datacentre/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YBJE6@sXBP2FJxB5-B-fTgAAAAw&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://news.ucar.edu/132774/new-ncar-wyoming-supercomputer-accelerate-scientific-discovery
[5] https://www.nextplatform.com/2019/08/16/how-cray-makes-ethernet-suited-for-hpc-and-ai-with-slingshot/
[6] https://www2.cisl.ucar.edu/resources/computational-systems/cheyenne
[7] https://www.theregister.com/2020/10/15/nvidia_ai_supercomputer_italy_2022/
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_datacentre/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YBJE6@sXBP2FJxB5-B-fTgAAAAw&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[9] https://www.nextplatform.com/2021/01/27/amd-cray-nvidia-behind-massive-ncar-supercomputer-upgrade/
[10] https://whitepapers.theregister.com/
Bandwidth
Er, yeah, that 200Gb/s figure is in no way the entire internal bandwidth of the system. It's basically the base speed per port.
As the linked-to article and this [1]paper [PDF] and HPE's [2]own bumf says, that's the link speed of the interconnect. You can have, eg, a 64-port switch with each port doing 200Gb/s per direction (four lanes of 50Gb/s).
NCAR also said the "HPE Slingshot bandwidth is 200 Gb/sec per port per direction."
C.
[1] https://spcl.inf.ethz.ch/Publications/.pdf/sensi-slingshot.pdf
[2] https://www.hpe.com/us/en/compute/hpc/slingshot-interconnect.html
Any experts in the house?
"It’ll have a total RAM capacity of a whopping 692TB. The connectivity between the nodes is powered by HPE’s Slingshot interconnect architecture, boasting a bandwidth of 200GB per second."
I'm not great at bandwidth <=> storage relationships but does the bandwidth sound like a limiting factor to those in the know? Obviously it depends upon unit of work though.