New Chinese exascale supercomputer runs 'brain-scale AI'
- Reference: 1647013651
- News link: https://www.theregister.co.uk/2022/03/11/china_exascale_sunway_ai/
- Source link:
While there have been few architectural details to date, a [2]paper [PDF], published today, outlines the compute, memory, and other aspects, in addition to showing off the capabilities via a system-spanning AI workload for a pre-trained language model with 14.5 trillion parameters with mixed precision performance of over one exaflop.
The system has "as many as 96,000 nodes" the paper reveals, based on the Sunway SW26010-PRO compute units (manycore with built-in custom accelerators) with custom memory configuration and a homegrown network fabric.
[3]
Although the exascale achievement results using the supercomputing standard "Top 500" benchmark were verified, though [4]not published , it is important to note that this "brain-scale" workload is not itself running at full exascale capability. Generally in supercomputing performance measurements, the standard is 64-bit floating point (FP64) but this work was based on mixed precision. The new Sunway system can handle FP64, FP16, and BF16 and can trade those around during training for maximum efficiency.
[5]
[6]
Even though mixed precision fails to make this a true sustained exascale workload in traditional terms, it does show evidence of some impressive hardware/software co-design thinking, especially as the supercomputing world wraps its collective head around how AI/ML is supposed to [7]integrate with "old school" modeling and simulation.
The Chinese team provides detailed chip and node-level details for tuning HPC systems for AI, including scheduling, memory, and I/O operation optimizations and a unique parallelization strategy that mixes parallel models and then cuts down on compute time and memory use. They also developed a distinct load balancer and strategy for using mixed precision efficiently.
[8]Biden administration effectively slaps bans on seven Chinese supercomputer companies for military links
[9]US Air Force boots up not one but two AMD-powered supercomputers after five years of Intel Haswell CPUs
[10]China prototypes pre-exascale super trio with its own non-US chips
[11]Uncle Sam bungs rich tech giants quarter of a billion bucks for exascale super R&D
"This is an unprecedented demonstration of algorithm and system co-design on the convergence of AI and HPC," the paper's authors say.
The model and optimization set, called BaGuaLu, "enables decent performance and scalability on extremely large models by combining hardware-specific optimizations, hybrid parallel strategies, and mixed precision training," the team adds.
[12]
The authors, which include Alibaba employees in addition to academics from major Chinese universities, add that with current capabilities, a 174-trillion parameter model train is within the realm of possibility.
For avid readers of the architecture-centric [13]The Next Platform , you can be sure there is a deep dive into the chewy bits of the architecture later today. Information about the machine's architecture has been light but there is finally some detail to sink our teeth into. ®
Get our [14]Tech Resources
[1] https://www.nextplatform.com/2021/10/26/china-has-already-reached-exascale-on-two-separate-systems/
[2] https://keg.cs.tsinghua.edu.cn/jietang/publications/PPOPP22-Ma%20et%20al.-BaGuaLu%20Targeting%20Brain%20Scale%20Pretrained%20Models%20w.pdf
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YiuAN0MDqW5eJAruFN8P5gAAAEQ&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://www.nextplatform.com/2021/11/15/why-did-china-keep-its-exascale-supercomputers-quiet/
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YiuAN0MDqW5eJAruFN8P5gAAAEQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YiuAN0MDqW5eJAruFN8P5gAAAEQ&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[7] https://www.theregister.com/2021/12/09/reg_debate_agnostic_against_2_thursday/
[8] https://www.theregister.com/2021/04/09/biden_china_supercomputers/
[9] https://www.theregister.com/2021/02/10/us_airforce_supercomputer_hpe_amd/
[10] https://www.theregister.com/2016/07/12/china_pre_exascale_prototypes/
[11] https://www.theregister.com/2017/06/15/us_exascale_funding/
[12] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/hpc&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YiuAN0MDqW5eJAruFN8P5gAAAEQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[13] https://www.nextplatform.com/
[14] https://whitepapers.theregister.com/
So, American Trade wars work?
The US started a trade war that seems to have given the Chinese the push they needed to become self sufficient. Trumpists must be so proud of their achievement.
96000 nodes... That must result in an impressive electricity bill.
With 37+ MegaCores and lets say 25W per core (including memory and fabric, etc.) it would make a 925 MW installation. Wow... almost a GigaWatt computer.
No wonder they keep building power generation capacity. You need one per super computer. And all that for running a language model. A normal brain is about 25W too. They could just employ 37 million people to run the model and get better language response.