Nvidia just made a killing on AI – where is everyone else?
- Reference: 1693290914
- News link: https://www.theregister.co.uk/2023/08/29/nvidia_q2_ai_market/
- Source link:
Demand for the tech titan's GPUs drove its revenues to [1]new heights , as enterprises, cloud providers, and hyperscalers all scrambled to stay relevant under the new AI world order.
But while Nvidia's execs expect to extract multi-billion dollar gains from this demand over the next several quarters, the question on many minds is whether Nvidia and partners can actually build enough GPUs to satisfy demand, and what happens if they can't.
[2]
In a call with financial analysts, Nvidia CFO Colette Kress [3]assured Wall Street that the graphics processor giant was working closely with partners to reduce cycle times and add supply capacity. When pressed for specifics, Kress repeatedly dodged the question, arguing that Nv's gear involved so many suppliers it was difficult to say how much capacity they'd be able to bring to bear and when.
[4]
[5]
A Financial Times [6]report , meanwhile, suggested Nvidia plans to, at a minimum, triple the production of its top-spec H100 accelerator in 2024 to between 1.5 and 2 million units, up from roughly half a million this year. While this is great news for Nvidia's bottom line, if true, some companies aren't waiting around for Nvidia to catch up, and instead are looking to alternative architectures.
Unmet demand breeds opportunity
One of the most compelling examples is United Arab Emirates G42 Cloud, which [7]tapped Cerebras Systems to build nine AI supercomputers capable of a combined 36 exaflops of sparse FP16 performance for a mere $100 million a piece.
Cerebras's [8]accelerators are wildly different from the GPUs that power Nvidia's HGX and DGX systems. Rather than packing four or eight GPUs into a rack mount chassis, Cerebra's accelerators are enormous dinner-plate-sized sheets of silicon packing 850,000 cores and 40GB of SRAM. The chipmaker claims just 16 of these accelerators are required to achieve 1 exaflop of sparse FP16 performance, a feat that, by our estimate, would require north of 500 Nvidia H100s.
And for others willing to venture out beyond Nvidia's walled garden, there's no shortage of alternatives. Last we heard, Amazon is [9]using Intel's Gaudi AI training accelerators to supplement its own custom Trainium chips — though it isn't clear in what volumes.
[10]
Compared to Nvidia's A100, Intel's Gaudi2 processors, which [11]launched last May, claim to deliver roughly twice the performance, at least in the ResNet-50 image classification model and BERT natural language processing models. And for those in China, Intel recently [12]introduced a cut down version of the chip for sale in the region. Intel is expected to launch an even more powerful version of the processor, predictably called Gaudi3, to compete with Nvidia's current-gen H100 sometime next year.
Then of course, there's AMD, which, having enjoyed a recent string of high-profile wins in the supercomputing space, has turned its attention to the AI market.
At its Datacenter and AI event in June, AMD [13]detailed its Instinct MI300X, which is slated to start shipping by the end of the year. The accelerator packs 192GB of speed HBM3 memory and eight CDNA 3 GPU dies into a single package.
[14]
Our sister site The Next Platform [15]estimates the chip will deliver roughly 3 petaflops of FP8 performance. While 75 percent of a Nvidia's H100 in terms of performance, the MI300X's offers 2.4x higher memory capacity, which could allow customers to get away with using fewer GPUs to train their models.
The prospect of a GPU that can not only deliver compelling performance, but which you can actually buy , clearly has piqued some interest. During AMD's Q2 earnings call this month, CEO Lisa Su [16]boasted that the company's AI engagements had grown seven fold during the quarter. "In the datacenter alone, we expect the market for AI accelerators to reach over $150 billion by 2027," she said.
[17]AMD hopes for rebound in the second half of the year as Q2 profits plunge 94%
[18]Cerebras's Condor Galaxy AI supercomputer takes flight carrying 36 exaFLOPS
[19]Chinese web giants go on $5B Nvidia shopping spree to fuel AI ambitions
[20]Good thing Nvidia makes number-crunching GPUs – it'll need them to count its cash
Barriers to adoption
So if Nvidia thinks right now it's only addressing a third of demand for its AI-focused silicon, why aren't its rivals stepping up to fill in the gap and cash in on the hype?
The most obvious issue is that of timing. Neither AMD nor Intel will have accelerators capable of challenging Nvidia's H100, at least in terms of performance, ready for months. However, even after that, customers will still have to contend with less mature software.
Then there's the fact that Nvidia's rivals will be fighting for the same supplies and manufacturing capacity that Nv wants to secure or has already secured. For example, AMD [21]relies on TSMC just as Nvidia does for chip fabrication. Though semiconductor demand [22]is in a slump as fewer people are interested in snapping up PCs, phones, and the like lately, there is significant demand for server accelerators to train models and power machine learning applications.
But back to the code: Nvidia's close knit hardware and software ecosystem has been around for years. As a result, there's a lot of code, including many of the most popular AI models, optimized for Nv's industry-dominating CUDA framework.
That's not to say rival chip houses aren't trying to change this dynamic. Intel's OneAPI includes tools to help users to [23]convert code written for Nvidia's CUDA to SYCL, which can then run on Intel's suite of AI platforms. Similar efforts have been made to convert CUDA workloads to run on AMD's Instinct GPU family using the HIP API.
Many of these same chipmakers are also soliciting the help of companies like Hugging Face, which develops tools for building ML apps, to reduce the barrier to running popular models on their hardware. [24]These investments recently drove Hugging's valuation to over $4 billion.
Other chip outfits, like Cerebras, have looked to side step this particular issue by developing custom AI models for its hardware, which customers can leverage rather than having to start from scratch. Back in March, Cerebras [25]announced Cerebras-GPT, a collection of seven LLMs ranging from 111 million to 13 billion parameters in size.
For more technical customers with the resources to devote to developing, optimizing, or porting legacy code to newer, less mature architectures, opting for an alternative hardware platform may be worth the potential cost savings or reduced lead times. Both Google and Amazon have already gone down this route with their TPU and Trainium accelerators, respectively.
However, for those that lack these resources, embracing infrastructure without a proven software stack - no matter how performant it may be - could be seen as a liability. In that case, Nvidia will likely remain the safe bet. ®
Get our [26]Tech Resources
[1] https://www.nextplatform.com/2023/08/24/nvidia-theres-a-new-kid-in-datacenter-town/
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZO3Bw66409qwZ1iR@LUYLQAAAYQ&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://www.theregister.com/2023/08/24/nvidia_fy2024_q2_profits/
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZO3Bw66409qwZ1iR@LUYLQAAAYQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZO3Bw66409qwZ1iR@LUYLQAAAYQ&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.ft.com/content/c7e9cfa9-3f68-47d3-92fc-7cf85bcb73b3
[7] https://www.theregister.com/2023/07/20/cerebras_condor_galaxy_supercomputer/
[8] https://www.nextplatform.com/2022/11/16/cerebras-wants-its-piece-of-an-increasingly-heterogenous-hpc-world/
[9] https://aws.amazon.com/ec2/instance-types/dl1/
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZO3Bw66409qwZ1iR@LUYLQAAAYQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[11] https://www.theregister.com/2022/05/10/intels_habana_unit_reveals_new/
[12] https://www.theregister.com/2023/07/13/intel_guadi_china/
[13] https://www.theregister.com/2023/06/13/amd_epyc_announcement/
[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZO3Bw66409qwZ1iR@LUYLQAAAYQ&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[15] https://www.nextplatform.com/2023/06/14/the-third-time-charm-of-amds-instinct-gpu/
[16] https://www.theregister.com/2023/08/02/amd_q2_results/
[17] https://www.theregister.com/2023/08/02/amd_q2_results/
[18] https://www.theregister.com/2023/07/20/cerebras_condor_galaxy_supercomputer/
[19] https://www.theregister.com/2023/08/11/chinese_web_giants_nvidia/
[20] https://www.theregister.com/2023/08/24/nvidia_fy2024_q2_profits/
[21] https://www.theregister.com/2023/07/24/amd_looks_beyond_tsmc/
[22] https://www.theregister.com/2023/04/26/semiconductor_slump_worse_than_feared/
[23] https://www.nextplatform.com/2022/05/20/intel-takes-the-sycl-to-nvidias-cuda-with-migration-tool/
[24] https://www.theregister.com/2023/08/24/hugging_face_investment/
[25] https://www.cerebras.net/blog/cerebras-gpt-a-family-of-open-compute-efficient-large-language-models/
[26] https://whitepapers.theregister.com/
I suspect that another reason not all chipmakers are scrambling to make big investments in "AI" silicon capacity is that it's not yet proven that "AI" tools are all that valuable to end users.
So far, I've seen a whole lot of really fun toys. I can get neat ideas for my D&D campaign from ChatGPT, and I can remove objects I don't like from photos, and I can get nice drawings from descriptions (sort of, with lots of limitations, and the copyright implications are problematic). And lots more fun toys.
But where are the real world-killer applications? The ones that could remake the world, move hundreds of gigabucks, by automating away entire classes of jobs? AI customer support chatbots are nearly useless. Self-driving cars are not happening. AI-powered search is comically unreliable. Automatically generated news articles, legal documents, summaries, medical documents, whatever, they all must be carefully double-checked by a human that's qualified in the field. And we don't know how to fix these problems; we're just hoping really hard that they'll go away with a bigger model.
On top of that, it's not yet settled that scraping the Internet for training data is even legal, from a copyright perspective.
Do you really want to invest ten billions to build a new chip fab that will be operative five years from now, when it could well be that in a couple years the hype will die down and all that's left will be a few cool photoshop plugins and some games?
The end of the call centre
That's what they are for. Finally BigCorps can stop wasting money on frail flesh to follow on-screen scripts. You will never speak to a real person again on a support line.
AMD's lackluster GPGPUs
AMD has a horrible track record with its GPGPU products, MxGPU (GPUs with SR-IOV support) went nowhere, while the few existing products are poorly supported by drivers and rarely work as expected. And its Instinct GPGPU accelerators also suffer from poor driver support, and worse of all, for older models (e.g., Mi25) the drivers aren't even available from AMD itself.
This is pretty much the way AMD has been caring for the GPGPU market in the past. And there's no sign that things are going to improve.
It's no surprise Nvidia is the top dog amongst GPGPU suppliers.