How Apple's M1 uses high-bandwidth memory to run like the clappers
- Reference: 1605789012
- News link: https://www.theregister.co.uk/2020/11/19/apple_m1_high_bandwidth_memory_performance/
- Source link:
The company claims its [1]M1 Arm chip delivers up to 3.5x faster CPU performance, up to 6x faster GPU performance, up to 15x faster machine learning, and up to 2x longer battery life than previous-generation Macs, which use Intel x86 CPUs.
Let's take a closer look at how Apple uses high-bandwidth memory in the M1 system-on-chip (SoC) to deliver this rocket boost.
High-bandwidth memory (HBM) avoids the traditional CPU socket-memory channel design by pooling memory connected to a processor via an interposer layer. HBM combines memory chips and gives them closer and faster access to the CPU as the distance to the processor is only a few micrometer units. This on its own speeds data transfers.
The M1, Apple's first Mac SoC, is built by chip foundry TSMC using 16 billion transistors with 5nm technology. It includes an eight-core CPU, an eight-core GPU, a 16-core neural engine, storage controller, image signal processor, and media code/decode engines.
This Apple diagram of the M1 SoC shows two blocks of DRAM:
Apple M1 unified memory architecture
The SoC has access to 16GB of unified memory. This uses 4266 MT/s LPDDR4X SDRAM (synchronous DRAM) and is mounted with the SoC using a system-in-package (SiP) design. A SoC is built from a single semiconductor die whereas a SiP connects two or more semiconductor dies.
SDRAM operations are synchronised to the SoC processing clock speed. Apple describes the SDRAM as a single pool of high-bandwidth, low-latency memory, allowing apps to share data between the CPU, GPU, and Neural Engine efficiently.
In other words, this memory is shared between the three different compute engines and their cores. The three don't have their own individual memory resources, which would need data moved into them. This would happen when, for example, an app executing in the CPU needs graphics processing – meaning the GPU swings into action, using data in its memory.
The downside of this design is that expandability is traded for performance. Users cannot simply add more memory to the configuration; they cannot plug more memory DIMMs into carriers as there are no carriers and DIMM technology isn't used.
We can envisage a future in which all storage controllers, [2]SmartNICs , and DPUs could use Arm SoCs with a pool of unified memory to run their workloads much faster than traditional x86 controllers, which are hampered by memory sockets and DIMMs.
For instance, [3]Nebulon's Storage Processing Unit (SPU) uses dual Arm processors. Conceivably this could move to a unified memory design, giving Nebulon additional power to run its storage processing workload, and so exceed x86-powered storage controllers in performance, cost, and efficiency terms even more than it does now. ®
Get our [4]Tech Resources
[1] https://www.theregister.com/2020/11/10/apple_silicon_debut/
[2] https://blocksandfiles.com/2020/10/28/amd-xilinx-smartnic-data-centre/
[3] https://blocksandfiles.com/2020/10/27/nebulon-spu-costs-vs-san-and-hci/
[4] https://whitepapers.theregister.com/
I was going to say something similar. Conventional DRAM is still faster than NAND so it makes sense to start virtual memory there, and off load to disk if needed.
Isn't this what AMD are already doing with their "Infinity Cache" - currently only 128MB on their latest GPUs but you could easily see how that could be expanded (manufacturing tech permitting).
Various of their previous GPUs have featured HBM but I gather it was prohibitively expensive and now only appears on their datacenter/workstation products.
HBM is a DDR memory technology with the associated high latency so not suited for a L4 cache. But such an external memory controller could well be useful for Optane type memory holding the operating system's page file.
Bandwidth and latency for an HBM2e pseudo-channel are roughly comparable with DDR5-5600 (32-bit wide). So you can think of an HBM stack as having 8x the performance of a DDR5 DIMM (though it has 16 effectively independent pseudo-channels). Currently HBM devices are available in up to 8GB capacity, but like all things DRAM that will inevitable increase with time.
expandability? when has a mac ever had that?
2010 Mac Pro
Mac pro
My 2010 was a genuine “Trigger’s broom”. Third vid., second set of full RAM, better than Apple RAID controller driving 12TB across four drives in 0/1 configuration , hybrid system drive in second DVD slot.... etc.
Recently moved on up to a 2012 as I needed High Sierra support without using a VM.
You can install extra RAM in many Macs, sometimes the diskdrive too.
It's been a while
This Mac user would certainly appreciate a simple box you put additional things into (or replace existing things with better/working things) :-/.
Sounds like a sound plan to ensure those pesky customers don't avoid the Apple Tax by buying more memory later, even if their use case changes and demands it, or they simply realise that what they thought would be sufficient actually isn't.
Better over-spec at the start just in case $$$
Except it also has a substantial benefit from the engineering side. Did you even bother reading the article?
You haven't been able to do anything with the notebooks for more than five years now. It is annoying. However, apart from being able to run more and more VMs at once, RAM use on MacOS has been reasonably constant for the last 10 or so.
Who knows, maybe Apple will let people swap the SoC in a year or so from now. For a price, of course. But in the meantime there's no denying that they have a fairly compelling value proposition: improved performance and significantly improved battery life. That said, I certainly won't be switching to Big Sur until I know what the restrictions are and when they've really fixed all the bugs. I've skipped versions in the past when it was clear they were too buggy.
Glad to see Apple continuing their 'innovation' by copying what Silicon Graphics did in their workstations, what, 25 years ago?
Unified Memory
Isn't that just a new name for 'Shared Memory'?
I'm sure that a lot of readers of this esteemed site remember that.
You know when the Grapics Card used a great chunk of the CPU memory because the Graphics Card makers were cheapskates.
Things improved when graphics cards started to get really fast RAM.
At least Apple have made the memory bandwidth really, really large. IMHO, that's where a good deal of the performance comes from.
This CPU will give a few other chip designers a lot to think about.
But hey... Apple can't innovate can they? (sic)
Re: Unified Memory
The thing is, HBM stacks have 16 effectively independent data paths (each > 20 GB/s for HBM2e) so the system architect can assign some to the CPU and some to the GPU. It looks like that design has 2 stacks, so 32x 20GB/s should be enough to go round.
(By the way, yes there are applications that will saturate the bandwidth of 2 HBM stacks though you will find them running on GPGPUs for HPC).
Edit: by the way, something like this was done by Intel with the truly weird and wonderful Kaby Lake G - Intel CPU, AMD GPU, and HBM all in the same package. There is an important difference though: in that product the HBM was only attached to the GPU, and the CPU retained a traditional external memory interface.
Re: Unified Memory
Unified memory I believe creates the same addressing scheme across users unlike shared memory, where it can allocate from the pool, but the allocation cannot be moved, it can be copied and freed.
This is the difference between shared and unified memory. Unified is a super set of shared - they theoretically can see all the memory and a handle alone can move data across cores/users.
This gives a significant performance advantage for offload and heterogenous compute loads.
Re: Unified Memory
Hmm, that's a pretty convincing explanation of the unified vs shared mem difference. I've always struggled to understand the difference between OpenCL's USM and SVM, maybe this is the key...
If I get this right, the difference is that in a heterogenous shared memory scenario, the memory appears at different locations in each device's address space so pointers are not translatable. For example if the CPU builds a linked list in shared memory, the accelerator cannot just dereference the pointers.
So what is the neural engine for?
Apple are making a big deal about ML capabilities on the M1 chippery but what use does this have in a commercial laptop used for gaming/music production/video editing/general use?
Is it used for everyday purposes I'm simply not aware of? I can't imagine Apple would bother with it for no reason but other than possibly Siri, I'm struggling to think why this is a core part of the system.
Anyone got a good answer?
Re: So what is the neural engine for?
Photo touchup and editing algos uses inferencing a lot. Edge detection, contour, face detect come to mind. These can leverage the NPU, just like GPUs do for graphics.
Others are for applications that customise per user - such as your usage patterns. These can use inferencing as well.
I'd imagine games could use it for some aspects of game play. This is a tiny subset of examples.
Just like graphics/GPU, the CPU could do it, but it is far more efficient and faster with an NPU.
Re: So what is the neural engine for?
Siri speech recognition. The part where you say: "Hey Siri" is being processed locally on your phone, not in a datacenter like the rest of the "conversation."
Performance tricks
I believe that had the DRAM been stored off-chip the M1's performance numbers would've been a lot less flattering.
But chances AMD and Intel will stomp up something similar soon, before Q3 2021 is my guess.
Re: Performance tricks
"The SoC has access to 16GB of unified memory. This uses 4266 MT/s LPDDR4X SDRAM (synchronous DRAM) and is mounted with the SoC using a system-in-package (SiP) design. A SoC is built from a single semiconductor die whereas a SiP connects two or more semiconductor dies."
The DRAM is not on the same chip. Or die.
Sowing confusion
Apple really isn't helping by calling this [1]high‑bandwidth, low‑latency memory , because, despite it being LPDDR4X, people are likely to confuse it with [2]High Bandwidth Memory , a JEDEC standard. Indeed, a poor translation on Apple Finnish store ( [3]since corrected ) actually said "High Bandwidth Memory".
On the "unified memory is old hat" theme, yes indeed: there's a [4]now-expired Apple patent concerning it from 1996.
[1] https://www.apple.com/mac/m1
[2] https://www.jedec.org/document_search?search_api_views_fulltext=jesd235
[3] https://www.apple.com/fi/shop/buy-mac/mac-mini/applen-m1-siru,-jossa-on-8-ytiminen-prosessori-ja-8-ytiminen-näytönohjain-256gt#
[4] https://patents.google.com/patent/US6031964A/en
So now Apple had made it's entire RAM space shared
I seem to recall not so long ago that Intel had a big problem with its [1]CPU architecture that could allow programs to access kernel memory, and many people were all in tizzy about it.
Now, Apple brings a system-on-a-chip that shares all its memory space.
Is there no problem with that ?
[1] https://www.theregister.com/2018/01/02/intel_cpu_design_flaw/
Long term, I think we will see expansion options
With good memory management, perhaps we could expect little or no noticeable performance decrease where there is a mix of integrated and DIMM memory. That said, it is starting to look like what memory you have goes further on these new Macs (various demos of users maxing out the 8GB MacBooks; lots of vigorous debate around this).
The MacBook Air 2020 M1 sitting behind me (humble base model I've just bought)? Erm, runs like the clappers. Office under Rosetta? About 20 seconds when launching each application for the first time. Near as damnit instant thereafter (proving there is a pre-process step). So, clappers for Office also. I'm a big believer in the capitalist principle that competition is good. This is that massive (and much needed) boot up Intel's arse.
Re: Long term, I think we will see expansion options
This is that massive (and much needed) boot up Intel's arse.
Especially with AMD firmly taking the high end performance crown from them with Zen3, this is probably not a fun time for Intel.
As you say competition is good. Last time Intel were getting clobbered by AMD (Athlon 64 vs Pentium IV), Intel went back to the drawing board and came back with the Core architecture, so hopefully they'll take this opportunity to do the same.
Great Block Diagram
I like the block diagram that illustrates the construction of this processor. Very informative. Apparently there are boxes connected through a box called 'fabric'. I shall remember this technique when I do my next design.
Reading between the lines it looks like the key design points are that the memory is physically close to the processing elements which allows the memory to be synchronized with the processor clocks. There's also no mention of (L1) cache, suggesting that the memory is effectively the cache. So at the price of some trickery in the memory controller design (maybe just alternating access between main processors and GPUs?) you get rid of all the overhead of managing cache misses.
(Shared memory is as old as the hills and then some. I got lumbered with a university project in the mid-1970s that rehabilitated a GPU that was designed to share memory with the main processor. I had to make it standalone and provide it with an interface to a different system. Logically straightforward but very tedious because the technology used was prehistoric (it was "transistorized", though). The GPU was from an old English Electric computer, it was really very well designed for somthing that old.)
How long can it be before this approach gets extended with a memory controller for off-SIP memory so that the high bandwidth on-SIP memory is just another layer of cache? One year? Two? In the lab now?