Mysterious 16-Tile & 32-Tile AMX Implementations Talked Up For More Performance
([Hardware] 5 Hours Ago
Advanced Matrix Extensions)
- Reference: 0001648576
- News link: https://www.phoronix.com/news/16-Tile-32-Tile-AMX-Performance
- Source link:
Last October on Phoronix we were first to point out there being some [1]mysterious intrigue around a new x86 implementation from a "corporate entity other than Intel/AMD" with the developer involved not elaborating on the matter. Now that same developer has posted another semi-cryptic message to the Linux kernel mailing list to talk about 16-tile and 32-tile Advanced Matrix Extensions (AMX) implementations. To date Intel only has 8-tile AMX.
Christian Ludloff has in the past has worked for Google, AMD, TI, and others as an x86 architecture expert and maintainer of the Sandpile.org x86 CPU information site. He was the one that relayed the message last year about an x86 "corporate entity other than Intel/AMD" and today his irregular activity on the Linux kernel mailing list is around larger tile implementations of AMX.
Intel Xeon CPU implementations with Advanced Matrix Extensions (AMX) to date have only featured eight tiles. But now Ludloff is talking about CPUs with 16 tile and 32 tile configurations in the name of increased performance. This is also ahead of AMX's evolution into the next-generation [2]AI Compute Extensions "ACE" for both future AMD and Intel CPUs.
"If x86 AMX/ACE with >8 tiles does not concern you, then you can stop reading.
----------------------- 8< -----------------------
I have been asked to document x86 AMX/ACE behavior of 16-tile and 32-tile implementations.
Background
----------
The goal of a 16-tile or 32-tile HW implementation is to achieve extra performance by running SW that takes advantage of the extra tiles. In particular, said SW is expected to mostly run compute kernels, minimize process switches, and avoid tile usage in interrupt handlers for storage and networking. The use of VMs is possible, but again, rapid switching amongst guests is not a target scenario.
...
An implementation with support for 16 and 32 tiles has been running at large scale for some time."
The [3]mailing list post goes on to document 16-tile and 32-tile implementations of AMX within the confines of the specification. It also goes on to point out various oversights in the AMX and APX and ACE specifications.
Still no word on who has been asking Ludloff to relay these technical details to the public mailing lists.
[1] https://www.phoronix.com/news/x86-Opcodes-Not-AMD-Or-Intel
[2] https://www.phoronix.com/search/AI+Compute+Extensions
[3] https://lore.kernel.org/lkml/CAKSQd8WxM75DeZvovYXkt2c60ftHeEW_gpf0qTxaLF8i8kjm=w@mail.gmail.com/
Christian Ludloff has in the past has worked for Google, AMD, TI, and others as an x86 architecture expert and maintainer of the Sandpile.org x86 CPU information site. He was the one that relayed the message last year about an x86 "corporate entity other than Intel/AMD" and today his irregular activity on the Linux kernel mailing list is around larger tile implementations of AMX.
Intel Xeon CPU implementations with Advanced Matrix Extensions (AMX) to date have only featured eight tiles. But now Ludloff is talking about CPUs with 16 tile and 32 tile configurations in the name of increased performance. This is also ahead of AMX's evolution into the next-generation [2]AI Compute Extensions "ACE" for both future AMD and Intel CPUs.
"If x86 AMX/ACE with >8 tiles does not concern you, then you can stop reading.
----------------------- 8< -----------------------
I have been asked to document x86 AMX/ACE behavior of 16-tile and 32-tile implementations.
Background
----------
The goal of a 16-tile or 32-tile HW implementation is to achieve extra performance by running SW that takes advantage of the extra tiles. In particular, said SW is expected to mostly run compute kernels, minimize process switches, and avoid tile usage in interrupt handlers for storage and networking. The use of VMs is possible, but again, rapid switching amongst guests is not a target scenario.
...
An implementation with support for 16 and 32 tiles has been running at large scale for some time."
The [3]mailing list post goes on to document 16-tile and 32-tile implementations of AMX within the confines of the specification. It also goes on to point out various oversights in the AMX and APX and ACE specifications.
Still no word on who has been asking Ludloff to relay these technical details to the public mailing lists.
[1] https://www.phoronix.com/news/x86-Opcodes-Not-AMD-Or-Intel
[2] https://www.phoronix.com/search/AI+Compute+Extensions
[3] https://lore.kernel.org/lkml/CAKSQd8WxM75DeZvovYXkt2c60ftHeEW_gpf0qTxaLF8i8kjm=w@mail.gmail.com/