KTransformers 0.7 Expands AVX-512 Support To Benefit AMD EPYC Servers
([AI] 5 Hours Ago
KTransformers 0.7)
- Reference: 0001652335
- News link: https://www.phoronix.com/news/KTransformers-0.7
- Source link:
KTransformers as the framework for heterogeneous LLM inference and fine-tune optimizations is out today with its v0.7 feature release.
With KTransformers 0.7 there is now full [1]AVX-512 support for LoRA fine-tuning without depending upon Advanced Matrix Extensions ( [2]AMX ) also being present. This AVX-512-only without AMX benefits AMD EPYC Zen 4 / Zen 5 / Zen 6 servers with excellent AVX-512 support while lacking AMX and also older Intel Xeon processors with AVX-512 prior to the introduction of AMX with Sapphire Rapids.
The [3]merge request noted the testing on AMD hardware and the foxus on AVX-512 without AMX platforms. The KTransformers runtime will automatically select the proper CPU implementation and in turn allowing MoE expert training to happen on a wider range of large-memory servers.
KTransformers 0.7 also adds VLM fine-tuning support, native FP8 LoRA support, improved DeepSeek V4 deployment, and CPU activation reuse.
More details on KTransformers 0.7 for those using it for LLM inference optimizations and fine-tuning can find all the details via the release announcement on [4]GitHub .
[1] https://www.phoronix.com/search/AVX-512
[2] https://www.phoronix.com/search/AMX
[3] https://github.com/kvcache-ai/ktransformers/pull/2141
[4] https://github.com/kvcache-ai/ktransformers/releases/tag/v0.7.0
With KTransformers 0.7 there is now full [1]AVX-512 support for LoRA fine-tuning without depending upon Advanced Matrix Extensions ( [2]AMX ) also being present. This AVX-512-only without AMX benefits AMD EPYC Zen 4 / Zen 5 / Zen 6 servers with excellent AVX-512 support while lacking AMX and also older Intel Xeon processors with AVX-512 prior to the introduction of AMX with Sapphire Rapids.
The [3]merge request noted the testing on AMD hardware and the foxus on AVX-512 without AMX platforms. The KTransformers runtime will automatically select the proper CPU implementation and in turn allowing MoE expert training to happen on a wider range of large-memory servers.
KTransformers 0.7 also adds VLM fine-tuning support, native FP8 LoRA support, improved DeepSeek V4 deployment, and CPU activation reuse.
More details on KTransformers 0.7 for those using it for LLM inference optimizations and fine-tuning can find all the details via the release announcement on [4]GitHub .
[1] https://www.phoronix.com/search/AVX-512
[2] https://www.phoronix.com/search/AMX
[3] https://github.com/kvcache-ai/ktransformers/pull/2141
[4] https://github.com/kvcache-ai/ktransformers/releases/tag/v0.7.0