Intel Updates LLM-Scaler-vLLM Build For vLLM 0.26 & Other Improvements
([Intel] 11 Minutes Ago
llm-scaler-vllm)
- Reference: 0001655975
- News link: https://www.phoronix.com/news/Intel-LLM-Scaler-vLLM-0.26.0-b1
- Source link:
For those looking for the smoothest experience to get vLLM up and running on Intel Arc (Pro) graphics hardware, out today is the newest beta of Intel's LLM-Scaler-vLLM project for a Docker-based setup of vLLM ready to go on Intel GPUs.
The new Intel LLM-Scaler-vLLM 0.26.0-b1 beta has upgraded against upstream vLLM 0.26. The upstream vLLM 0.26 release added DeepSeek-V4 kernel support, improved KV offloading, JIT warm-up infrastructure, and a variety of other performance optimization work.
In addition to re-basing against vLLM 0.26, the new LLM-Scaler-vLLM also improves the time-to-first-token for Qwen models. There is also improved FP8 KV cache performance as well as fixing various bugs.
Downloads and more details on the new LLM-Scaler-vLLM release via [1]GitHub .
[1] https://github.com/intel/llm-scaler/releases/tag/vllm-0.26.0-b1
The new Intel LLM-Scaler-vLLM 0.26.0-b1 beta has upgraded against upstream vLLM 0.26. The upstream vLLM 0.26 release added DeepSeek-V4 kernel support, improved KV offloading, JIT warm-up infrastructure, and a variety of other performance optimization work.
In addition to re-basing against vLLM 0.26, the new LLM-Scaler-vLLM also improves the time-to-first-token for Qwen models. There is also improved FP8 KV cache performance as well as fixing various bugs.
Downloads and more details on the new LLM-Scaler-vLLM release via [1]GitHub .
[1] https://github.com/intel/llm-scaler/releases/tag/vllm-0.26.0-b1