
Intel's LLM Scaler software stack has been updated to support advanced features like Multi-token Prediction and LoRA fine-tuning on its Arc Pro B60 and B70 GPUs, expanding the range of generative AI workloads that can run on Intel hardware.
The platform now covers 60+ models across text, vision, audio, and video generation, with recent releases (August 2026) also adding per-block quantization support to improve efficiency.
What happened
Intel released llm-scaler-vllm:0.21.0-b2 in August 2026, adding Multi-token Prediction (MTP) and LoRA Serving support for Qwen3.6-27B, Qwen3.6-35B-A3B, gemma-4-31B-it, and gemma-4-26B-A4B-it models, as well as support for per-block quantized models Qwen3.6-27B-FP8 and Qwen3.6-35B-A3B-FP8.
Why it matters
LLM Scaler enables developers and enterprises to run state-of-the-art generative AI models—for text, image, and video—natively on Intel Arc Pro B60 and B70 GPUs using standard frameworks like vLLM and ComfyUI, rather than relying solely on NVIDIA or other vendors' hardware.
What to watch
The platform now supports 60+ models across text generation (including DeepSeek, Qwen, Llama, Mistral, Gemma), multi-modal vision models, embedding and reranking models, and omni models for image, video, and audio generation. Omni Serving provides OpenAI-API-compatible endpoints for image and audio tasks.
Ask the AI about this article →
LLM Scaler represents Intel's strategy to enable its Arc Pro GPUs as viable alternatives for generative AI workloads. Rather than building proprietary software, the platform leverages open-source frameworks (vLLM, ComfyUI, SGLang Diffusion, Xinference) to ensure compatibility and ease of adoption for developers already familiar with these tools. The August 2026 release adds Multi-token Prediction and LoRA Serving—two features that improve both inference speed and fine-tuning capabilities—extending the platform's utility for both deployment and model customization.
The breadth of supported models has grown substantially: the platform now covers major open-weight models from OpenAI (GPT-OSS), Qwen, DeepSeek, Llama, Mistral, Gemma, and others, spanning text, vision, embedding, and omni (multimodal) architectures. Quantization options (INT4, FP8, MXFP4) allow operators to trade off precision for memory and speed, a critical capability for running larger models on GPU memory-constrained systems. The Omni Serving mode, which exposes image and audio generation via OpenAI-compatible APIs, lowers integration friction for applications already built on that interface.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Israeli startup DataAgent Ltd
SK Hynix presented a custom HBM concept at SEMICON Taiwan 2026, where compute functions are placed in the base…

Nvidia reported earnings that were both remarkable and boring, reflecting its focus on avoiding a consolidated…

Anthropic has agreed to a $35bn cloud-computing contract with Lambda, a Nvidia-backed cloud provider

The Supreme Court of Japan has included about ¥60 million in its fiscal 2027 budget request for AI-related exp…

The Consumer Affairs Agency said Tuesday it will use generative AI to analyze about 900,000 annual consultatio…
