
What happened
On the same Xeon 6980P silicon and socket count as MLPerf v6.0, Intel reported a 2.4x rise in Llama 3.1 8B Server throughput and 56% higher Offline throughput in v6.1, achieved through software alone.
Why it matters
Intel says optimizations can stretch the useful life of servers customers already run, letting them handle larger, more varied AI models without buying new hardware.
What to watch
The gains apply to specific submitted configurations, so real-world results will hinge on each customer's setup and software stack. Partner submissions rose from 29 to 39, with Oracle, Red Hat, Quanta Cloud Technology and Supermicro joining.
WHO IT HITSEnterprise infrastructure teams running Xeon-based servers can potentially defer hardware refreshes if they adopt the software updates. Buyers evaluating AI inference platforms may find Intel's growing partner validation useful, though results depend on their own configurations.
Summaries like this, in your inbox every morning.
MLPerf Inference is a widely watched set of benchmarks run by MLCommons that lets companies compare how fast different hardware handles AI tasks. In the v6.0 round, Intel's submissions focused on a narrower set of Xeon and Arc Pro configurations. For v6.1, Intel expanded its Xeon participation from two benchmarked SKUs to five, increasing CPU inference results from 24 to 35, and its Arc Pro B70 submissions now cover more models including Llama 2 70B and gpt-oss-120B.
The company also pointed to broader ecosystem involvement, with Oracle making its first Intel-based submission and Red Hat delivering its first Xeon CPU inference submission. Quanta Cloud Technology and Supermicro provided the first partner submissions using Intel Arc Pro B70 GPUs. Intel says its Xeon improvements are being upstreamed into widely used AI frameworks so customers can benefit on deployed infrastructure, and its Arc Pro B70 optimization work is intended to advance software for future Intel GPU products.
The stakes for Intel's argument rest on whether these software gains translate into real cost savings for businesses running AI inference. That likely depends on how closely a customer's setup matches the benchmarked configurations and whether they adopt the updated software. If the improvements hold in typical enterprise environments, they could help delay hardware replacement cycles, a meaningful consideration for IT teams managing tight budgets.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Alphabet is Berkshire's growth pick — nearly 106 million shares bought across its two tickers, lifting its por…

Dell Technologies and Axelera AI announced a collaboration on the new Europa AI Processing Unit architecture f…

Dell took almost $61 billion in AI server orders last quarter and ended with a $95 billion backlog

Denver-based Crusoe survived the crypto winter of 2022 and 2023, which sent many bitcoin miners to the brink o…

Bank of America CEO Brian Moynihan told investors that AI has helped the bank avoid layoffs

In a post on X and other social platforms, Meta CEO Mark Zuckerberg said labs that fail to "focus on alignment…
