
Hugging Face released 207 optimized WebGPU kernels for faster AI in browsers. The kernels are 2.57x faster than ORT WebGPU on average in tests.
A tool called Fleet lets users benchmark the kernels on their own hardware.
The collection is Apache-2.0 licensed and available now.
What happened
Hugging Face released @huggingface/kernels, a library for running optimized WebGPU kernels from the Hugging Face Hub, along with an initial collection of 207 kernels. The kernels are published as individual, versioned packages with correctness tests and benchmark cases, and the release also includes Fleet, an in-browser benchmarking tool.
Why it matters
The kernels are designed to make AI inference (when a model produces an answer) in a web browser faster and easier. Compared to ORT WebGPU on an Apple M4 GPU, the collection was 2.57x faster by geometric mean and 1.90x faster at the median across 809 comparable test cases, with some individual operations running over 10,000x faster. This foundational layer allows higher-level runtimes to dispatch more efficient GPU operations.
What to watch
The kernels are Apache-2.0 licensed and available now. The package can be installed with `npm install @huggingface/kernels@preview` and requires a browser with WebGPU support. Hugging Face plans to connect these kernels to higher-level model tooling and expand operation coverage.
Ask the AI about this article →
Hugging Face's release targets a bottleneck in running AI models in the browser. While WebGPU offers a portable API for GPU operations, performance can vary dramatically depending on the device and how an operation is implemented. By making these low-level operations (kernels) individually testable and versioned, they create a stable foundation for higher-level runtimes to build upon.
The initial benchmark results against ONNX Runtime Web show a significant performance advantage in many cases. The team emphasizes that these figures are for individual operations, not complete models, and that exact performance will vary across different hardware. This is precisely why they launched Fleet, to gather a broader picture of real-world performance through crowdsourced testing.
Looking ahead, the strategy is to build a shared ecosystem. The kernels are published on the Hub alongside those for other platforms like CUDA and ROCm, and the team is working with the ONNX Runtime team to upstream these improvements. This suggests the goal is not just to create a separate product, but to improve the overall infrastructure for browser-based AI, making it faster and more reliable for all developers.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Aranya Inc., a startup founded last year, launched today with $11 million in funding
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Phonely Ltd. launched Alma, a large language AI model built for voice agents and trained on over 10 million re…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Sarah O’Connor's book 'We Are Not Machines' explores how mechanization and AI have transformed the workforce…
