
Baseten, an AI infrastructure platform, is now available as a supported Inference Provider on Hugging Face Hub, allowing developers to easily run models like DeepSeek V4 Flash and Kimi K3 through the Hub's interface and SDKs.
Users can either use their own Baseten API key (billed directly to their Baseten account) or route requests through Hugging Face (charged at standard provider rates with no additional markup), and Hugging Face PRO members receive $2 in monthly Inference credits that work across all providers.
What happened
Baseten, an AI infrastructure platform, is now integrated into Hugging Face Hub as a supported Inference Provider. Users can access Baseten-hosted models directly through the Hub's website, Python and JavaScript SDKs, and integrated agent tools. Initial support covers conversational and text-generation tasks, with models including DeepSeek V4 Flash, Kimi K3, and GLM-5.2.
Why it matters
Developers can now run a broader range of open-weight LLMs and other AI models through Hugging Face without setting up separate infrastructure. Users can choose between routing requests through Hugging Face (billed to their HF account with no added markup) or using their own Baseten API key. Hugging Face PRO subscribers receive $2 in monthly Inference credits usable across providers.
What to watch
Support for additional task types beyond text generation will roll out soon. The full list of Baseten-supported models is available at https://huggingface.co/baseten, and users can customize provider preferences in their account settings.
Hugging Face has announced that Baseten, an AI infrastructure platform covering serverless AI, training, and related services, is now available as a supported Inference Provider on the Hugging Face Hub. The integration enables developers to run a broad spectrum of models—from LLMs to text-to-speech—directly through Hugging Face's platform.
Baseten's initial integration on the Hub focuses on conversational and text-generation tasks, with support for popular open-weight models including DeepSeek V4 Flash, Kimi K3, and GLM-5.2. Users can access these models through multiple pathways: the Hugging Face website UI (where compatible providers appear on model pages and can be reordered by user preference), the Python SDK (huggingface_hub >= 1.26.1), the JavaScript SDK (@huggingface/inference), and integrated agent harnesses such as Pi, OpenCode, Hermes Agents, and OpenClaw. Code examples provided in the body show developers can authenticate using their Hugging Face token and automatically route requests to Baseten—for instance, calling DeepSeek V4 Flash via the model identifier "deepseek-ai/DeepSeek-V4-Flash-0731:baseten".
Hugging Face offers two billing modes for Baseten. In the "custom key" mode, developers use their own Baseten API key and are billed directly by Baseten. In the "routed by HF" mode, users authenticate via their Hugging Face account and pay the standard Baseten provider API rates with no additional markup from Hugging Face; future revenue-sharing agreements with providers are possible. Hugging Face PRO subscribers receive $2 worth of Inference credits each month that work across all providers, while free users have access to a smaller inference quota. The body also notes that support for additional task types beyond text generation will roll out soon, and a full list of Baseten-supported models is available at https://huggingface.co/baseten.
Hugging Face Hub has expanded its ecosystem of Inference Providers to include Baseten, an AI infrastructure platform offering serverless AI and training services. This integration allows developers to access a wider range of models—particularly open-weight LLMs—without the friction of setting up separate accounts and infrastructure. The body describes two pathways for using Baseten through Hugging Face: direct billing (using a personal Baseten API key) or routed billing (through the Hugging Face account with no added markup). This dual approach removes barriers for both casual users and those with existing provider relationships.
The integration spans multiple entry points: the website UI (where providers appear on model pages ranked by user preference), client SDKs in Python and JavaScript, and popular agent harnesses including Pi, OpenCode, Hermes Agents, and OpenClaw. This breadth of integration points suggests that Hugging Face is positioning itself as a unified interface where developers can compare and switch between multiple inference providers without rewriting application code. The mention that Baseten will add support for additional task types (beyond the initial conversational and text-generation tasks) indicates this partnership is intended to deepen over time.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
NVIDIA and partners released multiple open-source AI models optimized for local execution throughout August, i…

Gregory Kurtzer, founder of Rocky Linux and co-founder of CentOS, released OpenWALDO, an open-source project d…

Meta released a new open-weight AI model this week—one with publicly accessible core components that users can…

Nvidia released Nemotron 3.5 Lightning, a compact open-weights model with 31.6 billion total parameters (only…

A portable, open-source tool called Se-harness (seh) generates unified instruction files (AGENTS.md) that work…

OpenAI announced ChatGPT Business Premium seats on Monday, priced at $125/month (or $100/month if billed annua…

The AI news that matters, in one minute each morning.
Sign up free