AIToday
Open-Source AIHugging Face BlogPublished: Aug 7, 2026, 01:00 JST4 min read

Baseten joins Hugging Face Hub as supported inference provider

Baseten joins Hugging Face Hub as supported inference provider

Key takeaway

  • Baseten, an AI infrastructure platform, is now available as a supported Inference Provider on Hugging Face Hub, allowing developers to easily run models like DeepSeek V4 Flash and Kimi K3 through the Hub's interface and SDKs.

  • Users can either use their own Baseten API key (billed directly to their Baseten account) or route requests through Hugging Face (charged at standard provider rates with no additional markup), and Hugging Face PRO members receive $2 in monthly Inference credits that work across all providers.

3 Key Points

  1. What happened

    Baseten, an AI infrastructure platform, is now integrated into Hugging Face Hub as a supported Inference Provider. Users can access Baseten-hosted models directly through the Hub's website, Python and JavaScript SDKs, and integrated agent tools. Initial support covers conversational and text-generation tasks, with models including DeepSeek V4 Flash, Kimi K3, and GLM-5.2.

  2. Why it matters

    Developers can now run a broader range of open-weight LLMs and other AI models through Hugging Face without setting up separate infrastructure. Users can choose between routing requests through Hugging Face (billed to their HF account with no added markup) or using their own Baseten API key. Hugging Face PRO subscribers receive $2 in monthly Inference credits usable across providers.

  3. What to watch

    Support for additional task types beyond text generation will roll out soon. The full list of Baseten-supported models is available at https://huggingface.co/baseten, and users can customize provider preferences in their account settings.

In Depth

Read the full story

Hugging Face has announced that Baseten, an AI infrastructure platform covering serverless AI, training, and related services, is now available as a supported Inference Provider on the Hugging Face Hub. The integration enables developers to run a broad spectrum of models—from LLMs to text-to-speech—directly through Hugging Face's platform.

Baseten's initial integration on the Hub focuses on conversational and text-generation tasks, with support for popular open-weight models including DeepSeek V4 Flash, Kimi K3, and GLM-5.2. Users can access these models through multiple pathways: the Hugging Face website UI (where compatible providers appear on model pages and can be reordered by user preference), the Python SDK (huggingface_hub >= 1.26.1), the JavaScript SDK (@huggingface/inference), and integrated agent harnesses such as Pi, OpenCode, Hermes Agents, and OpenClaw. Code examples provided in the body show developers can authenticate using their Hugging Face token and automatically route requests to Baseten—for instance, calling DeepSeek V4 Flash via the model identifier "deepseek-ai/DeepSeek-V4-Flash-0731:baseten".

Hugging Face offers two billing modes for Baseten. In the "custom key" mode, developers use their own Baseten API key and are billed directly by Baseten. In the "routed by HF" mode, users authenticate via their Hugging Face account and pay the standard Baseten provider API rates with no additional markup from Hugging Face; future revenue-sharing agreements with providers are possible. Hugging Face PRO subscribers receive $2 worth of Inference credits each month that work across all providers, while free users have access to a smaller inference quota. The body also notes that support for additional task types beyond text generation will roll out soon, and a full list of Baseten-supported models is available at https://huggingface.co/baseten.

Context & Analysis

Hugging Face Hub has expanded its ecosystem of Inference Providers to include Baseten, an AI infrastructure platform offering serverless AI and training services. This integration allows developers to access a wider range of models—particularly open-weight LLMs—without the friction of setting up separate accounts and infrastructure. The body describes two pathways for using Baseten through Hugging Face: direct billing (using a personal Baseten API key) or routed billing (through the Hugging Face account with no added markup). This dual approach removes barriers for both casual users and those with existing provider relationships.

The integration spans multiple entry points: the website UI (where providers appear on model pages ranked by user preference), client SDKs in Python and JavaScript, and popular agent harnesses including Pi, OpenCode, Hermes Agents, and OpenClaw. This breadth of integration points suggests that Hugging Face is positioning itself as a unified interface where developers can compare and switch between multiple inference providers without rewriting application code. The mention that Baseten will add support for additional task types (beyond the initial conversational and text-generation tasks) indicates this partnership is intended to deepen over time.

FAQ

How do I get started with Baseten on Hugging Face?
You can access Baseten models through the Hugging Face website UI (model pages will show Baseten as a compatible provider), the Python SDK (huggingface_hub >= 1.26.1), the JavaScript SDK (@huggingface/inference), or integrated agent harnesses like Pi, OpenCode, and Hermes Agents. You authenticate using your Hugging Face token, and requests are routed to Baseten automatically.
How much does it cost to use Baseten through Hugging Face?
For routed requests (when you authenticate via Hugging Face Hub), you pay the standard Baseten provider API rates with no additional markup from Hugging Face. For direct requests using your own Baseten API key, you are billed on your Baseten account. Hugging Face PRO users receive $2 worth of Inference credits every month usable across providers.
What models does Baseten support on Hugging Face right now?
Initial support covers conversational and text-generation tasks, with popular open-weight LLMs including DeepSeek V4 Flash, Kimi K3, and GLM-5.2. The full list of supported models is available at https://huggingface.co/baseten, and additional task types will roll out soon.
Hugging Face BlogRead Original Article

Get the latest Open-Source AI news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleApple researchers propose method to lock open-weight models against unauthorized fine-tuning

The AI news that matters, in one minute each morning.

Sign up free