AIToday
Open-Source AIHugging Face BlogPublished: Aug 7, 2026, 01:00 JST

Baseten joins Hugging Face Hub as supported inference provider

Baseten joins Hugging Face Hub as supported inference provider

3 Key Points

  1. What happened

    Baseten, an AI infrastructure platform, is now integrated into Hugging Face Hub as a supported Inference Provider. Users can access Baseten-hosted models directly through the Hub's website, Python and JavaScript SDKs, and integrated agent tools. Initial support covers conversational and text-generation tasks, with models including DeepSeek V4 Flash, Kimi K3, and GLM-5.2.

  2. Why it matters

    Developers can now run a broader range of open-weight LLMs and other AI models through Hugging Face without setting up separate infrastructure. Users can choose between routing requests through Hugging Face (billed to their HF account with no added markup) or using their own Baseten API key. Hugging Face PRO subscribers receive $2 in monthly Inference credits usable across providers.

  3. What to watch

    Support for additional task types beyond text generation will roll out soon. The full list of Baseten-supported models is available at https://huggingface.co/baseten, and users can customize provider preferences in their account settings.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Hugging Face Hub has expanded its ecosystem of Inference Providers to include Baseten, an AI infrastructure platform offering serverless AI and training services. This integration allows developers to access a wider range of models—particularly open-weight LLMs—without the friction of setting up separate accounts and infrastructure. The body describes two pathways for using Baseten through Hugging Face: direct billing (using a personal Baseten API key) or routed billing (through the Hugging Face account with no added markup). This dual approach removes barriers for both casual users and those with existing provider relationships.

The integration spans multiple entry points: the website UI (where providers appear on model pages ranked by user preference), client SDKs in Python and JavaScript, and popular agent harnesses including Pi, OpenCode, Hermes Agents, and OpenClaw. This breadth of integration points suggests that Hugging Face is positioning itself as a unified interface where developers can compare and switch between multiple inference providers without rewriting application code. The mention that Baseten will add support for additional task types (beyond the initial conversational and text-generation tasks) indicates this partnership is intended to deepen over time.

FAQ
How do I get started with Baseten on Hugging Face?
You can access Baseten models through the Hugging Face website UI (model pages will show Baseten as a compatible provider), the Python SDK (huggingface_hub >= 1.26.1), the JavaScript SDK (@huggingface/inference), or integrated agent harnesses like Pi, OpenCode, and Hermes Agents. You authenticate using your Hugging Face token, and requests are routed to Baseten automatically.
How much does it cost to use Baseten through Hugging Face?
For routed requests (when you authenticate via Hugging Face Hub), you pay the standard Baseten provider API rates with no additional markup from Hugging Face. For direct requests using your own Baseten API key, you are billed on your Baseten account. Hugging Face PRO users receive $2 worth of Inference credits every month usable across providers.
What models does Baseten support on Hugging Face right now?
Initial support covers conversational and text-generation tasks, with popular open-weight LLMs including DeepSeek V4 Flash, Kimi K3, and GLM-5.2. The full list of supported models is available at https://huggingface.co/baseten, and additional task types will roll out soon.
Hugging Face BlogRead Original Article

Get the latest Open-Source AI news every morning

For example, today's edition would include:

  • NVIDIA Open Agent Safety Platform targets agent trust, Cisco joinsTop Companies AI · 6h ago
  • NVIDIA Open Agent Safety Platform targets AI agent trustTop Companies AI · 6h ago
  • Corporate America embraces cheaper open AI modelsTop Companies AI · 6h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleApple researchers propose method to lock open-weight models against unauthorized fine-tuning