Skip to content
worth noting Tools and apps

Baseten has joined Hugging Face Hub as an Inference Provider for LLMs

only one source so far

Baseten has become a supported Inference Provider on Hugging Face Hub for conversational and text-generation tasks with models such as Kimi K3, DeepSeek V4 Flash or GLM-5.2. It is available through Python and JavaScript SDKs and in agent tools such as OpenCode and Hermes Agents.

Baseten has been added as a supported Inference Provider on the Hugging Face Hub platform. According to Hugging Face, it is an infrastructure platform covering both serverless AI and model training, with a catalog of frontier models; the first phase of the integration brings support for conversational and text-generation tasks. Open models such as Kimi K3, DeepSeek V4 Flash or GLM-5.2 are thus available through Baseten, while support for other task types (e.g. text-to-speech) is planned to follow later, according to Hugging Face.

Models hosted on Baseten can be accessed through the Python library huggingface_hub (version 1.26.1 and higher) and the JavaScript library @huggingface/inference, or through an OpenAI-compatible interface at router.huggingface.co with authentication using a Hugging Face token. The integration is also incorporated into agent tools (Agent Harnesses) such as Pi, OpenCode, Hermes Agents or OpenClaw, so models hosted on Baseten can be used directly in these tools without additional integration code.

There are two payment methods: either using your own Baseten API key, in which case the user is billed directly by Baseten, or routing through a Hugging Face account, in which case standard provider rates apply without a markup from Hugging Face. According to Hugging Face, users with a Hugging Face PRO subscription receive a monthly credit worth 2 dollars for inference across providers, while logged-in users without a subscription have a limited free quota.

What changed

Why it matters

Developers and companies building applications on open-weight models gain another inference provider available through the unified Hugging Face interface, without having to change code when switching between providers. Routing through Hugging Face without a markup and direct integration with existing agent tools lower the barrier to trying models hosted on Baseten.

Two audiences, two different impacts

What this means

01

For individuals

Developers can call models hosted on Baseten (e.g. DeepSeek V4 Flash, Kimi K3, GLM-5.2) directly through the Hugging Face SDK or an OpenAI-compatible interface, including in tools such as OpenCode or Hermes Agents, without writing additional integration code.

What to do Try calling a model through router.huggingface.co with the :baseten flag and compare it with your current provider.
More practical updates →
02

For a business

Companies can choose whether to pay Baseten directly using their own API key or route requests through a Hugging Face account without a markup, expanding their options for managing inference costs and reducing dependence on a single provider.

Development
What to decide Compare billing through your own Baseten API key with routing through a Hugging Face account and choose an option based on your cost preferences.
More business impacts →
Baseten DeepSeek Hugging Face inference API LLM serverless AI

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
Hugging Face Blog primary source · first detected Baseten on Hugging Face Inference Providers 🔥