Baseten has joined Hugging Face Hub as an Inference Provider for LLMs
Baseten has become a supported Inference Provider on Hugging Face Hub for conversational and text-generation tasks with models such as Kimi K3, DeepSeek V4 Flash or GLM-5.2. It is available through Python and JavaScript SDKs and in agent tools such as OpenCode and Hermes Agents.
Baseten has been added as a supported Inference Provider on the Hugging Face Hub platform. According to Hugging Face, it is an infrastructure platform covering both serverless AI and model training, with a catalog of frontier models; the first phase of the integration brings support for conversational and text-generation tasks. Open models such as Kimi K3, DeepSeek V4 Flash or GLM-5.2 are thus available through Baseten, while support for other task types (e.g. text-to-speech) is planned to follow later, according to Hugging Face.
Models hosted on Baseten can be accessed through the Python library huggingface_hub (version 1.26.1 and higher) and the JavaScript library @huggingface/inference, or through an OpenAI-compatible interface at router.huggingface.co with authentication using a Hugging Face token. The integration is also incorporated into agent tools (Agent Harnesses) such as Pi, OpenCode, Hermes Agents or OpenClaw, so models hosted on Baseten can be used directly in these tools without additional integration code.
There are two payment methods: either using your own Baseten API key, in which case the user is billed directly by Baseten, or routing through a Hugging Face account, in which case standard provider rates apply without a markup from Hugging Face. According to Hugging Face, users with a Hugging Face PRO subscription receive a monthly credit worth 2 dollars for inference across providers, while logged-in users without a subscription have a limited free quota.
Why it matters
Developers and companies building applications on open-weight models gain another inference provider available through the unified Hugging Face interface, without having to change code when switching between providers. Routing through Hugging Face without a markup and direct integration with existing agent tools lower the barrier to trying models hosted on Baseten.
Two audiences, two different impacts
What this means
For individuals
Developers can call models hosted on Baseten (e.g. DeepSeek V4 Flash, Kimi K3, GLM-5.2) directly through the Hugging Face SDK or an OpenAI-compatible interface, including in tools such as OpenCode or Hermes Agents, without writing additional integration code.
For a business
Companies can choose whether to pay Baseten directly using their own API key or route requests through a Hugging Face account without a markup, expanding their options for managing inference costs and reducing dependence on a single provider.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.