Baseten on Hugging Face Inference Providers 🔥
read at source ↗ huggingface.co
Baseten on Hugging Face Inference Providers 🔥
Source: HuggingFace Date: 2026-08-06 URL: https://huggingface.co/blog/baseten
Summary
Baseten joined Hugging Face’s Inference Providers roster, serving DeepSeek V4 Flash, Kimi K3, and GLM-5.2 through either HF-routed calls (no separate API key) or direct Baseten credentials, with SDK support in huggingface_hub and @huggingface/inference and compatibility with agent harnesses including OpenCode. HF Pro subscribers get $2/month in inference credit usable across any provider on the roster.
Implications
Feeds open vs local model clocks and cost/economy of AI together: this is infrastructure normalizing access to the current large open-weight tier (Kimi K3 at ~1.56TB, DeepSeek V4 Flash at 304B) for people who can’t self-host it — the “open ≠ local” gap the landscape has been tracking gets partly bridged by inference-as-a-service rather than by the models getting smaller. It’s also an agent-layer convergence data point: explicit agent-harness compatibility (OpenCode, Hermes Agents) as a listed integration feature means inference providers are now building for agentic callers as a first-class use case, not just chat completion.