2026-07-20

2026-07-20-kimi-k3-open-weights-escrowed

models

Summary

Moonshot AI launched Kimi K3 on 07-16 via the Kimi app, Kimi Code, and the Kimi API: ~2.8T total parameters, Mixture-of-Experts with 16 of 896 experts active per token (~1.8%, ≈50B active) under a “Stable LatentMoE” framework, 1M-token context, multimodal (text/image/video), reported ~2.5× K2’s overall scaling efficiency. VentureBeat calls it “the largest open-source model ever built”; Simon Willison covered it day-one; Moonshot paused new K3 subscriptions on 07-20 under a demand surge. Press framed it as a “new DeepSeek shock,” closing the open-vs-closed gap to roughly one generation.

The catch is the whole point: it is live but not yet open. The API and product are live now; the open weights are promised by 07-27. Right now K3 is API-and-benchmarks — the artifact that makes it matter for a local-first practitioner is escrowed for a week. This is a self-imposed dated escrow (demand management, not a government recall like Fable), but structurally the same announce-then-ship shape as the GPT-5.6 preview→GA gate.

Verify-don’t-trust: benchmarks are lab-reported; the 07-27 weight date is a stated target that could slip (cf. Gemini 3.5 Pro, five weeks of “next month”).

Implications

Feeds [open coding/agentic frontier], [open ≠ local], and the [announce-now-ship-later] cluster (this week’s 07-27/28 convergence: MCP final, K3 weights, Grok 4.6 tease).

  • Open ≠ local, increasingly. The open frontier is scaling GLM-5.2 (753B) → Inkling (975B) → Kimi K3 (2.8T). Active params stay efficient (~40–50B across all three), but total params — the disk-and-RAM constraint — climb into cloud-only territory. K3 at 4-bit ≈ 1.4TB; even an aggressive 1.5-bit community quant lands ~500GB, nowhere near an M3 Max’s 36GB. Local fit for reference hardware: ❌❌. The open-weight and locally-runnable axes are diverging — the frontier stays open (you may download it) while ceasing to be local (you can’t run it). Reference machines get derivatives and sub-30B bases, not the open frontier.
  • The credibility spectrum. K3 sits mid-spectrum in this week’s announce-then-ship cluster: firmer than a Musk Grok-4.6 tease (the model is API-live and benchmarked), softer than the MCP RC (which has beta SDKs). The falsifiable content: does 07-27 land (weights downloadable) or join the escrow tail with Gemini 3.5 Pro?
  • Watch: 07-27 weight release; community MXFP4/GGUF quant sizes; whether K3 contests GLM-5.2 as the open coding head once weights exist (cloud-tier for reference hardware regardless); whether Moonshot warrants promotion to a tracked voice/org (the Kimi family is already tracked; K3 is its largest signal yet).

← all signals