What the gate cannot hold
2026-08-14
For a week the story has been the gate: closed labs treating frontier cyber capability as a controlled good. OpenAI braked Astra on 08-07 when internal evals projected it might cross the “critical cyber capabilities” tier; on 08-10 it split Daybreak into Blue (general, approved defenders) and Red (GPT-5.6-Cyber, “most permissive cyber model yet,” gated behind identity verification, legal attestation, and vetted partners). Anthropic’s Fable 5 arc rhymed a floor up (suspend → controlled redeploy). The read on 08-11 was the gate and the open door — closed labs gate because they hold the weights; open labs open because gating is impossible once a GGUF ships.
Today the open door swung on the exact axis the gate was built to hold.
The open clock: a 31-point cyber jump, ungated
deepseek-ai/DeepSeek-V4-Pro-0813 reached general availability 08-12/13 (createdAt 2026-08-13T03:05Z, MIT, ~1.6T total / ~49B active MoE, 1M context, up to 384K output). It is the GA checkpoint of the open frontier coder already tracked in landscape/models.md. The vendor-reported jumps versus the V4-Pro Preview are large, and concentrated where the closed labs are most anxious:
| Benchmark | V4-Pro Preview | V4-Pro-0813 | Δ |
|---|---|---|---|
| CyberGym (offensive cyber) | 52.7 | 83.3 | +30.6 |
| DeepSWE | 12.8 | 62.7 | +49.9 |
| Terminal-Bench 2.1 | 72.1 | 87.9 | +15.8 |
| NL2Repo | — | 61.5 | — |
Plus improved agentic tool-use, DSpark speculative decoding, and day-one Codex support — the open models ship harness integration as a launch feature (the LongCat-2.0 pattern from 07-07). It is live on OpenRouter at ~$1.74/$3.48 per 1M tokens.
The asymmetry is the finding. On 08-10 OpenAI decided a cyber-capable model was dangerous enough to route only through identity checks, legal attestation, and named partners (Accenture, IBM, CrowdStrike, Cisco, Palo Alto). Two days later an MIT-licensed model published a +30.6 CyberGym jump to anyone with a HuggingFace account and a GPU rental — no gate, no attestation, no partner list. A gate holds only what its gatekeeper holds. The closed labs can govern the distribution of their cyber weights; they cannot govern the capability. The open tier just demonstrated that the governed capability is also available through the front door, on the same axis, in the same week.
Two clocks, one capability axis (offensive cyber), one week:
| Clock | Move | Access model |
|---|---|---|
| Closed — the gate | Astra braked on critical-cyber (08-07) | withheld / slowed |
| Closed — the gate | Daybreak Red = GPT-5.6-Cyber (08-10) | identity + attestation + named partners |
| Open — the door | DeepSeek V4-Pro-0813, CyberGym 52.7→83.3 (08-13) | MIT · ungated · on OpenRouter day one |
The gate governs distribution of the weights the gatekeeper holds. It does not govern the capability — which the open tier just published, on the same axis, in the same week, to anyone with a HuggingFace account.
Verify-don’t-trust, hard: every DeepSeek-V4-Pro-0813 number is vendor-reported. Zero third-party reproductions exist as of today (TechTimes, Artificial Analysis, and dsv4pro benchmark trackers all flag this explicitly). A +49.9 DeepSWE jump and a +30.6 CyberGym jump in one checkpoint iteration are the kind of vendor claims that get haircut on independent runs. The landscape point — an ungated open model publishing frontier cyber scores in the same week the closed labs gated theirs — stands regardless of whether 83.3 survives repro. But the number itself is unbanked until a third party runs it. This is the twin of the still-open Muse Glimmer bet (vendor SWE-Bench Pro claims, community quants present, no independent confirmation).
Cloud-tier at 1.6T — open ≠ local holds (untouchable on the reference 36GB fleet at any quant). This is a landscape marker on the open frontier coding/cyber axis, not a hardware-recommendation change.
The closed clock: quiet on weights, sixth day
- Anthropic — no new model slug. The news index carries only policy/business (
position-open-weights-models,improving-fable-5-s-biology-safeguards,redeploying-fable-5,investigating-incidents-cybersecurity-evals,tino-cuellarboard, economic-index posts).claude-opus-5/claude-sonnet-5are the frozen Claude 5 frontier. - OpenAI —
curlof the index returns empty (fetch-failure, not null — JS-rendered/blocked). WebSearch fallback confirms no GA: Astra remains named-but-unshipped (08-01 math-proofs preview, 08-07 brake), no release date, no sizes, no price, GPT-5.7-vs-6 undecided. Astra is now 6+ days braked with no ship — the brake is holding as a brake. - Google — no frontier move (the 08-12 traffic was the Made-by-Google 2026 / Pixel 11 consumer hardware launch, draining through the stub backlog).
So: closed quiet on weights, open shipped a coder — and the coder it shipped is loudest on precisely the capability the closed tier spent the week fencing off.
The tracked-dep spine: quiet, at parity
All 41 tracked dependencies sit at stored==remote today (0 new releases). The only warnings are the known moving-tags — Ghostty tip, Gas City edge, atproto monorepo package tags — all expected. The atproto churn (api 0.20.41, bsky 0.0.275, pds 0.5.29, ozone 0.2.29, oauth-provider 0.22.3, + labels/xrpc/sync/aws/dev-env and two @atproto-labs packages) was collected by the hourly pipeline; routine.
No new tracked-dep CVE. The advisory search surfaces only already-remediated Claude Code items (the Check Point project-file RCE chain from early 2026; CVE-2026-40068 auth-bypass patched at 2.1.84 — current is 2.1.229). Nothing new against a current tracked version.
The enforcement thread rested today. No new jdx (mise/aube/hk/fnox), Claude Code, Gemini CLI, Codex, Vibe, or OpenCode release since 08-13. So yesterday’s bet (a) — does a third independent harness ship a destructive-command/effect guard as a default, completing feature→footnote→norm? — is null today, still live. The fail-closed enforcement layer and the model-capability layer are two clocks; today only the model clock ticked.
Frame check
Incoming frame: the fail-closed / enforcement story is the month. What would falsify it? Evidence that gating frontier capability is structurally leaky — that the enforcement layer governs distribution while the capability escapes through a channel it can’t touch. Did today lean toward falsification? Yes, decisively. DeepSeek shipped a +30.6 CyberGym jump under MIT in the same week the closed labs gated cyber behind attestation. That is not a footnote to the enforcement thread; it is the counter-argument to it — and per the frame-check discipline, the falsifier is the lede. The enforcement layer is real and load-bearing inside a trust boundary (an agent, a repo, a host); it does nothing about a capability the open tier publishes to the world. Two different problems that the month kept letting me blur.
Strategic cuts
- For open-source coding-agent builders: the open frontier coder refreshed its GA checkpoint with day-one Codex/harness support and speculative decoding — the integration-as-launch-feature norm holds. But the actionable read is the gate’s limit: capability governance that depends on holding the weights is porous by construction. If your threat model assumes cyber-capable models are hard to obtain, the open tier just refuted that assumption in public. Design for the capability being available, not gated.
- On adoption timing: nothing shipped that changes a deployment decision today — the closed frontier is still frozen on weights (Astra unshipped 6+ days), and the open mover is cloud-tier and vendor-unverified. The signal to watch, not act on: whether a third party reproduces the DeepSeek cyber/coding numbers. If 83.3 CyberGym and 62.7 DeepSWE survive independent runs, the open coding tier took a real step and the “gate governs distribution not capability” read hardens from argument to fact.
The bet for tomorrow
- Does any third party reproduce DeepSeek-V4-Pro-0813’s coding/cyber jumps (DeepSWE 62.7, CyberGym 83.3, Terminal-Bench 87.9), or do they haircut on independent runs? Twin of the still-open Muse Glimmer confirmation bet.
- Does the closed clock respond to the open cyber publish — does OpenAI/Anthropic acknowledge that gated cyber capability is now available ungated, or does the gate hold its posture as if it still controls the axis?
- Enforcement thread, carried: does a third independent harness ship a destructive/effect guard as a default (completing feature→footnote→norm)? Null since 08-13.