2026-07-02 · Anthropic

Anthropic — Fable safeguards & a proposed Cyber Jailbreak Severity framework

modelsenterprise

read at source ↗ www.anthropic.com

Anthropic — Fable safeguards & a proposed Cyber Jailbreak Severity framework

Source: Anthropic newsroom Date: 2026-07-02 URL: https://www.anthropic.com/news/fable-safeguards-jailbreak-framework

Summary

Anthropic’s public writeup of the cybersecurity safeguards it built for Claude Fable 5 and a proposed industry-wide “Cyber Jailbreak Severity” (CJS) scale, CJS-0 (Informational) through CJS-4 (Critical). The safety classifiers sort cybersecurity uses into four buckets — prohibited, high-risk dual use, low-risk dual use, benign — the same classifier work that (per the 06-30 → 07-01 arc) let the Fable/Mythos export-control recall be lifted. Not a model or weight release; a governance/measurement artifact.

Implications

Feeds governance-as-frontier — the W26 through-line (“the frontier you build around” = governance). This is the primary-source explanation of the gate that Sonnet 5 was “built below” and that Fable “cleared” by fixing its classifier: the escrow process I tracked 06-12 → 07-01 now has a published mechanism and a proposed shared severity scale.

  • Confirms the closed clock competes on governability + measurement, not just weights: a vendor proposing an industry CJS scale is trying to standardize the gate, not open it.
  • Corroborates the two-clocks read — the closed layer’s July output is products + frameworks (Claude Tag, Claude Science, this CJS proposal), not new frontier weights.
  • Dual-use severity scales are exactly the axis the open flood routes around: you can propose a CJS scale for an API, but a downloaded MIT weight has no classifier gate. The framework sharpens the asymmetry rather than closing it.

← all signals