ModelRiskIndex

Incident database / 2026-05-11

FAR.AI: DeepSeek V4 Pro safeguards collapse at 98–100% across three attack strategies

Public jailbreaks, authority manipulation, response prefillDeepSeek V4 Pro

FAR.AI's stress test broke DeepSeek V4 Pro's safeguards at 100% with public jailbreaks across CBRN, cyber, and terrorism domains in roughly 15 minutes, 99.6% via authority manipulation, and 99.6% via response prefill — and a jailbreak written for V3.2 worked on V4 Pro unmodified. Neo Research separately raised the StrongREJECT jailbreak rate from 0.6% to 77.8% with a single 2023-era roleplay template that peer models resisted.

Outcome. No provider response; open weights preclude post-release safeguard fixes. Recorded as the best-documented safeguard failure among tracked models.

Sources
Sources (2)
How to cite thisCC BY 4.0 — reuse freely, attribution required

Plain

ModelRiskIndex. "FAR.AI: DeepSeek V4 Pro safeguards collapse at 98–100% across three attack strategies." Incident database, 2026-05-11. https://modelriskindex.com/incidents/deepseek-v4-pro-farai-collapse
BibTeX
@misc{mri-2026-05-11,
  title  = {FAR.AI: DeepSeek V4 Pro safeguards collapse at 98–100% across three attack strategies},
  author = {{ModelRiskIndex}},
  year   = {2026},
  note   = {Incident database, 2026-05-11},
  url    = {https://modelriskindex.com/incidents/deepseek-v4-pro-farai-collapse}
}

Please cite the dated entry rather than the site root. Every assessment here is a point-in-time judgment bound to evidence retrieved on a specific date — an undated citation asserts something the data does not.

All incidents