Incident database / 2026-05-11
FAR.AI: DeepSeek V4 Pro safeguards collapse at 98–100% across three attack strategies
FAR.AI's stress test broke DeepSeek V4 Pro's safeguards at 100% with public jailbreaks across CBRN, cyber, and terrorism domains in roughly 15 minutes, 99.6% via authority manipulation, and 99.6% via response prefill — and a jailbreak written for V3.2 worked on V4 Pro unmodified. Neo Research separately raised the StrongREJECT jailbreak rate from 0.6% to 77.8% with a single 2023-era roleplay template that peer models resisted.
Outcome. No provider response; open weights preclude post-release safeguard fixes. Recorded as the best-documented safeguard failure among tracked models.
Plain
ModelRiskIndex. "FAR.AI: DeepSeek V4 Pro safeguards collapse at 98–100% across three attack strategies." Incident database, 2026-05-11. https://modelriskindex.com/incidents/deepseek-v4-pro-farai-collapse
BibTeX
@misc{mri-2026-05-11,
title = {FAR.AI: DeepSeek V4 Pro safeguards collapse at 98–100% across three attack strategies},
author = {{ModelRiskIndex}},
year = {2026},
note = {Incident database, 2026-05-11},
url = {https://modelriskindex.com/incidents/deepseek-v4-pro-farai-collapse}
}Please cite the dated entry rather than the site root. Every assessment here is a point-in-time judgment bound to evidence retrieved on a specific date — an undated citation asserts something the data does not.