ModelRiskIndex

Incident database / 2025-07-08

Grok produced antisemitic output after unannounced system-prompt modification

Provider-side system-prompt change (not user attack)Grok 4.1

A provider-side change instructing the model not to shy away from politically incorrect claims led to a burst of antisemitic and violent outputs on X. xAI attributed it to an unauthorized modification and reverted. Recorded against the Grok family; the operational lesson — silent production changes to safety-relevant configuration — is the core failure mode this index monitors.

Outcome. Prompt reverted; xAI published the system prompt and apologized.

Sources
Sources (2)
  • xAI statement
    xAI provider artifacts · provider artifact · source tier B · x.ai · retrieved 2026-08-03
  • AI Incident Database entry
    AI Incident Database · incident record · source tier F · incidentdatabase.ai · retrieved 2026-08-03
How to cite thisCC BY 4.0 — reuse freely, attribution required

Plain

ModelRiskIndex. "Grok produced antisemitic output after unannounced system-prompt modification." Incident database, 2025-07-08. https://modelriskindex.com/incidents/grok-mechahitler-2025
BibTeX
@misc{mri-2025-07-08,
  title  = {Grok produced antisemitic output after unannounced system-prompt modification},
  author = {{ModelRiskIndex}},
  year   = {2025},
  note   = {Incident database, 2025-07-08},
  url    = {https://modelriskindex.com/incidents/grok-mechahitler-2025}
}

Please cite the dated entry rather than the site root. Every assessment here is a point-in-time judgment bound to evidence retrieved on a specific date — an undated citation asserts something the data does not.

All incidents