ModelRiskIndex

Rankings / xAI

Grok 4.1

Tier 120/100draft — pending re-verificationlegacy

grok-4-1 via api.x.ai

legacy. No longer appears in OpenRouter rankings as of 2026-08-04; xAI's current lineup there is Grok 4.3/4.5/4.20. Full re-verification of the successor lineup pending. Successor: Grok 4.3 / Grok 4.5 (per OpenRouter model listings).

Tier assessment
Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.xAI's first model card, published with Grok 4.1 in November 2025.
  • Published safety evals. Safety evaluations for this model are published.Published alongside the 4.1 model card and risk-management framework.
  • Documented safety policy. A documented safety, usage, or acceptable-use policy governs the model.
  • Enterprise data controls. Customer data is not used for training by default, or a documented opt-out exists.
Tier 2 requirements
  • External pre-deployment testing. No disclosed external pre-deployment testing.
  • Third-party certification. No verifiable third-party certification.Certification claims not independently verifiable at time of entry — no public trust portal with audit artifacts.
  • Versioning with changelogs. No versioning discipline or changelog for model changes.
  • Stated deprecation policy. No stated deprecation policy.

Missing for Tier 2: external pre-deployment testing, third-party certification, versioning with changelogs, stated deprecation policy. The tier is computed from this checklist — satisfying these requirements moves the badge, automatically.

Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governancepartial

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

Enterprise API states a no-train default; consumer X integration defaults to data use for training with opt-out buried in settings — one of the widest enterprise/consumer gaps tracked.

Receipts (1)
  • xAI privacy policy
    xAI provider artifacts · provider artifact · source tier B · x.ai · retrieved 2026-08-03
    To develop and improve our Service and to conduct research: For example to develop new product features, to train our models, to identify usage trends, to operate and expand our business activities, to identify new customers, and for data analysis.

Operational stabilityweak

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Repeated silent behavior changes in production — including system-prompt modifications acknowledged only after public incidents — and no changelog discipline.

Receipts (1)
  • xAI statement on July 2025 Grok incident
    xAI provider artifacts · provider artifact · source tier B · x.ai · retrieved 2026-08-03
    See linked incident record for the July 2025 unauthorized-modification episode.

Adversarial resistanceweak

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistanceweak

Public jailbreaks land quickly after each release; permissive persona modes materially weaken the model's own stated policies, and independent indexes place it near the bottom among frontier labs.

Prompt injection (agentic)weak

No published agentic-injection defenses or agent red-teaming results; deep X-platform integration creates a large indirect-injection surface.

Receipts (2)
  • F5 Labs CASI leaderboard
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-03
  • xAI API documentation
    xAI provider artifacts · provider artifact · source tier B · docs.x.ai · retrieved 2026-08-03

Transparencypartial

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Published its first model card with Grok 4.1 and a risk-management framework — real movement — but no external pre-deployment testing and a history of unannounced system-prompt changes.

Receipts (1)
  • Grok 4.1 model card
    xAI provider artifacts · provider artifact · source tier B · data.x.ai · retrieved 2026-08-03
    In line with our Risk Management Framework (RMF), we measure safety-relevant behaviors across three categories: abuse potential, concerning propensities, and dual-use capabilities.

Compliance postureweak

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

Certification claims are not independently verifiable at time of entry; no public trust portal with audit artifacts.

Receipts (1)

Governance & evidence

Where your data goes

  • US

No regional pinning — the provider chooses where data is processed.

Enterprise vs consumer terms

Enterprise vs consumer gap: wide. Enterprise API states a no-train default while the consumer X integration defaults to training on user data with an opt-out buried in settings — one of the widest gaps tracked.CONSUMERENTERPRISE / APIWORSE TERMS →
Wide gap

Enterprise API states a no-train default while the consumer X integration defaults to training on user data with an opt-out buried in settings — one of the widest gaps tracked.

Change cadence

insufficient history1 tracked change · last 2025-11-17 · 261d since Insufficient history to estimate a cadence.261d

One tracked change, on 2025-11-17 — 261 days before the as-of date (2026-08-05). A single event cannot establish a cadence, so days-since is shown without a baseline.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

6 evidence refs2 distinct sources1 independent
  • CF5 Labs CASI/ARS leaderboardindependent eval×1 reference
  • BxAI provider artifactsprovider artifact×5 references

retrieved 2026-08-03

Compliance & deployment

Trains on customer data by default
No
SOC 2
not verified
ISO/IEC 42001
not verified
HIPAA eligible
not verified
Retention window
Enterprise API: stated no-train default; consumer: training default with opt-out
Data residency
US
EU AI Act
Did not sign the EU GPAI Code of Practice safety chapter in full; posture unclear.
Deprecation policy
None stated
Available via
xAI API · Azure AI Foundry

Incident history

2025-07-08
Grok produced antisemitic output after unannounced system-prompt modification
Provider-side system-prompt change (not user attack)

A provider-side change instructing the model not to shy away from politically incorrect claims led to a burst of antisemitic and violent outputs on X. xAI attributed it to an unauthorized modification and reverted. Recorded against the Grok family; the operational lesson — silent production changes to safety-relevant configuration — is the core failure mode this index monitors.

Outcome: Prompt reverted; xAI published the system prompt and apologized.

Sources (2)
  • xAI statement
    xAI provider artifacts · provider artifact · source tier B · x.ai · retrieved 2026-08-03
  • AI Incident Database entry
    AI Incident Database · incident record · source tier F · incidentdatabase.ai · retrieved 2026-08-03

Change timeline

2025-11-17
xAI publishes its first model card with Grok 4.1
model-cardinfo

First formal model card from xAI, alongside a risk-management framework. Moves Grok from Tier 0 to Tier 1 under this index's definitions.

Evidence (1)
  • Grok 4.1 model card
    xAI provider artifacts · provider artifact · source tier B · data.x.ai · retrieved 2026-08-03

Compare this model: Grok 4.1 + open compare view →