ModelRiskIndex

Rankings / OpenAI

GPT-5.1

Tier 280/100legacy

gpt-5.1-2025-11-13 via api.openai.com

legacy. Superseded through a rapid cadence (5.2, 5.3-Codex, 5.4, 5.5, 5.6). Variant shutdowns already underway: gpt-5.1-codex/codex-max shut down 2026-07-23; gpt-5.1-chat-latest shuts down 2026-08-10. The base dated snapshot remains live but deprecation should be expected (~8-month snapshot-to-shutdown cadence observed for GPT-5 base). Successor: GPT-5.6 family (gpt-5.6-sol flagship, GA 2026-07-09).

Usage share 0.04% · OpenRouter rankings API (daily token share, 2026-08-04)

Re-verified against live sources 2026-08-05. OpenAI's platform docs migrated to developers.openai.com; citations updated. ISO/IEC 42001 was 'not verified' at seeding and is now confirmed via the trust portal.

Tier assessment
Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.
  • Published safety evals. Safety evaluations for this model are published.
  • Documented safety policy. A documented safety, usage, or acceptable-use policy governs the model.
  • Enterprise data controls. Customer data is not used for training by default, or a documented opt-out exists.
Tier 2 requirements
  • External pre-deployment testing. Independent external parties tested the model before deployment, and this is disclosed.US CAISI and UK AISI pre-deployment access disclosed for the GPT-5 family; UK AISI published its own GPT-5.5 evaluation.
  • Third-party certification. The operating organization holds verifiable third-party certification (e.g. SOC 2, ISO/IEC 42001).
  • Versioning with changelogs. Model versions are explicitly identified and changes are changelogged.Dated snapshots and a deprecations page; alias-level behavior changes are not always changelogged.
  • Stated deprecation policy. A deprecation policy with notice windows is published.
Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governancestrong

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

API inputs not used for training by default; ~30-day abuse-monitoring retention; ZDR by approval for eligible endpoints. Litigation caveat: the NYT-case blanket preservation order was lifted 2025-09-26 (ZDR customers were always excluded), but a court affirmed production of 20M de-identified chat logs in Jan 2026 and a July 2026 sanctions motion disputes OpenAI's log handling — retention promises are currently entangled with active litigation.

Receipts (2)
  • How we use your data (API platform docs)
    OpenAI provider artifacts · provider artifact · source tier B · developers.openai.com · retrieved 2026-08-05
    As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us).
  • Preservation order narrowed (NYT v. OpenAI)
    Technology press reporting · reporting · source tier F · engadget.com · retrieved 2026-08-05
    However, this latest decision means the AI giant no longer has to preserve chat logs as of September 26, except for some.

Operational stabilitypartial

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Dated snapshots and a maintained deprecations page with stated replacements — genuinely good hygiene — but the observed cadence is aggressive: GPT-5.1's codex variants went from release to shutdown in about eight months, and alias-level behavior changes still ship without changelog entries.

Receipts (1)
  • OpenAI API deprecations page
    OpenAI provider artifacts · provider artifact · source tier B · developers.openai.com · retrieved 2026-08-05
    This page lists all API deprecations, along with recommended replacements.

Adversarial resistancepartial

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistancepartial

Safe-completions training documented in the GPT-5 system card; the GPT-5.x family sits mid-tier on F5's CASI board (GPT-5.5 at 78.47 in July 2026, against Claude leaders above 89). GPT-5.1 itself has rotated off the current board — family-level scores stand in.

Prompt injection (agentic)partial

Instruction-hierarchy training is documented, but on Gray Swan's agentic prompt-injection benchmark GPT-5.1 recorded a 21.9% attack success rate — versus 12.5% for Gemini 3 Pro and 4.7% for Claude Opus 4.5.

Receipts (4)
  • GPT-5 system card (safe-completions, jailbreaks, instruction hierarchy)
    OpenAI provider artifacts · provider artifact · source tier B · cdn.openai.com · retrieved 2026-08-05
    All of the GPT-5 models additionally feature safe-completions, our latest approach to safety training to prevent disallowed content.
  • F5 Labs CASI/ARS leaderboards (July 2026 update)
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-05
    Anthropic maintains dominance at the very top (93+ range), while OpenAI and Alibaba models show strong upward momentum in the 78-83 range, and new providers like nVidia are entering the competitive field.
  • Gray Swan agentic injection benchmark results (GPT-5.1 ASR 21.9%)
    Technology press reporting · reporting · source tier F · the-decoder.com · retrieved 2026-08-05
    A benchmark by the security firm Gray Swan found that a single "very strong" prompt injection attack breaks through Opus 4.5's safeguards 4.7 percent of the time.
    Reporting on Gray Swan benchmark data published alongside the Opus 4.5 release, Nov 2025.
  • GPT-5 system card (prompt injections section)
    OpenAI provider artifacts · provider artifact · source tier B · cdn.openai.com · retrieved 2026-08-05
    To mitigate this, we use a multilayered defense stack including teaching models to ignore prompt injections in web or connector contents.

Transparencystrong

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

System cards for the family (now hosted on a dedicated deployment-safety hub), a published Model Spec, and disclosed pre-deployment access for US CAISI and UK AISI — continued through GPT-5.5, whose cyber-capability evaluation UK AISI published itself.

Receipts (2)
  • OpenAI deployment safety hub (system cards, external evaluations)
    OpenAI provider artifacts · provider artifact · source tier B · deploymentsafety.openai.com · retrieved 2026-08-05
    Sharing the technical work we do to make our systems safe, including how deployed models perform in evaluations, the risks we measure, and the steps we take to improve over time.
  • UK AISI: evaluation of GPT-5.5 cyber capabilities
    UK AI Security Institute publications · independent eval · source tier C · aisi.gov.uk · retrieved 2026-08-05
    Separately, we conducted expert red-teaming on GPT-5.5’s cyber safeguards.

Compliance posturestrong

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

Trust portal lists SOC 2 Type 2, SOC 3, CSA STAR, ISO/IEC 42001:2023, ISO 27001/27017/27018/27701, PCI DSS, FedRAMP 20x; HIPAA BAAs available for ZDR/modified-retention-eligible API endpoints.

Receipts (2)
  • OpenAI trust portal
    OpenAI provider artifacts · provider artifact · source tier B · trust.openai.com · retrieved 2026-08-05
    Our products are covered in our SOC 2 Type 2 report and have been evaluated by an independent third-party auditor to confirm that our controls align with industry standards for security, confidentiality, privacy and availability.
  • HIPAA BAA eligibility (help center)
    OpenAI provider artifacts · provider artifact · source tier B · help.openai.com · retrieved 2026-08-05
    Most API services are covered, with a few exceptions listed here

Governance & evidence

Where your data goes

  • US
  • EU
  • UK
  • AU
  • CA
  • JP
  • IN
  • SG
  • KR
  • AE

Regional pinning available — customers can pin processing to a chosen region.

US default; regional storage in EEA/CH, UK, AU, CA, JP, IN, SG, KR, UAE — most non-US regions require eligibility approval.

Enterprise vs consumer terms

Enterprise vs consumer gap: narrow. API and enterprise tiers carry clean no-train defaults; consumer ChatGPT terms differ but the gap is narrower than peer providers'.CONSUMERENTERPRISE / APIWORSE TERMS →
Narrow gap

API and enterprise tiers carry clean no-train defaults; consumer ChatGPT terms differ but the gap is narrower than peer providers'.

Change cadence

within cadence4 tracked changes · typical interval ~7d 0d since last change (2026-08-05) Low confidence: only 4 events0d

4 tracked changes. The typical (median) interval between them is ~7 days (low confidence: a median of only 3 intervals). The last change was on 2026-08-05, 0 days before the as-of date (2026-08-05) — about 0.0× the typical interval. That is within its historical cadence.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

12 evidence refs5 distinct sources2 independent
  • CF5 Labs CASI/ARS leaderboardindependent eval×1 reference
  • CUK AI Security Institute publicationsindependent eval×1 reference
  • BOpenAI provider artifactsprovider artifact×7 references
  • EOpenRouter model rankingsusage data×1 reference
  • FTechnology press reportingreporting×2 references

retrieved 2026-08-05

Compliance & deployment

Trains on customer data by default
No
SOC 2
Yes
ISO/IEC 42001
Yes
HIPAA eligible
Yes
Retention window
~30 days (API abuse monitoring); ZDR by approval for eligible endpoints; litigation-related preservation applies to flagged accounts and an Apr–Sep 2025 corpus
Data residency
US default; regional storage in EEA/CH, UK, AU, CA, JP, IN, SG, KR, UAE (most non-US regions require eligibility approval)
EU AI Act
Full signatory of the EU GPAI Code of Practice (all three chapters, July 2025); enforcement phase began Aug 2026 with reported gaps in training-data transparency disclosure.
Deprecation policy
Published deprecations page with dated timelines and stated replacement models
Available via
OpenAI API · Microsoft Foundry (Azure AI Foundry)

Incident history

2025-09-18
ShadowLeak: zero-click data exfiltration via ChatGPT Deep Research email integration
Indirect prompt injection (zero-click, service-side exfiltration)

Researchers demonstrated a service-side indirect injection: a crafted email caused the Deep Research agent, when later asked to summarize the inbox, to exfiltrate mailbox data to an attacker-controlled URL without user interaction. Patched by OpenAI after disclosure. Tagged to the GPT-5 family as the underlying agent model.

Outcome: Patched following responsible disclosure.

Sources (1)

Change timeline

2026-08-05
Verification pass: top-3 entries re-verified against live sources; usage shares corrected
methodologynotice

Every citation for GPT-5.1, Claude Sonnet 4.5, and Gemini 3 Pro was fetched and checked. Grades held; corrections were citations, migrated documentation domains (claude.com, developers.openai.com), and facts: ISO/IEC 42001 confirmed for OpenAI (previously unverified), lifecycle markers added (all three models are now legacy or retired), and seeded usage-share estimates replaced with measured OpenRouter daily data — under which every tracked model now sits below 1% share.

Evidence (2)
2026-07-16
Developer reports of behavior shift on gpt-5.1 alias without changelog entry
versionnotice

Multiple developer reports of changed refusal behavior on the floating gpt-5.1 alias while dated snapshots remained stable. No corresponding entry on the deprecations/changelog pages. Illustrates the alias-vs-snapshot distinction this index tracks.

Evidence (1)
  • OpenAI developer community thread
    OpenAI developer community · reporting · source tier F · community.openai.com · retrieved 2026-07-16
    Curated manually; endpoint fingerprinting (Phase 2) will verify future occurrences.
2026-07-09
GPT-5.6 family GA; GPT-5.1 variant shutdowns scheduled and executed
versionwarning

GPT-5.6 (Sol/Terra/Luna) reached general availability as OpenAI's flagship line. Per the deprecations page: gpt-5.1-codex and codex-max shut down 2026-07-23, gpt-5.1-chat-latest shuts down 2026-08-10, and GPT-5 base snapshots retire 2026-12-11. The dated gpt-5.1-2025-11-13 snapshot remains live, for now.

Evidence (1)
  • OpenAI API deprecations page
    OpenAI provider artifacts · provider artifact · source tier B · developers.openai.com · retrieved 2026-08-05
2026-07-09
NYT v. OpenAI: sanctions motion over log handling; retention promises remain litigation-entangled
policynotice

After the blanket preservation order was narrowed in Sept 2025 and a court affirmed production of 20M de-identified chat logs in Jan 2026, NYT and Daily News filed a sanctions motion alleging OpenAI deleted logs subject to preservation. API zero-data-retention customers remain excluded, but buyers relying on OpenAI's retention story should track the case.

Evidence (1)

Compare this model: GPT-5.1 + open compare view →