ModelRiskIndex

Rankings / OpenAI

GPT-5.6 (Sol · Terra · Luna)

Tier 180/100

gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna via api.openai.com — no dated snapshots exist

Usage share 6.3% · OpenRouter rankings API (daily token share, 2026-08-04)

Current OpenAI flagship line. Launch was government-gated: a June 2026 executive order created a voluntary pre-release review process, and the White House asked OpenAI to restrict the June 26 preview to ~20 vetted US partners before the July 9 GA — recorded in the change feed. Tier note: this family fails Tier 2 solely on versioning (no dated snapshots), an assessment the checklist makes mechanical rather than editorial.

Tier assessment
Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.
  • Published safety evals. Safety evaluations for this model are published.
  • Documented safety policy. A documented safety, usage, or acceptable-use policy governs the model.
  • Enterprise data controls. Customer data is not used for training by default, or a documented opt-out exists.
Tier 2 requirements
  • External pre-deployment testing. Independent external parties tested the model before deployment, and this is disclosed.Six named external testers including UK AISI and US CAISI, disclosed in the system card.
  • Third-party certification. The operating organization holds verifiable third-party certification (e.g. SOC 2, ISO/IEC 42001).
  • Versioning with changelogs. No versioning discipline or changelog for model changes.No dated snapshots exist for any GPT-5.6 variant — versions cannot be pinned and updates under the floating IDs would be externally indistinguishable. Earlier GPT families satisfied this requirement; 5.6 does not.
  • Stated deprecation policy. A deprecation policy with notice windows is published.

Missing for Tier 2: versioning with changelogs. The tier is computed from this checklist — satisfying these requirements moves the badge, automatically.

Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governancestrong

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

Unchanged from the verified GPT-5.1 posture: no training on API business data by default, ~30-day abuse-monitoring retention, ZDR by approval. The NYT-litigation caveat on OpenAI's retention story (see the GPT-5.1 entry and change feed) applies family-wide.

Receipts (1)
  • How we use your data (API platform docs)
    OpenAI provider artifacts · provider artifact · source tier B · developers.openai.com · retrieved 2026-08-05
    Eligible customers may have their customer content excluded from these abuse monitoring logs, subject to the limitations below, by getting approved for the Zero Data Retention or Modified Abuse Monitoring controls.

Operational stabilitypartial

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Stated deprecation policy is good (≥6 months notice for GA models) — but no dated snapshots exist for any 5.6 variant. Users cannot pin a version, the gpt-5.6 alias floats to Sol, and a model update under these IDs would be indistinguishable from the outside. A regression from the dated-snapshot discipline of earlier GPT families.

Receipts (2)
  • GPT-5.6 Sol model page (single undated snapshot)
    OpenAI provider artifacts · provider artifact · source tier B · developers.openai.com · retrieved 2026-08-05
    The gpt-5.6 alias routes requests to GPT-5.6 Sol.
  • OpenAI API deprecations page (notice policy)
    OpenAI provider artifacts · provider artifact · source tier B · developers.openai.com · retrieved 2026-08-05
    Unless safety or compliance concerns require a faster timeline, we provide the following minimum notice periods before model retirement: Generally available models: At least 6 months.

Adversarial resistancepartial

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistancepartial

The most heavily safeguarded OpenAI release to date — first family rated Preparedness-High for both cyber and bio across all sizes, with activation classifiers and 700K+ GPU-hours of automated red-teaming — yet UK AISI found universal jailbreaks in Sol unlocking autonomous cyber-exploitation pre-deployment (specific exploits mitigated), and public jailbreak claims landed within days of GA.

Prompt injection (agentic)partial

Rare quantitative per-variant disclosure: injection resistance 1.000 on Connectors but 0.910 (Sol) / 0.946 (Terra) / 0.897 (Luna) on search-plus-function-calling — the surface where agents actually operate, with the mid-tier Terra oddly beating the flagship. GPT-Red automated red-teaming (added to the card 2026-08-03) still finds working injection attacks.

Receipts (4)
  • GPT-5.6 system card (jailbreak robustness, Preparedness classifications)
    OpenAI provider artifacts · provider artifact · source tier B · deploymentsafety.openai.com · retrieved 2026-08-05
    Under our Preparedness Framework, we are treating Sol, Terra and Luna as High capability in both Cybersecurity and Biological and Chemical risk.
  • Fortune: UK AISI universal jailbreaks in GPT-5.6 Sol
    Technology press reporting · reporting · source tier F · fortune.com · retrieved 2026-08-05
    OpenAI markets its latest model, GPT-5.6 Sol, as its most secure to date, but the British government researchers who tested it prior to release say the model’s guardrails are susceptible to jailbreaks that can unlock dangerous cyber capabilities.
  • F5 Labs CASI/ARS leaderboards
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-05
    Anthropic maintains dominance at the very top (93+ range), while OpenAI and Alibaba models show strong upward momentum in the 78-83 range, and new providers like nVidia are entering the competitive field.
    Not yet listed on the July 6 board (predates GA); GPT-5.x family has tested mid-tier.
  • GPT-5.6 evaluations with challenging prompts (per-variant injection scores)
    OpenAI provider artifacts · provider artifact · source tier B · deploymentsafety.openai.com · retrieved 2026-08-05
    The model is currently particularly strong at finding prompt injection attacks that utilize human-like strategies and can probe deeper into the defender model’s weaknesses.

Transparencystrong

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Arguably the strongest disclosure package of any tracked model: preview and GA system cards, six named external testers (UK AISI, US CAISI, SecureBio, Irregular, METR, Apollo Research), quantitative per-variant safety scores, and post-release red-teaming results appended to the card.

Receipts (2)
  • GPT-5.6 system card
    OpenAI provider artifacts · provider artifact · source tier B · deploymentsafety.openai.com · retrieved 2026-08-05
    As part of our ongoing partnership with the UK AI Security Institute (UK AISI), we provided UK AISI with early access to GPT-5.6 Sol to support independent, pre-deployment evaluation of cyber capabilities.
  • GPT-5.6 preview system card (limited-preview period)
    OpenAI provider artifacts · provider artifact · source tier B · deploymentsafety.openai.com · retrieved 2026-08-05
    At their request, we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government, before releasing more broadly.

Compliance posturestrong

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

SOC 2 Type 2, ISO/IEC 42001:2023, ISO 27001/27017/27018/27701, CSA STAR, PCI DSS, FedRAMP 20x per the trust portal; HIPAA BAAs for eligible API configurations.

Receipts (1)
  • OpenAI trust portal
    OpenAI provider artifacts · provider artifact · source tier B · trust.openai.com · retrieved 2026-08-05
    Our products are covered in our SOC 2 Type 2 report and have been evaluated by an independent third-party auditor to confirm that our controls align with industry standards for security, confidentiality, privacy and availability.

Governance & evidence

Where your data goes

  • US
  • EU
  • UK
  • AU
  • CA
  • JP
  • IN
  • SG
  • KR
  • AE

Regional pinning available — customers can pin processing to a chosen region.

US default; regional storage options per the GPT-5.1 entry, eligibility-gated outside the US.

Enterprise vs consumer terms

Enterprise vs consumer gap: narrow. Same posture as GPT-5.1: API/enterprise no-train defaults are clean while consumer ChatGPT terms differ modestly.CONSUMERENTERPRISE / APIWORSE TERMS →
Narrow gap

Same posture as GPT-5.1: API/enterprise no-train defaults are clean while consumer ChatGPT terms differ modestly.

Change cadence

insufficient history1 tracked change · last 2026-06-26 · 40d since Insufficient history to estimate a cadence.40d

One tracked change, on 2026-06-26 — 40 days before the as-of date (2026-08-05). A single event cannot establish a cadence, so days-since is shown without a baseline.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

11 evidence refs4 distinct sources1 independent
  • CF5 Labs CASI/ARS leaderboardindependent eval×1 reference
  • BOpenAI provider artifactsprovider artifact×8 references
  • EOpenRouter model rankingsusage data×1 reference
  • FTechnology press reportingreporting×1 reference

retrieved 2026-08-05

Compliance & deployment

Trains on customer data by default
No
SOC 2
Yes
ISO/IEC 42001
Yes
HIPAA eligible
Yes
Retention window
~30 days (API abuse monitoring); ZDR by approval; NYT-litigation preservation caveats apply to flagged accounts
Data residency
US default; regional storage options per the GPT-5.1 entry (eligibility-gated outside the US)
EU AI Act
Full EU GPAI Code of Practice signatory; enforcement-phase transparency gaps reported July 2026.
Deprecation policy
≥6 months notice for GA models; ~3 months specialized; as little as 2 weeks for previews
Available via
OpenAI API · Microsoft Foundry (Azure AI Foundry)

Incident history

2026-07-10
UK AISI found universal jailbreaks in GPT-5.6 Sol unlocking autonomous cyber-exploitation
Universal jailbreaks (pre-deployment external testing)

During pre-deployment testing, UK AISI developed universal jailbreaks — often within hours — that unlocked Sol's Preparedness-High cyber capabilities for autonomous exploitation. Disclosed in the system card; OpenAI reproduced and mitigated the specific jailbreaks before GA. Notable both for the finding and for the disclosure: an adverse pre-release result published by the vendor.

Outcome: Specific jailbreaks mitigated pre-GA; public jailbreak claims against Sol continued post-release.

Sources (2)

Change timeline

2026-06-26
GPT-5.6 launch gated by White House request under June executive order
regulationnotice

Under a voluntary pre-release review process created by a June 2 executive order, the White House asked OpenAI to restrict the GPT-5.6 preview to roughly twenty government-vetted US partners, citing Sol's cyber capabilities. Broad rollout cleared on July 9. The first instance of US government pre-release gating shaping a frontier model launch.

Evidence (2)

Compare this model: GPT-5.6 (Sol · Terra · Luna) + open compare view →