ModelRiskIndex

Rankings / Anthropic

Claude Sonnet 4.5

Tier 280/100legacy

claude-sonnet-4-5-20250929 via api.anthropic.com

legacy. Still served, but listed under Legacy models after Sonnet 4.6 (Feb 2026) and Sonnet 5 (June 2026). Tentative retirement floor 2026-09-29 under Anthropic's 60-day-notice deprecation policy. Successor: Claude Sonnet 5 (claude-sonnet-5, released 2026-06-30).

Usage share 0.13% · OpenRouter rankings API (daily token share, 2026-08-04)

Re-verified against live sources 2026-08-05. Anthropic's documentation and privacy portals migrated to claude.com domains in 2026; citations updated accordingly.

Tier assessment
Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.
  • Published safety evals. Safety evaluations for this model are published.
  • Documented safety policy. A documented safety, usage, or acceptable-use policy governs the model.
  • Enterprise data controls. Customer data is not used for training by default, or a documented opt-out exists.
Tier 2 requirements
  • External pre-deployment testing. Independent external parties tested the model before deployment, and this is disclosed.UK AISI and Apollo Research pre-deployment testing disclosed in the system card; US CAISI absent for this model.
  • Third-party certification. The operating organization holds verifiable third-party certification (e.g. SOC 2, ISO/IEC 42001).
  • Versioning with changelogs. Model versions are explicitly identified and changes are changelogged.Dated pinned snapshots and release notes; safety-filter and system-prompt tuning is not always changelogged.
  • Stated deprecation policy. A deprecation policy with notice windows is published.
Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governancestrong

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

API inputs/outputs deleted from backend within ~30 days by default, not used for training; ZDR agreements and enterprise retention controls available. Consumer gap: claude.ai Free/Pro/Max training toggle was pre-selected On in a mandatory choice (deadline 2025-10-08) with 5-year retention when enabled — API and commercial tiers explicitly excluded.

Receipts (2)
  • Anthropic privacy center: organization data retention
    Anthropic provider artifacts · provider artifact · source tier B · privacy.claude.com · retrieved 2026-08-05
    For Anthropic API users, we automatically delete inputs and outputs on our backend within 30 days of receipt or generation
  • Consumer terms update (training choice)
    Anthropic provider artifacts · provider artifact · source tier B · anthropic.com · retrieved 2026-08-05
    We are also extending data retention to five years, if you allow us to use your data for model training.

Operational stabilitypartial

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Dated pinned snapshots, documented lifecycle states (Active/Legacy/Deprecated/Retired), 60-day minimum retirement notice, and published deprecation-preservation commitments; behavior-affecting safety-filter and system-prompt tuning is still not always changelogged.

Receipts (1)
  • Anthropic model deprecations documentation
    Anthropic provider artifacts · provider artifact · source tier B · platform.claude.com · retrieved 2026-08-05
    Anthropic provides a recommended replacement and assigns a retirement date.

Adversarial resistancepartial

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistancestrong

Constitutional-classifier safeguards under an ASL-3 precautionary deployment; historically top-tier on F5's CASI index (95.86), and Anthropic models still hold the top three CASI slots as of July 2026. Sonnet 4.5 itself has rotated off the current board in favor of Sonnet 5.

Prompt injection (agentic)partial

Anthropic's model report states Sonnet 4.5 achieved the lowest successful prompt-injection rate in an external red-team exercise across MCP, computer-use, and tool-use scenarios — best-in-class, but injections still succeed at nonzero rates.

Receipts (4)
  • Anthropic model report (Sonnet 4.5 section, ASL-3 deployment)
    Anthropic provider artifacts · provider artifact · source tier B · anthropic.com · retrieved 2026-08-05
    Based on our assessments of the model’s demonstrated capabilities, we determined that Claude Sonnet 4.5 did not meet the “notably more capable” threshold, described in our Responsible Scaling Policy, and decided to deploy Claude Sonnet 4.5 under the ASL-3 Standard.
  • F5 Labs CASI/ARS leaderboards (July 2026 update)
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-05
    Claude Sonnet 4.5 — 95.86 CASI, 49.6% performance, $18.78 CoS: delivers top-tier security without the traditional capability tradeoff.
    Sonnet 4.5 historical CASI 95.86; current board leads with Claude Sonnet 5 (93.08), Haiku 4.5, Opus 4.8.
  • Anthropic model report (external injection red-team results)
    Anthropic provider artifacts · provider artifact · source tier B · anthropic.com · retrieved 2026-08-05
    In an externally conducted red team exercise that evaluated 23 models from multiple AI developers, Claude Sonnet 4.5 achieved the lowest rate of successful prompt injection attacks.
  • Gray Swan × UK AISI agent red-teaming (background: every model tested was broken)
    Gray Swan agent red-teaming arena · independent eval · source tier C · grayswan.ai · retrieved 2026-08-05
    While all models were broken many times for all behaviors, measuring the attack success rate (ASR) by calculating Total Breaks / Total Chats gives a sense of each model’s relative robustness
    Mar–Apr 2025 challenge predates Sonnet 4.5 (covers Claude 3.5/3.7 era); cited as field context, not model-specific evidence.

Transparencystrong

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Dedicated system card, model report with published eval results, RSP/ASL framework, and disclosed external pre-deployment testing by UK AISI and Apollo Research. US CAISI was notably absent for this model.

Receipts (2)
  • Claude Sonnet 4.5 system card (PDF)
    Anthropic provider artifacts · provider artifact · source tier B · assets.anthropic.com · retrieved 2026-08-05
    We shared late pre-deployment snapshots of Claude Sonnet 4.5 with two organizations that are building up specialized functions for model behavior evaluation: The UK AI Security Institute (AISI) and the independent nonprofit Apollo Research.
  • Anthropic transparency hub
    Anthropic provider artifacts · provider artifact · source tier B · anthropic.com · retrieved 2026-08-05
    The following are summaries of key safety evaluations from our Claude Sonnet 4.5 system card

Compliance posturestrong

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

SOC 2 Type I & II, ISO 27001:2022, ISO/IEC 42001:2023, HIPAA-ready configuration with BAA — scoped to commercial products (Claude for Work, API), not consumer plans.

Receipts (1)
  • Anthropic certifications (privacy center)
    Anthropic provider artifacts · provider artifact · source tier B · privacy.claude.com · retrieved 2026-08-05
    maintains the following compliance credentials: HIPAA-ready configuration (BAA available) ISO 27001:2022 (Information Security Management) ISO/IEC 42001:2023 (AI Management Systems) SOC 2 Type I & Type II

Governance & evidence

Where your data goes

  • US

Regional pinning available — customers can pin processing to a chosen region.

First-party API residency unstated (US provider home jurisdiction presumed); regional endpoints via Bedrock and Vertex.

Enterprise vs consumer terms

Enterprise vs consumer gap: wide. claude.ai Free/Pro/Max training toggle was pre-selected On (5-year retention when enabled) while API and commercial tiers are explicitly excluded.CONSUMERENTERPRISE / APIWORSE TERMS →
Wide gap

claude.ai Free/Pro/Max training toggle was pre-selected On (5-year retention when enabled) while API and commercial tiers are explicitly excluded.

Change cadence

within cadence3 tracked changes · typical interval ~156d 0d since last change (2026-08-05) Low confidence: only 3 events0d

3 tracked changes. The typical (median) interval between them is ~156 days (low confidence: a median of only 2 intervals). The last change was on 2026-08-05, 0 days before the as-of date (2026-08-05) — about 0.0× the typical interval. That is within its historical cadence.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

11 evidence refs4 distinct sources2 independent
  • CF5 Labs CASI/ARS leaderboardindependent eval×1 reference
  • CGray Swan agent red-teaming arenaindependent eval×1 reference
  • BAnthropic provider artifactsprovider artifact×8 references
  • EOpenRouter model rankingsusage data×1 reference

retrieved 2026-08-05

Compliance & deployment

Trains on customer data by default
No
SOC 2
Yes
ISO/IEC 42001
Yes
HIPAA eligible
Yes
Retention window
API: deleted within ~30 days; ZDR agreements available; trust-and-safety-flagged content up to 2 years
Data residency
Regional endpoints via Bedrock (global+regional) and Vertex (global/multi-region/regional); first-party API residency unstated
EU AI Act
Signed the EU GPAI Code of Practice (July 2025).
Deprecation policy
Documented lifecycle with 60-day minimum retirement notice; deprecation-preservation commitments published
Available via
Anthropic API · AWS Bedrock · Google Vertex AI · Microsoft Foundry (Azure, preview Nov 2025)

Incident history

2025-11-13
State-sponsored actor used Claude Code to automate intrusion campaign
Agentic misuse / social-engineering the model's safety context

Anthropic disclosed a campaign in which a state-sponsored group manipulated Claude Code into automating reconnaissance and exploitation against ~30 targets by role-playing as a legitimate security firm — an early, well-documented case of agentic-scale misuse of a frontier coding agent.

Outcome: Accounts banned, targets notified, public disclosure with TTPs. Cited here as evidence that agentic deployment risk is a property of the deployment pattern, not only the model.

Sources (1)
  • Anthropic disclosure
    Anthropic provider artifacts · provider artifact · source tier B · anthropic.com · retrieved 2026-08-03

Change timeline

2026-08-05
Verification pass: top-3 entries re-verified against live sources; usage shares corrected
methodologynotice

Every citation for GPT-5.1, Claude Sonnet 4.5, and Gemini 3 Pro was fetched and checked. Grades held; corrections were citations, migrated documentation domains (claude.com, developers.openai.com), and facts: ISO/IEC 42001 confirmed for OpenAI (previously unverified), lifecycle markers added (all three models are now legacy or retired), and seeded usage-share estimates replaced with measured OpenRouter daily data — under which every tracked model now sits below 1% share.

Evidence (2)
2026-06-30
Claude Sonnet 5 released; Sonnet 4.5 moves to Legacy
versioninfo

Anthropic released Claude Sonnet 5 (claude-sonnet-5) with a system card. Sonnet 4.5 is now listed under Legacy models with a tentative retirement floor of 2026-09-29 under the 60-day-notice policy.

Evidence (2)
2025-09-28
Anthropic consumer terms: training default switched to opt-in-by-default
policywarning

Consumer claude.ai accounts were transitioned to a training-permitted default with a five-year retention window (opt-out available). API and enterprise tiers unchanged. Widens the enterprise/consumer gap tracked under the data-handling vector.

Evidence (1)

Compare this model: Claude Sonnet 4.5 + open compare view →