ModelRiskIndex

The heartbeat

What changed, and when we knew

A reverse-chronological record of provider policy changes, model card publications, version events, and incidents across tracked models. Phase 1: curated weekly by hand. Phase 2 adds automated policy diffing, endpoint fingerprinting, and behavioral drift detection — with alerting on the models you pin.

2026-08-05
Correction: two of our own claims did not survive verification

Binding claims to verbatim source text caught two errors in our published data. (1) We stated NVIDIA's hosted API trial carries a no-training clause; the Trial Terms in fact reserve use of submitted content 'to improve NVIDIA products and services, including AI models' — the opposite. The data-governance summary and retention field are corrected and the clause is now quoted directly. (2) Our Gemini transparency entries credited the model card with naming external testers (UK AISI, Apollo, Vaultis, Dreadnode); the archived card names none of them, so that requirement now rests on Google's release materials and is flagged as needing a better primary citation. Published rather than silently fixed, per the methodology.

Evidence (1)
  • NVIDIA API Trial Terms of Service (§3.3)
    NVIDIA Nemotron provider artifacts · provider artifact · source tier B · assets.ngc.nvidia.com · retrieved 2026-08-05
    3.3 NVIDIA will collect the following data, without identifying specific users, to operate and improve the API Services and other products and services: (i) session metrics (e.g., the amount of processing power consumed, type of request made); (ii) error logs and execution logs relating to your session (e.g., whether your request was executed successfully); (iii) your feedback and ranking of specific API Services; and (iv) User Content and Generated Content to improve NVIDIA products and services, including AI models.
2026-08-05
PDF extraction pipeline: system cards are now verifiable evidence
methodologyinfoecosystem

Frontier-lab system cards, model cards and technical reports are published as PDFs, which the archival monitor could hash but not read — so every claim citing one sat permanently unverified. A pypdf-based extraction pipeline now archives their text through the same internal ingest API the monitor uses, making the Claude Opus 5, Claude Sonnet 5, Claude Sonnet 4.5, GPT-5, Gemini, Grok 4.1, Nemotron and arXiv PDFs quotable and checkable for the first time.

Evidence (1)
2026-08-05
Quote-level provenance: claims are now bound to verbatim source text
methodologynoticeecosystem

Evidence references can now carry the exact sentence supporting a claim, verified against the archived snapshot of the cited page and re-checked after every monitor run. A provider editing the sentence a grade rests on now flags that claim as invalidated instead of leaving a live-but-hollow link. Coverage is published openly at /api/v1/verification and is being backfilled — an unbound claim is unverified, not wrong.

Evidence (1)
2026-08-05
Methodology v0.3: framework collapsed from six vectors to five
methodologynoticeecosystem

Jailbreak and prompt-injection resistance now present as one Adversarial resistance vector, graded to the weaker facet, with both measures still shown on each model page. Composite scores are now computed over five vectors, and the risk rosette is a five-slot wheel. No underlying assessment changed — this is a presentation and scoring change, published like any other.

Evidence (1)
2026-08-05
Tracked set expanded from 12 to 25 models; OpenRouter top-10 coverage complete
methodologynoticeecosystem

Thirteen entries added with fetched evidence: DeepSeek V4 Flash and Pro, Xiaomi MiMo v2.5, Tencent HY3, GPT-5.6 (Sol/Terra/Luna), GLM-5.2, MiniMax M3, Nemotron 3 Ultra, Step 3.7 Flash, Kimi K3, Claude Sonnet 5, Claude Opus 5, and Gemini 3.1 Pro. This closes the coverage gap flagged on 2026-08-05: the current OpenRouter top 10 is now fully assessed. Five of the thirteen enter at Tier 0; two carry under-review grades where no security evidence exists in either direction.

Evidence (1)
2026-08-05
Verification pass: top-3 entries re-verified against live sources; usage shares corrected

Every citation for GPT-5.1, Claude Sonnet 4.5, and Gemini 3 Pro was fetched and checked. Grades held; corrections were citations, migrated documentation domains (claude.com, developers.openai.com), and facts: ISO/IEC 42001 confirmed for OpenAI (previously unverified), lifecycle markers added (all three models are now legacy or retired), and seeded usage-share estimates replaced with measured OpenRouter daily data — under which every tracked model now sits below 1% share.

Evidence (2)
2026-08-05
Coverage gap: the entire OpenRouter top 10 is currently untracked
methodologywarningecosystem

Measured rankings show ~60% of OpenRouter token volume flowing through models this index does not yet cover — led by DeepSeek V4 Flash (~12%), Xiaomi MiMo v2.5, Tencent HY3, GPT-5.6 Luna, and GLM-5.2 — plus successors to tracked models (Claude Sonnet/Opus 5, Kimi K3, Gemini 3.1 Pro). Tracked-set expansion is planned; publishing the gap is preferable to implying coverage that does not exist.

Evidence (1)
2026-08-04
Methodology v0.2: computed tiers and first-class N/A grades

Tiers are now derived from a per-requirement checklist rather than assigned, and not-applicable / under-review became first-class grade states excluded from the composite score. Deployer-property vectors on open-weight models (data handling, compliance) moved from graded to not-applicable, changing the composite scores of the tagged models.

Evidence (1)
2026-08-04
ModelRiskIndex methodology v0.1 published
methodologyinfoecosystem

Initial publication of the six-vector scoring framework, tier definitions, and evidence requirements. All initial grades are Phase 1 aggregation: sourced from public data, not first-party probes.

Evidence (1)
  • Methodology page
    ModelRiskIndex methodology · first-party · source tier A · modelriskindex.com · retrieved 2026-08-04
2026-08-02
EU AI Act GPAI obligations: one year in force
regulationnoticeecosystem

GPAI transparency and copyright obligations have applied since August 2025; high-risk system obligations continue phasing in. Deployer documentation duties are the driver for most enterprise buyers tracked here.

Evidence (1)
2026-07-31
DeepSeek V4 Flash retrained and swapped into the live alias in place
versionwarningDeepSeek V4 Flash

DeepSeek replaced the deepseek-v4-flash alias with a retrained build (0731) — 'the calling method remains unchanged' per their own release note — with drastically different agentic behavior (it now beats V4 Pro on all nine agent benchmarks). The dated open weights were published separately; API users were switched without action. The largest-share model on OpenRouter is also the clearest current example of the silent-swap failure mode.

Evidence (1)
2026-07-24
Claude Opus 5 released under ASL-3 with four named external testers

System card names UK AISI, Trajectory Labs, 10a Labs, and Gray Swan; includes an adverse capability finding (agentic cyber-range success against weakly-secured networks) published against interest. Day-one availability on Bedrock, Vertex, and Foundry. The Opus 4.5 entry moves to legacy.

Evidence (1)
  • Claude Opus 5 announcement
    Anthropic provider artifacts · provider artifact · source tier B · anthropic.com · retrieved 2026-08-05
2026-07-16
Developer reports of behavior shift on gpt-5.1 alias without changelog entry
versionnoticeGPT-5.1

Multiple developer reports of changed refusal behavior on the floating gpt-5.1 alias while dated snapshots remained stable. No corresponding entry on the deprecations/changelog pages. Illustrates the alias-vs-snapshot distinction this index tracks.

Evidence (1)
  • OpenAI developer community thread
    OpenAI developer community · reporting · source tier F · community.openai.com · retrieved 2026-07-16
    Curated manually; endpoint fingerprinting (Phase 2) will verify future occurrences.
2026-07-09
GPT-5.6 family GA; GPT-5.1 variant shutdowns scheduled and executed
versionwarningGPT-5.1

GPT-5.6 (Sol/Terra/Luna) reached general availability as OpenAI's flagship line. Per the deprecations page: gpt-5.1-codex and codex-max shut down 2026-07-23, gpt-5.1-chat-latest shuts down 2026-08-10, and GPT-5 base snapshots retire 2026-12-11. The dated gpt-5.1-2025-11-13 snapshot remains live, for now.

Evidence (1)
  • OpenAI API deprecations page
    OpenAI provider artifacts · provider artifact · source tier B · developers.openai.com · retrieved 2026-08-05
2026-07-09
NYT v. OpenAI: sanctions motion over log handling; retention promises remain litigation-entangled
policynoticeGPT-5.1

After the blanket preservation order was narrowed in Sept 2025 and a court affirmed production of 20M de-identified chat logs in Jan 2026, NYT and Daily News filed a sanctions motion alleging OpenAI deleted logs subject to preservation. API zero-data-retention customers remain excluded, but buyers relying on OpenAI's retention story should track the case.

Evidence (1)
2026-07-06
Tencent HY3 GA: license loosened to Apache 2.0, free tier drives #1 weekly usage
versionnoticeTencent Hunyuan HY3

Official release relicensed the weights from the preview's community license to Apache 2.0, and a two-week free OpenRouter tier pushed HY3 to #1 by weekly usage (6.13T tokens) — a ~8.6% share model with, at time of entry, zero independent security testing in either direction.

Evidence (1)
2026-06-30
Claude Sonnet 5 released; Sonnet 4.5 moves to Legacy

Anthropic released Claude Sonnet 5 (claude-sonnet-5) with a system card. Sonnet 4.5 is now listed under Legacy models with a tentative retirement floor of 2026-09-29 under the 60-day-notice policy.

Evidence (2)
2026-06-26
GPT-5.6 launch gated by White House request under June executive order

Under a voluntary pre-release review process created by a June 2 executive order, the White House asked OpenAI to restrict the GPT-5.6 preview to roughly twenty government-vetted US partners, citing Sol's cyber capabilities. Broad rollout cleared on July 9. The first instance of US government pre-release gating shaping a frontier model launch.

Evidence (2)
2026-03-09
gemini-3-pro-preview shut down; model ID silently aliased to Gemini 3.1 Pro
versionwarningGemini 3 Pro

Roughly four months after launch and without reaching a stable GA ID, Google shut down gemini-3-pro-preview and repointed the ID to gemini-3.1-pro-preview. Requests to the old ID now reach a different model — the canonical silent-version-swap failure mode this index tracks.

Evidence (1)
  • Gemini API changelog
    Google / DeepMind provider artifacts · provider artifact · source tier B · ai.google.dev · retrieved 2026-08-05
2025-11-24
Claude Opus 4.5 released with system card and external pre-deployment testing
model-cardinfoClaude Opus 4.5

Release accompanied by a detailed system card including third-party evaluation results — the disclosure pattern Tier 2 requires.

Evidence (1)
2025-11-18
Gemini 3 Pro launched with model card
versioninfoGemini 3 Pro

Flagship release with published model card and Frontier Safety Framework coverage.

Evidence (1)
  • Gemini 3 Pro model card
    Google / DeepMind provider artifacts · provider artifact · source tier B · storage.googleapis.com · retrieved 2026-08-03
2025-11-17
xAI publishes its first model card with Grok 4.1
model-cardinfoGrok 4.1

First formal model card from xAI, alongside a risk-management framework. Moves Grok from Tier 0 to Tier 1 under this index's definitions.

Evidence (1)
  • Grok 4.1 model card
    xAI provider artifacts · provider artifact · source tier B · data.x.ai · retrieved 2026-08-03
2025-09-29
DeepSeek V3.2 released; hosted alias updated in place
versionnoticeDeepSeek V3.2

Open weights published with a technical report. The hosted deepseek-chat alias was cut over to the new version in place — users of the first-party API changed models without an account-level action.

Evidence (1)
2025-09-28
Anthropic consumer terms: training default switched to opt-in-by-default

Consumer claude.ai accounts were transitioned to a training-permitted default with a five-year retention window (opt-out available). API and enterprise tiers unchanged. Widens the enterprise/consumer gap tracked under the data-handling vector.

Evidence (1)