The heartbeat
What changed, and when we knew
A reverse-chronological record of provider policy changes, model card publications, version events, and incidents across tracked models. Phase 1: curated weekly by hand. Phase 2 adds automated policy diffing, endpoint fingerprinting, and behavioral drift detection — with alerting on the models you pin.
Binding claims to verbatim source text caught two errors in our published data. (1) We stated NVIDIA's hosted API trial carries a no-training clause; the Trial Terms in fact reserve use of submitted content 'to improve NVIDIA products and services, including AI models' — the opposite. The data-governance summary and retention field are corrected and the clause is now quoted directly. (2) Our Gemini transparency entries credited the model card with naming external testers (UK AISI, Apollo, Vaultis, Dreadnode); the archived card names none of them, so that requirement now rests on Google's release materials and is flagged as needing a better primary citation. Published rather than silently fixed, per the methodology.
Evidence (1)
- NVIDIA API Trial Terms of Service (§3.3)
“3.3 NVIDIA will collect the following data, without identifying specific users, to operate and improve the API Services and other products and services: (i) session metrics (e.g., the amount of processing power consumed, type of request made); (ii) error logs and execution logs relating to your session (e.g., whether your request was executed successfully); (iii) your feedback and ranking of specific API Services; and (iv) User Content and Generated Content to improve NVIDIA products and services, including AI models.”
Frontier-lab system cards, model cards and technical reports are published as PDFs, which the archival monitor could hash but not read — so every claim citing one sat permanently unverified. A pypdf-based extraction pipeline now archives their text through the same internal ingest API the monitor uses, making the Claude Opus 5, Claude Sonnet 5, Claude Sonnet 4.5, GPT-5, Gemini, Grok 4.1, Nemotron and arXiv PDFs quotable and checkable for the first time.
Evidence (1)
Evidence references can now carry the exact sentence supporting a claim, verified against the archived snapshot of the cited page and re-checked after every monitor run. A provider editing the sentence a grade rests on now flags that claim as invalidated instead of leaving a live-but-hollow link. Coverage is published openly at /api/v1/verification and is being backfilled — an unbound claim is unverified, not wrong.
Evidence (1)
Jailbreak and prompt-injection resistance now present as one Adversarial resistance vector, graded to the weaker facet, with both measures still shown on each model page. Composite scores are now computed over five vectors, and the risk rosette is a five-slot wheel. No underlying assessment changed — this is a presentation and scoring change, published like any other.
Evidence (1)
Thirteen entries added with fetched evidence: DeepSeek V4 Flash and Pro, Xiaomi MiMo v2.5, Tencent HY3, GPT-5.6 (Sol/Terra/Luna), GLM-5.2, MiniMax M3, Nemotron 3 Ultra, Step 3.7 Flash, Kimi K3, Claude Sonnet 5, Claude Opus 5, and Gemini 3.1 Pro. This closes the coverage gap flagged on 2026-08-05: the current OpenRouter top 10 is now fully assessed. Five of the thirteen enter at Tier 0; two carry under-review grades where no security evidence exists in either direction.
Evidence (1)
Every citation for GPT-5.1, Claude Sonnet 4.5, and Gemini 3 Pro was fetched and checked. Grades held; corrections were citations, migrated documentation domains (claude.com, developers.openai.com), and facts: ISO/IEC 42001 confirmed for OpenAI (previously unverified), lifecycle markers added (all three models are now legacy or retired), and seeded usage-share estimates replaced with measured OpenRouter daily data — under which every tracked model now sits below 1% share.
Measured rankings show ~60% of OpenRouter token volume flowing through models this index does not yet cover — led by DeepSeek V4 Flash (~12%), Xiaomi MiMo v2.5, Tencent HY3, GPT-5.6 Luna, and GLM-5.2 — plus successors to tracked models (Claude Sonnet/Opus 5, Kimi K3, Gemini 3.1 Pro). Tracked-set expansion is planned; publishing the gap is preferable to implying coverage that does not exist.
Tiers are now derived from a per-requirement checklist rather than assigned, and not-applicable / under-review became first-class grade states excluded from the composite score. Deployer-property vectors on open-weight models (data handling, compliance) moved from graded to not-applicable, changing the composite scores of the tagged models.
Evidence (1)
Initial publication of the six-vector scoring framework, tier definitions, and evidence requirements. All initial grades are Phase 1 aggregation: sourced from public data, not first-party probes.
Evidence (1)
GPAI transparency and copyright obligations have applied since August 2025; high-risk system obligations continue phasing in. Deployer documentation duties are the driver for most enterprise buyers tracked here.
Evidence (1)
DeepSeek replaced the deepseek-v4-flash alias with a retrained build (0731) — 'the calling method remains unchanged' per their own release note — with drastically different agentic behavior (it now beats V4 Pro on all nine agent benchmarks). The dated open weights were published separately; API users were switched without action. The largest-share model on OpenRouter is also the clearest current example of the silent-swap failure mode.
Evidence (1)
System card names UK AISI, Trajectory Labs, 10a Labs, and Gray Swan; includes an adverse capability finding (agentic cyber-range success against weakly-secured networks) published against interest. Day-one availability on Bedrock, Vertex, and Foundry. The Opus 4.5 entry moves to legacy.
Evidence (1)
Multiple developer reports of changed refusal behavior on the floating gpt-5.1 alias while dated snapshots remained stable. No corresponding entry on the deprecations/changelog pages. Illustrates the alias-vs-snapshot distinction this index tracks.
Evidence (1)
- OpenAI developer community threadCurated manually; endpoint fingerprinting (Phase 2) will verify future occurrences.
GPT-5.6 (Sol/Terra/Luna) reached general availability as OpenAI's flagship line. Per the deprecations page: gpt-5.1-codex and codex-max shut down 2026-07-23, gpt-5.1-chat-latest shuts down 2026-08-10, and GPT-5 base snapshots retire 2026-12-11. The dated gpt-5.1-2025-11-13 snapshot remains live, for now.
Evidence (1)
After the blanket preservation order was narrowed in Sept 2025 and a court affirmed production of 20M de-identified chat logs in Jan 2026, NYT and Daily News filed a sanctions motion alleging OpenAI deleted logs subject to preservation. API zero-data-retention customers remain excluded, but buyers relying on OpenAI's retention story should track the case.
Evidence (1)
Official release relicensed the weights from the preview's community license to Apache 2.0, and a two-week free OpenRouter tier pushed HY3 to #1 by weekly usage (6.13T tokens) — a ~8.6% share model with, at time of entry, zero independent security testing in either direction.
Evidence (1)
Anthropic released Claude Sonnet 5 (claude-sonnet-5) with a system card. Sonnet 4.5 is now listed under Legacy models with a tentative retirement floor of 2026-09-29 under the 60-day-notice policy.
Under a voluntary pre-release review process created by a June 2 executive order, the White House asked OpenAI to restrict the GPT-5.6 preview to roughly twenty government-vetted US partners, citing Sol's cyber capabilities. Broad rollout cleared on July 9. The first instance of US government pre-release gating shaping a frontier model launch.
Roughly four months after launch and without reaching a stable GA ID, Google shut down gemini-3-pro-preview and repointed the ID to gemini-3.1-pro-preview. Requests to the old ID now reach a different model — the canonical silent-version-swap failure mode this index tracks.
Evidence (1)
Release accompanied by a detailed system card including third-party evaluation results — the disclosure pattern Tier 2 requires.
Evidence (1)
Flagship release with published model card and Frontier Safety Framework coverage.
Evidence (1)
First formal model card from xAI, alongside a risk-management framework. Moves Grok from Tier 0 to Tier 1 under this index's definitions.
Evidence (1)
Open weights published with a technical report. The hosted deepseek-chat alias was cut over to the new version in place — users of the first-party API changed models without an account-level action.
Evidence (1)
Consumer claude.ai accounts were transitioned to a training-permitted default with a five-year retention window (opt-out available). API and enterprise tiers unchanged. Widens the enterprise/consumer gap tracked under the data-handling vector.