ModelRiskIndex

Rankings / Google

Gemini 3.1 Pro

Tier 270/100

gemini-3.1-pro-preview via generativelanguage.googleapis.com

Usage share 0.28% · OpenRouter rankings API (daily token share, 2026-08-04)

Google's current flagship as of Aug 2026 (no 3.5 Pro exists; newer releases are Flash-tier). Succeeded Gemini 3 Pro via the alias repoint recorded in the change feed.

Tier assessment
Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.
  • Published safety evals. Safety evaluations for this model are published.
  • Documented safety policy. A documented safety, usage, or acceptable-use policy governs the model.
  • Enterprise data controls. Customer data is not used for training by default, or a documented opt-out exists.
Tier 2 requirements
  • External pre-deployment testing. Independent external parties tested the model before deployment, and this is disclosed.Frontier Safety Framework evaluation is disclosed in the model card; the archived card does not name external assessors, so the external-testing element rests on Google's release materials rather than the card itself.
  • Third-party certification. The operating organization holds verifiable third-party certification (e.g. SOC 2, ISO/IEC 42001).
  • Versioning with changelogs. Model versions are explicitly identified and changes are changelogged.Changelog exists, but the preview-only ID and the inherited silent alias repoint weaken version stability in practice.
  • Stated deprecation policy. A deprecation policy with notice windows is published.
Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governancepartial

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

Same Gemini API posture verified for the 3 Pro entry: paid tier not used for improvement (55-day abuse logs, ZDR by approval), unpaid tier used for improvement including human review, consumer Gemini Apps default to data use.

Receipts (2)
  • Gemini API terms of service
    Google / DeepMind provider artifacts · provider artifact · source tier B · ai.google.dev · retrieved 2026-08-05
    To help with quality and improve our products, human reviewers may read, annotate, and process your API input and output.
  • Gemini API zero-data-retention documentation
    Google / DeepMind provider artifacts · provider artifact · source tier B · ai.google.dev · retrieved 2026-08-05
    When your request for ZDR for a particular project is approved, all user content (prompts and responses) and identifiable metadata (such as IP addresses and Google Account IDs) are cleared prior to logging.

Operational stabilitypartial

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Still a preview-suffixed ID nearly six months after release, and it inherited traffic via a silent alias repoint from its predecessor. Vertex stable-version lifecycle discipline exists but has yet to apply to this flagship line.

Receipts (2)
  • Gemini API changelog
    Google / DeepMind provider artifacts · provider artifact · source tier B · ai.google.dev · retrieved 2026-08-05
    The gemini-3-pro-preview now points to gemini-3.1-pro-preview
  • Vertex AI model versions and lifecycle
    Google / DeepMind provider artifacts · provider artifact · source tier B · docs.cloud.google.com · retrieved 2026-08-05
    The following tables list the available models and their retirement dates.

Adversarial resistancepartial

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistancepartial

Own model card with full Frontier Safety Framework evaluation across five risk domains; but F5's monthly CASI board places the Gemini 3.x line mid-to-low tier and strikingly volatile — the Gemini 3 Pro preview moved 55.7 CASI points across a four-month period, against a top-10 threshold near 86.

Prompt injection (agentic)partial

Inherits Google's layered defenses (model hardening against indirect injection, agentic security architecture); researcher demonstrations against Gemini-connected surfaces continue to land.

Receipts (5)
  • Gemini 3.1 Pro model card
    Google / DeepMind provider artifacts · provider artifact · source tier B · deepmind.google · retrieved 2026-08-05
    Our Frontier Safety Framework includes rigorous evaluations that address risks of severe harm from frontier models, covering five risk domains: CBRN (chemical, biological, radiological and nuclear information risks), cyber, harmful manipulation, machine learning R&D and misalignment.
  • F5 Labs CASI/ARS leaderboards
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-05
  • F5 Labs: The Small Model Cliff (Gemini volatility analysis)
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-05
    Gemini 3 Pro preview moved 55.7 points across the period, ending close to the volatile cohort.
  • Advancing Gemini's security safeguards (DeepMind)
    Google / DeepMind provider artifacts · provider artifact · source tier B · deepmind.google · retrieved 2026-08-05
    This model hardening has significantly boosted Gemini’s ability to identify and ignore injected instructions, lowering its attack success rate.
  • Lessons from Defending Gemini Against Indirect Prompt Injections (white paper)
    Google / DeepMind provider artifacts · provider artifact · source tier B · storage.googleapis.com · retrieved 2026-08-05
    However, such adversarial training will not render the model immune to indirect prompt injection and successful attacks remain possible, particularly with increased adversarial effort, novel techniques, or highly tailored exploits.

Transparencystrong

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Dedicated model card published at release (2026-02-19) with a full Frontier Safety Framework evaluation; Google's external-testing arrangements (UK AISI early access, independent assessors) carried forward from Gemini 3.

Receipts (1)
  • Gemini 3.1 Pro model card
    Google / DeepMind provider artifacts · provider artifact · source tier B · deepmind.google · retrieved 2026-08-05
    Our Frontier Safety Framework includes rigorous evaluations that address risks of severe harm from frontier models, covering five risk domains: CBRN (chemical, biological, radiological and nuclear information risks), cyber, harmful manipulation, machine learning R&D and misalignment.

Compliance posturestrong

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

Google Cloud's accredited ISO/IEC 42001:2023 certification covers Vertex generative AI, alongside SOC 1/2/3, ISO 27001 family, and HIPAA support.

Receipts (1)
  • Google Cloud ISO/IEC 42001 certification
    Google / DeepMind provider artifacts · provider artifact · source tier B · cloud.google.com · retrieved 2026-08-05
    Google Cloud Platform, Google Workspace, and Gemini (App) are certified as ISO/IEC 42001:2023 compliant.

Governance & evidence

Where your data goes

  • US
  • EU

Regional pinning available — customers can pin processing to a chosen region.

Vertex regional/jurisdictional endpoints (US, EU; expanding) with in-region ML processing; the global endpoint carries no residency guarantee.

Enterprise vs consumer terms

Enterprise vs consumer gap: wide. Paid API tier is not used for improvement, but the unpaid tier is (including human review) and consumer Gemini Apps default to data use.CONSUMERENTERPRISE / APIWORSE TERMS →
Wide gap

Paid API tier is not used for improvement, but the unpaid tier is (including human review) and consumer Gemini Apps default to data use.

Change cadence

insufficient history1 tracked change · last 2026-08-05 · 0d since Insufficient history to estimate a cadence.0d

One tracked change, on 2026-08-05 — 0 days before the as-of date (2026-08-05). A single event cannot establish a cadence, so days-since is shown without a baseline.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

12 evidence refs3 distinct sources1 independent
  • CF5 Labs CASI/ARS leaderboardindependent eval×2 references
  • BGoogle / DeepMind provider artifactsprovider artifact×9 references
  • EOpenRouter model rankingsusage data×1 reference

retrieved 2026-08-05

Compliance & deployment

Trains on customer data by default
No
SOC 2
Yes
ISO/IEC 42001
Yes
HIPAA eligible
Yes
Retention window
Paid tier: 55-day abuse-monitoring logs; ZDR by approval; unpaid tier: content used for improvement incl. human review
Data residency
Vertex regional/jurisdictional endpoints (US, EU; expanding) with in-region ML processing; global endpoint carries no residency guarantee
EU AI Act
Signed the EU GPAI Code of Practice (2025-07-30).
Deprecation policy
Published Vertex model lifecycle with retirement dates; preview-tier IDs excluded in practice
Available via
Gemini API · Google Vertex AI

Change timeline

2026-08-05
Correction: two of our own claims did not survive verification
methodologywarning

Binding claims to verbatim source text caught two errors in our published data. (1) We stated NVIDIA's hosted API trial carries a no-training clause; the Trial Terms in fact reserve use of submitted content 'to improve NVIDIA products and services, including AI models' — the opposite. The data-governance summary and retention field are corrected and the clause is now quoted directly. (2) Our Gemini transparency entries credited the model card with naming external testers (UK AISI, Apollo, Vaultis, Dreadnode); the archived card names none of them, so that requirement now rests on Google's release materials and is flagged as needing a better primary citation. Published rather than silently fixed, per the methodology.

Evidence (1)
  • NVIDIA API Trial Terms of Service (§3.3)
    NVIDIA Nemotron provider artifacts · provider artifact · source tier B · assets.ngc.nvidia.com · retrieved 2026-08-05
    3.3 NVIDIA will collect the following data, without identifying specific users, to operate and improve the API Services and other products and services: (i) session metrics (e.g., the amount of processing power consumed, type of request made); (ii) error logs and execution logs relating to your session (e.g., whether your request was executed successfully); (iii) your feedback and ranking of specific API Services; and (iv) User Content and Generated Content to improve NVIDIA products and services, including AI models.

Compare this model: Gemini 3.1 Pro + open compare view →