ModelRiskIndex

Rankings / Moonshot AI

Kimi K2

Tier 150/100draft — pending re-verificationopen weightslegacy

Kimi-K2-Instruct (open weights, self-hosted reference)

legacy. Open K2 checkpoints remain available; Moonshot usage and development have moved to K3, which launched hosted-first under a new custom license. Successor: Kimi K3 (released 2026-07-15, tracked separately).

Usage share 0.05% · OpenRouter rankings API (daily token share, 2026-08-04)

Tier assessment
Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.
  • Published safety evals. Safety evaluations for this model are published.Basic safety coverage in the technical report; not a dedicated safety eval suite.
  • Documented safety policy. A documented safety, usage, or acceptable-use policy governs the model.
  • Enterprise data controls. Not applicable to this artifact.
Tier 2 requirements
  • External pre-deployment testing. No disclosed external pre-deployment testing.
  • Third-party certification. No verifiable third-party certification.No certification pathway currently exists for open weights — an open methodology question, not a gap unique to this provider.
  • Versioning with changelogs. Model versions are explicitly identified and changes are changelogged.
  • Stated deprecation policy. A deprecation policy with notice windows is published.Weights remain available once released.

Missing for Tier 2: external pre-deployment testing, third-party certification. The tier is computed from this checklist — satisfying these requirements moves the badge, automatically.

Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governancen/a — deployer

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

Deployer property: open weights, self-hosted reference. Moonshot's hosted platform.moonshot.ai has separate PRC-jurisdiction terms not graded here.

Receipts (1)
  • Kimi K2 weights and license
    Moonshot AI provider artifacts · provider artifact · source tier B · huggingface.co · retrieved 2026-08-03
    Both the code repository and the model weights are released under the Modified MIT License

Operational stabilitystrong

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Immutable published checkpoints; versioned releases (K2, K2 Thinking) are explicit.

Receipts (1)
  • Kimi K2 weights and license
    Moonshot AI provider artifacts · provider artifact · source tier B · huggingface.co · retrieved 2026-08-03

Adversarial resistanceweak

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistanceweak

Bare weights jailbreak readily in independent testing; safety behavior is minimal relative to closed frontier peers.

Prompt injection (agentic)weak

Marketed for agentic use, which raises the stakes, but no published injection hardening or agent red-teaming results.

Receipts (2)
  • F5 Labs CASI leaderboard
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-03
  • Kimi K2 technical report
    Moonshot AI provider artifacts · provider artifact · source tier B · moonshotai.github.io · retrieved 2026-08-03

Transparencypartial

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Technical report and model card published; no safety evals or external testing.

Receipts (1)
  • Kimi K2 technical report
    Moonshot AI provider artifacts · provider artifact · source tier B · moonshotai.github.io · retrieved 2026-08-03

Compliance posturen/a — deployer

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

Deployer property: no certifications attach to the weights; deployer inherits all obligations.

Receipts (1)
  • Kimi K2 weights and license
    Moonshot AI provider artifacts · provider artifact · source tier B · huggingface.co · retrieved 2026-08-03

Governance & evidence

Where your data goes

  • SELF

No regional pinning — the provider chooses where data is processed.

Self-hosted reference deployment; Moonshot's hosted platform.moonshot.ai carries separate PRC-jurisdiction terms not graded here.

Enterprise vs consumer terms

Enterprise vs consumer gap: none measured. Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against. No consumer tier exists for the self-hosted reference deployment; data handling is deployer-controlled.CONSUMERENTERPRISE / APIWORSE TERMS →
No measured gap

No consumer tier exists for the self-hosted reference deployment; data handling is deployer-controlled.

Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against.

Change cadence

insufficient history1 tracked change · last 2026-08-04 · 1d since Insufficient history to estimate a cadence.1d

One tracked change, on 2026-08-04 — 1 days before the as-of date (2026-08-05). A single event cannot establish a cadence, so days-since is shown without a baseline.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

7 evidence refs3 distinct sources1 independent
  • CF5 Labs CASI/ARS leaderboardindependent eval×1 reference
  • BMoonshot AI provider artifactsprovider artifact×5 references
  • EOpenRouter model rankingsusage data×1 reference

retrieved 2026-08-03 — 2026-08-05

Compliance & deployment

Trains on customer data by default
No
SOC 2
not verified
ISO/IEC 42001
not verified
HIPAA eligible
not verified
Retention window
Deployer-controlled
Data residency
Deployer-controlled
EU AI Act
Open-weight GPAI provisions apply; deployer obligations dominate.
Deprecation policy
Weights remain available once released
Available via
Self-hosted · Multiple inference providers · platform.moonshot.ai (hosted, separate terms)

Change timeline

2026-08-04
Methodology v0.2: computed tiers and first-class N/A grades
methodologynotice

Tiers are now derived from a per-requirement checklist rather than assigned, and not-applicable / under-review became first-class grade states excluded from the composite score. Deployer-property vectors on open-weight models (data handling, compliance) moved from graded to not-applicable, changing the composite scores of the tagged models.

Evidence (1)

Compare this model: Kimi K2 + open compare view →