ModelRiskIndex

Rankings / StepFun

StepFun Step 3.7 Flash

Tier 00/100open weights

step-3.7-flash (open weights, Apache 2.0) via platform.stepfun.ai / multiple hosts

Usage share 2.7% · OpenRouter rankings API (daily token share, 2026-08-04)

Composite score 0/100 over the four applied vectors — not because the model is proven unsafe, but because nothing about its safety posture is documented or tested. The distinction matters and both halves are on the record here.

Tier assessment

Fails one or more Tier 1 requirements: no published model card or safety evals, or terms that permit training on customer API data by default with no opt-out.

Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.Capability-only card; no safety content.
  • Published safety evals. No safety evaluations are published for this model.
  • Documented safety policy. No documented safety or acceptable-use policy exists.
  • Enterprise data controls. Terms permit training on customer data by default with no documented opt-out.Privacy policy silent on training use of API data; no opt-out documented.
Tier 2 requirements
  • External pre-deployment testing. No disclosed external pre-deployment testing.
  • Third-party certification. No verifiable third-party certification.
  • Versioning with changelogs. No versioning discipline or changelog for model changes.
  • Stated deprecation policy. No stated deprecation policy.

Missing for Tier 1: published safety evals, documented safety policy, enterprise data controls. The tier is computed from this checklist — satisfying these requirements moves the badge, automatically.

Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governanceweak

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

The global platform's privacy policy states US-located servers (a genuine plus) and 3-year IP retention — but is entirely silent on whether API inputs and outputs are used for training, with no opt-out documented and the terms-of-use page returning 404. Silence on training is a fail for a buyer. Self-hosting the Apache 2.0 weights is deployer-controlled.

Receipts (1)

Operational stabilityweak

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Undated alias on the first-party API with no snapshot scheme (OpenRouter's date suffix is OpenRouter's, not StepFun's), no changelog, and no deprecation policy.

Receipts (1)

Adversarial resistanceunder review

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistanceunder review

A complete testing vacuum: no F5 listing, no academic or third-party security testing found, and no developer safety evals. For a top-10 model by usage, the absence is the finding.

Prompt injection (agentic)under review

Marketed for coding agents and search workflows; no injection results or hardening documentation exist from StepFun or any third party.

Receipts (3)

Transparencyweak

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Open weights, but the thinnest documentation of any tracked model: a capability-only model card, a blog post standing in for a technical report, no safety content of any kind, and no external testing.

Receipts (1)
  • Step 3.7 Flash weights and model card
    StepFun provider artifacts · provider artifact · source tier B · huggingface.co · retrieved 2026-08-05
    This project is open-sourced under the Apache 2.0 License

Compliance postureweak

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

No certifications mentioned anywhere in platform documentation.

Receipts (1)

Governance & evidence

Where your data goes

  • US

No regional pinning — the provider chooses where data is processed.

Global platform states US-located servers; the separate China platform's terms are not assessed.

Enterprise vs consumer terms

Enterprise vs consumer gap: not assessed. Not yet assessed. Absence of assessment is not absence of a gap.NOT ASSESSEDENTERPRISE / APICONSUMER
Not assessed

Not yet assessed. Absence of assessment is not absence of a gap.

Change cadence

no tracked changesNo tracked changes for StepFun Step 3.7 Flash. Absence of detection is not evidence of stability.none

No tracked changes for this model. That is absence of detection, not evidence of stability — it may mean the model is unwatched, not that it is unchanging.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

8 evidence refs3 distinct sources1 independent
  • CF5 Labs CASI/ARS leaderboardindependent eval×1 reference
  • BStepFun provider artifactsprovider artifact×6 references
  • EOpenRouter model rankingsusage data×1 reference

retrieved 2026-08-05

Compliance & deployment

Trains on customer data by default
not verified
SOC 2
not verified
ISO/IEC 42001
not verified
HIPAA eligible
not verified
Retention window
Not documented for prompts; IP addresses retained 3 years
Data residency
US servers (global platform); China platform terms not assessed
EU AI Act
No published EU AI Act posture.
Deprecation policy
none stated
Available via
platform.stepfun.ai · Self-hosted (Apache 2.0 weights) · NVIDIA NIM · Baseten · Multiple inference providers

Change timeline

No tracked changes yet for this model.

Compare this model: StepFun Step 3.7 Flash + open compare view →