Independent benchmarks, in context.

Explore what independent evaluations measure, which models they cover, and where the results apply.

Artificial Analysis

Artificial Analysis · video evaluations

Artificial Analysis · Text-to-video models. What it measures: Published evaluation protocol. Read the publisher’s evaluation method, task scope, and tested model version before applying its results to a tool. No tool score is inferred here.

Scope
Text-to-video models
What it measures
Published evaluation protocol

VBench team

VBench Leaderboard

VBench team · Text-to-video models. What it measures: Quality and consistency dimensions. A multi-dimensional video benchmark covering visual quality, motion and consistency.

Scope
Text-to-video models
What it measures
Quality and consistency dimensions

Artificial Analysis

Artificial Analysis · speech evaluations

Artificial Analysis · Text-to-speech providers. What it measures: Published evaluation protocol. Read the publisher’s evaluation method, task scope, and tested model version before applying its results to a tool. No tool score is inferred here.

Scope
Text-to-speech providers
What it measures
Published evaluation protocol

TTS-AGI & contributors

Speech synthesis evaluation sources

TTS-AGI & contributors · Text-to-speech models. What it measures: Published evaluation protocol. Read the publisher’s evaluation method, task scope, and tested model version before applying its results to a tool. No tool score is inferred here.

Scope
Text-to-speech models
What it measures
Published evaluation protocol

BigCode project

BigCodeBench

BigCode project · Code-generating models, not editors. What it measures: Practical coding task completion. Measures model performance on multi-library programming tasks. It is useful context, but it does not rank Cursor, Windsurf or Bolt as products.

Scope
Code-generating models, not editors
What it measures
Practical coding task completion

Artificial Analysis

Artificial Analysis · image evaluations

Artificial Analysis · Text-to-image models. What it measures: Published evaluation protocol. Read the publisher’s evaluation method, task scope, and tested model version before applying its results to a tool. No tool score is inferred here.

Scope
Text-to-image models
What it measures
Published evaluation protocol

Hugging Face & contributors

Open ASR Leaderboard

Hugging Face & contributors · Speech-recognition models only. What it measures: Transcription accuracy and speed. Plots word error rate against speed across ASR models. It does not measure meeting summaries, action items or integrations.

Scope
Speech-recognition models only
What it measures
Transcription accuracy and speed

Datacurve

DeepSWE

Datacurve · Coding agents run through mini-swe-agent. What it measures: Long-horizon engineering task success. Measures frontier coding agents on 113 original, long-horizon engineering tasks across 91 repositories and five languages.

Scope
Coding agents run through mini-swe-agent
What it measures
Long-horizon engineering task success