Independent benchmarks, in context.
Explore what independent evaluations measure, which models they cover, and where the results apply.
Artificial Analysis
Artificial Analysis · Text-to-video models. What it measures: Published evaluation protocol. Read the publisher’s evaluation method, task scope, and tested model version before applying its results to a tool. No tool score is inferred here.
- Scope
- Text-to-video models
- What it measures
- Published evaluation protocol
VBench team
VBench team · Text-to-video models. What it measures: Quality and consistency dimensions. A multi-dimensional video benchmark covering visual quality, motion and consistency.
- Scope
- Text-to-video models
- What it measures
- Quality and consistency dimensions
Artificial Analysis
Artificial Analysis · Text-to-speech providers. What it measures: Published evaluation protocol. Read the publisher’s evaluation method, task scope, and tested model version before applying its results to a tool. No tool score is inferred here.
- Scope
- Text-to-speech providers
- What it measures
- Published evaluation protocol
TTS-AGI & contributors
TTS-AGI & contributors · Text-to-speech models. What it measures: Published evaluation protocol. Read the publisher’s evaluation method, task scope, and tested model version before applying its results to a tool. No tool score is inferred here.
- Scope
- Text-to-speech models
- What it measures
- Published evaluation protocol
BigCode project
BigCode project · Code-generating models, not editors. What it measures: Practical coding task completion. Measures model performance on multi-library programming tasks. It is useful context, but it does not rank Cursor, Windsurf or Bolt as products.
- Scope
- Code-generating models, not editors
- What it measures
- Practical coding task completion
Artificial Analysis
Artificial Analysis · Text-to-image models. What it measures: Published evaluation protocol. Read the publisher’s evaluation method, task scope, and tested model version before applying its results to a tool. No tool score is inferred here.
- Scope
- Text-to-image models
- What it measures
- Published evaluation protocol
Hugging Face & contributors
Hugging Face & contributors · Speech-recognition models only. What it measures: Transcription accuracy and speed. Plots word error rate against speed across ASR models. It does not measure meeting summaries, action items or integrations.
- Scope
- Speech-recognition models only
- What it measures
- Transcription accuracy and speed
Datacurve
Datacurve · Coding agents run through mini-swe-agent. What it measures: Long-horizon engineering task success. Measures frontier coding agents on 113 original, long-horizon engineering tasks across 91 repositories and five languages.
- Scope
- Coding agents run through mini-swe-agent
- What it measures
- Long-horizon engineering task success