AI & models
Benchmark
A standardized test used to compare models on a particular skill — coding, math, reading, reasoning — by running them all on the same set of questions and scoring the results. It turns 'which model is better?' into a number you can line up side by side, which is handy but never the whole story.
Why it matters
Part of When AI is wrong → Benchmarks are how to read 'state-of-the-art' claims with a clear head — they're a useful signal, but the real test is whether a model does well on your actual task.