What Is AI Benchmarking?
What those AI scores actually mean
What Is AI Benchmarking?
What those AI scores actually mean

Every time a new AI model is released, you’ll see scores like “92% on MMLU” or claims such as “Tops HumanEval”. They’re standardized tests designed to measure how well an AI model performs on specific tasks.
What’s a Benchmark?
A benchmark is a set of questions with known correct answers. You run the model through all of them, check how many it gets right, and report the percentage. A model scoring 86% on MMLU means it got 86 out of 100 questions correct.
These tests are created by researchers, mostly from universities and AI labs. MMLU came from UC Berkeley. HumanEval was made by OpenAI. Google, Meta, and independent research groups have all published their own. There’s no central authority. But over time, certain benchmarks become the accepted standard simply because everyone uses them. If you release a model and don’t report MMLU scores, people will question how good it actually is. So, it’s just convention, not a rule.
What Gets Tested
Each benchmark targets a specific skill:

Each benchmark only tests one thing. So companies never report just one score. They run their model through a dozen benchmarks and publish the full table:

Model A knows more facts. Model B writes better code and does harder math. You’d pick differently depending on what you need.
Leaderboards
Leaderboards are public tables where models are ranked by benchmark scores. The most well-known one was the Open LLM Leaderboard on Hugging Face, but it was retired in early 2025 because benchmarks were getting saturated and the rankings stopped being useful. There’s no single official leaderboard. Instead, there are many independent ones like Chatbot Arena, LLM Stats, and various other community-run boards. Companies also publish their own benchmark tables when releasing a model, which always makes their model look great.
Why They’re Not the Full Picture
If benchmark questions accidentally end up in training data, the model isn’t thinking through the answer, but it’s just remembering it.
Companies can specifically train their model to score well on benchmarks without making it better at real tasks.
When every top model scores 95%+, the test can’t tell them apart anymore. So researchers make harder ones, scores drop, the cycle repeats.
Takeaway
Benchmarks give you a rough idea of how good a model is. If one scores 40% and another scores 85% on the same test, of course, the second one is better at that skill. But they won’t tell you which model works best for your case. For that, you just have to try it.
메타데이터
- post_id
- 3b7aa0b9c23b
- slug
- what-is-ai-benchmarking-3b7aa0b9c23b
- url
- https://medium.com/@r_chan/what-is-ai-benchmarking-3b7aa0b9c23b
- canonical_url
- https://medium.com/@r_chan/what-is-ai-benchmarking-3b7aa0b9c23b
- author_url
- https://medium.com/@r_chan
- status
- ok
- fetched_at
- 2026-07-10 21:50:49