← Back to list

How AI models claim that they are the BEST?

Every AI company says their model is the best.

Shrinivas.exe · 2026-07-05 06:34 · 3 claps · 1.5 min read
#llm #ai #claude #openai #benchmarking
Open on Medium ↗
Wiki topics: LLM · Large Language Models EVAL · Evaluation & Benchmarks AI · AI · General

How AI models claim that they are the BEST?

Every AI company says their model is the best.

But best at what, exactly?

I went down a rabbit hole trying to understand how AI models are actually ranked.

Here’s what I found:

An AI benchmark is just a standardised test.

There will be a set of fixed questions, known answers.

  • Massive Multitask Language Understanding (MMLU): 57 subjects — history, law, medicine, maths, physics, coding etc.

MCQs will be there from all the subjects to test the model knowledge across domains. Think of it as the JEE of AI.

  • HumanEval: 164 python coding problems. Model writes code, code gets executed against test cases pass/fail. Tests actual functional code generation — not “does it look right” but “does it run correctly”.

  • HellaSwag: Sentence completion. Given a situation, pick which continuation makes sense. Tests common sense reasoning.

  • GSM8K: Grade school maths word problems, 8,500 problems. Tests multi-step reasoning by tracking whether the formulae are applied in sequence.

  • LMSYS Chatbot Arena: This one is different because instead of fixed questions, real humans chat with 2 anonymous models simultaneously and vote which one they preferred. Crowdsourced human preference.

These were hard in 2021 for AI models, but by 2024 top models where scoring 90%+ on most of them.

A TEST everyone aces TELL YOU NOTHING.

So they built the FINAL BOSS.

  • Humanity’s Last Exam: 2,500 questions from 100+ subjects. Crowdsourced from PhD-level domain experts across the world.

Not the textbook questions. Questions that experts themselves struggle with.

Quantum mechanics. Algebraic topology. Organic synthesis. Ancient linguistics.

When it launched in early 2025, the best models scored in single digits. The same models scoring 90%+ on MMLU.

That gap tells you everything about what MMLU was actually measuring. The part that got me:

HLE’s answers are judged by another AI model not humans.

One model’s output checked against ground truth by a second model.

AI grading AI. At scale. Because humans can’t review 2,500 expert-level answers fast enough.

Benchmarks aren’t perfect either.

A model can score 90% on MMLU and still be useless for your actual problem.


메타데이터
post_id
ebe0b00338bb
slug
how-ai-models-claim-that-they-are-the-best-ebe0b00338bb
url
https://medium.com/@shrinivas.personal/how-ai-models-claim-that-they-are-the-best-ebe0b00338bb
canonical_url
https://medium.com/@shrinivas.personal/how-ai-models-claim-that-they-are-the-best-ebe0b00338bb
author_url
https://medium.com/@shrinivas.personal
status
ok
fetched_at
2026-07-10 21:56:20