Skip to content
CampeloLabs
← Blog

AI Benchmarks: What the Leaderboard Hides

Cicero Campelo

Cicero Campelo, CISSP
September 29, 2026 · 11 min read

Part of our guide to AI for startups.

A founder turning away from a wall of glowing leaderboard rankings to read one small private scorecard in her hand
Table of contents

An AI benchmark is a fixed set of tasks plus a rule for scoring the answers, used to compare models against each other. Founders reach for them at exactly one moment: when they have to pick a model and want a number that makes the decision for them.

The number will not make the decision for you. The scores you see quoted are largely self-reported by the labs that build the models, the tests themselves stop working at the moment they become popular, and any single benchmark is narrower than the headline makes it look. None of that makes them worthless. It means a leaderboard is evidence, not a verdict, and reading one well is a skill worth ten minutes of your week.

The clearest account of why comes from Rayan Krishnan, co-founder and CEO of the independent benchmarking company Vals AI, interviewed on a16z's podcast by Ben Horowitz and Jennifer Li. Vals exists because a team concluded in 2024 that public benchmarks were no longer sufficient to measure progress. That is a self-interested position, and it is also mostly correct.

The problem is self-reporting, not dishonesty

Model labs build excellent internal benchmarks. That is what drives progress. The issue starts when those results become the public evidence for how good a model is. Krishnan's framing is that there is an issue when the industry discusses "model capabilities in a way that's self-reported".

His example is Llama 4. On Vals' held-out private benchmarks, he says, the model underperformed. On the major public benchmarks, where the questions and the rubrics are open source, it looked excellent. "there's a huge disconnect between what was self-reported based on these open benchmarks", he says, and what his team measured.

The public record around that launch tells a parallel story about a different benchmark. In April 2025, Meta submitted a variant labelled Llama-4-Maverick-03-26-Experimental to the human-preference leaderboard then called LMArena, and it scored well. The publicly released weights were a different model, and when the unmodified version was tested it landed far down the board. Meta's VP of generative AI denied that the company had trained on test sets, and a Meta spokesperson said the leaderboard entry was a chat-optimized version the team had experimented with. LMArena then tightened its policy so that submissions have to match released weights.

Two different benchmarks, one mechanism. Once a test is public, the test is also training data, and the thing being measured is partly the ability to do well on that test. This is Goodhart's law with a marketing budget attached.

For a founder the practical read is narrow and useful: a score a lab published about its own model on a public benchmark is a claim, not a measurement. Treat it the way you treat any vendor's own numbers.

A benchmark stops being useful once it works

The second structural problem is that good benchmarks destroy themselves. A benchmark that everyone optimizes against gets saturated, and a saturated benchmark carries no information.

Vals' answer is to retire them. Krishnan describes an unofficial company motto, "always a higher peak", and the job as a permanent construction project: "It is our job to perpetually construct these next mountains for them to summit."

There is a second reason to retire a benchmark that gets less attention. A test also goes stale because the world moves. Krishnan's analogy is professional recertification: lawyers retake the bar, architects and doctors recertify, and "we should also expect models to be tested on the current state of the world". A legal benchmark built on superseded case law is measuring the wrong thing even if no model has saturated it.

So the age of a benchmark matters as much as the score on it. A leaderboard where the top ten models sit within two points of each other and of the ceiling is not telling you they are equivalent. It is telling you the test is finished.

Fewer questions, much harder grading

The shape of benchmarks has changed, and it changes how much a gap between two models means.

Early benchmarks were enormous and simple. ImageNet is millions of images mapped to one label each. Modern agentic benchmarks invert that. Krishnan's description of the trend is that "evaluations as they become more complex have a fewer sample size but a larger set of criteria or expectations of them". The task is now something like "generate me 50 fullstack web applications", graded against a long rubric, sometimes over runs that last hours or days.

That is the right direction for realism and it has a consequence founders should internalize. Fifty tasks is a small sample. A three-point gap between two models on a fifty-task agentic benchmark can be noise, run-to-run variance, or one rubric item. On a million-image benchmark, three points was real.

Read the sample size before you read the ranking. If the benchmark does not publish one, that is itself the finding.

Every leaderboard ranks exactly one thing

The most common mistake is reading a narrow leaderboard as a general ranking.

In July 2026, Moonshot AI's open-weight Kimi K3 took the top spot on Arena's frontend code leaderboard, ahead of Claude Fable 5 and GPT-5.6 Sol, while trailing those same models on broader intelligence indices. Both facts were true at once. Best model was never a coherent question. Best at frontend code, judged by human preference, in a given month, is a question with an answer. That board has since turned over again, which is the other half of the lesson: a leaderboard is a snapshot, and citing one without its date is citing nothing.

Arena (formerly LMArena) is worth understanding as a deliberately different instrument. Its co-founder and CEO, Anastasios Angelopoulos, describes it on 20VC as the platform "for measuring AI performance in the real world", and is explicit about the contrast: "we're not using static benchmarks", but rather "what happens when you put AI in the hands of real people".

That buys resistance to saturation, because the prompts keep changing. It also imports human preference, and Arena's own research is candid about what that brings in: its style-control work finds response length and markdown formatting move the scores enough that they have to be controlled for, and it has since added the same treatment for sentiment. A held-out expert rubric and a crowd vote are measuring different things, and neither one is measuring your product.

And capability is only one axis. Cost, latency, and how well a model holds up across a range of tasks are separate questions, and they are usually the ones that decide a build. Our write-up on choosing between the leading models covers how to hold those axes apart.

Check who is paid by whom

This is the part a security background makes hard to ignore. Independent assurance is only worth something when the assurer has nothing to gain from the result, and the AI benchmark market has not settled that question.

Horowitz reaches for an older parallel, calling the situation "a little bit reminiscent of the MPAA": a body drawing fuzzy lines that shift over time, with no crisp definition underneath. His objection to the current public benchmarks is two-sided and worth holding onto, that "it's hackable and then it's too narrow".

Krishnan reaches for auditing, and specifically for Enron: when the same firm audits a company and sells it consulting, the incentives collapse and "pay to pass the audit" becomes the product. His version of separation of duties was a very early decision to "never sell training data to labs", despite being pushed toward it, because he says a good deal of that industry has built "gimmick style benchmarks as a mechanism to sell their data". A benchmark that exists to advertise a data business is a brochure.

Before you weight a third-party score, ask three things. Who pays for the evaluation. Whether the evaluator sells anything else to the labs it grades. Whether the test set is held out or published. None of those require insider access, and any of them can disqualify a number.

The same logic runs up into policy. Horowitz's split is that government is good at setting and enforcing rules and poorly suited to running the tests, so the open question after "can the model do it" is "can you get the model to do it", and somebody independent has to answer both. Demis Hassabis has argued along similar lines for a US-led frontier standards body supported by an ecosystem of third-party auditors.

The only benchmark that picks your model

Here is the finding that should change what a founder does on Monday. Krishnan is blunt that it is still unclear whether the best OpenAI model or the best Anthropic model "is actually going to be best for your repository", and that "we've seen a lot of non-intuitive examples where you actually had to run the eval to figure out what's going to be the frontier performance for that repository".

His company's own product follows from it: Vals Smith turns a codebase's merged pull requests into a private coding benchmark, hiding the tests and checking whether a model resolves the task without the original developer's fix. The concept generalizes past code. You already own a repository of completed work with known-good answers, whether that is closed tickets, past contracts, resolved support threads, or shipped analyses. That archive is a benchmark nobody else has.

Harrison Chase, co-founder and CEO of LangChain, built a Sequoia talk around three lines from a post by Microsoft CEO Satya Nadella. The first is the one that matters here: "create your private evals because eval defines what good looks like inside the organization".

This is also why routing is harder than it looks. Krishnan's observation is that "the hardest part of routing is building the evals": deciding which model handles which request is trivial once you can score the outcomes and impossible before. The same gate sits in front of cost control, which is where our post on tokenmaxxing ends up, since you cannot tell an efficient model from a cheap one without a scorer.

Krishnan's strongest line is the one to keep: "a firm really is just its eval". If you cannot say what good output looks like precisely enough to test it, you cannot pick a model, price the work, or tell whether the last model upgrade helped. That capability, not the model, is the durable part, which is the argument running through our AI for startups pillar and the mechanics of which are in LLM evals for founders.

How to read a leaderboard in ten minutes

  1. Find out who ran it. A lab reporting on its own model is a claim. A third party with no other business with that lab is evidence.
  2. Check whether the questions and rubrics are public. If they are, assume some degree of training on them and discount accordingly.
  3. Look at the sample size. Small-n agentic benchmarks have wide error bars, and a few points of separation is often nothing.
  4. Check the date of the test, not just the date of the run. A benchmark built on a stale snapshot of the world measures the wrong world.
  5. Read the exact task. Frontend code judged by human preference does not generalize to your backend refactor.
  6. Check whether the top of the board is bunched near the ceiling. That is saturation, and it means the test is over.
  7. Treat the whole thing as a shortlist generator. Two or three candidates, then your own eval decides.

What to do this week

  • Write down the decision you are actually making. Which model for our support triage is answerable. Which model is best is not.
  • Pull 30 to 50 real, completed examples from your own archive: closed tickets, merged pull requests, past deliverables. Keep the known-good outcome.
  • Define what good looks like as a rubric someone else could apply without asking you. That document is the asset, more than the scores it produces.
  • Hold that set out. Do not paste it into a prompt, a fine-tune, or a vendor demo, or you have published your own benchmark.
  • Run your two or three leaderboard shortlist candidates against it and record accuracy, cost, and latency per task, not just a win rate.
  • Re-run it at the next model release. The interesting number is the delta, and you cannot get a delta without a fixed test.

If you want the operating system around this, from picking models to pricing the work they do, that is what AI Operating System for Startups is built to teach.

Sources

Frequently asked questions

Are AI benchmarks reliable?

They are reliable about what they measure and unreliable as a general verdict, and the gap between those two things is where founders get burned. Three structural problems apply to most public benchmarks. The headline scores are largely self-reported by the labs that build the models, so a published score is a vendor claim rather than an independent measurement. Their questions and rubrics are usually open source, which means the test is also training data and models are partly being scored on the ability to do well on that specific test. And they saturate: once every lab optimizes against a benchmark, the top models bunch near the ceiling and the ranking stops carrying information. Rayan Krishnan, co-founder and CEO of the independent benchmarking company Vals AI, points to Llama 4 as the illustration. On his team's held-out private benchmarks the model underperformed, while on the major public benchmarks it showed excellent capability. Use benchmarks to build a shortlist of two or three candidates, then decide with a test you built from your own work.

Why do AI models score differently on different benchmarks?

Because each benchmark ranks exactly one narrow thing, and the things are genuinely different. A benchmark is a set of tasks plus a rubric, so changing either changes the winner legitimately, with no gaming involved. In July 2026, Moonshot AI's open-weight Kimi K3 took the top spot on Arena's frontend code leaderboard ahead of Claude Fable 5 and GPT-5.6 Sol, while trailing those same models on broader intelligence indices and most agentic benchmarks. Both results were correct, and that board has since turned over again, so always read the date on a ranking. The method matters as much as the task: a held-out expert rubric and a crowd preference vote measure different things, and Arena's own style-control research finds that response length and formatting move its scores enough that they have to be controlled for, which is not the same as being more correct. Capability is also only one axis. Cost, latency and consistency across a range of tasks are separate measurements, and they usually decide a real build. Read the exact task and the grading method before you read the ranking.

What is the best AI benchmark for choosing a model?

The one you build from your own completed work, because no public benchmark was measured on your tasks. Krishnan's finding is that it remains unclear whether the best OpenAI model or the best Anthropic model will be the best model for your repository, and that his team keeps hitting non-intuitive cases where you had to run the eval to find out. Vals AI's own product follows the logic: Vals Smith turns a codebase's merged pull requests into a private benchmark, hiding the tests and checking whether a model resolves the task without the original developer's fix. The concept is not limited to code. Closed support tickets, past contracts, resolved engineering threads and shipped analyses are all archives of work with known-good answers, and that archive is a benchmark no competitor has. Harrison Chase, co-founder and CEO of LangChain, makes the same case at Sequoia by quoting Microsoft CEO Satya Nadella: create your private evals, because the eval defines what good looks like inside the organization. Public leaderboards are still useful upstream of that, as a shortlist generator.

Why do AI benchmarks keep getting retired?

For two reasons, and only one of them is saturation. The obvious one is that a benchmark everyone optimizes against eventually stops separating models, so it has to be replaced by a harder test. Krishnan describes that as a permanent construction project, with an unofficial company motto of always a higher peak and the job being to perpetually build the next mountains for labs to summit. The less obvious reason is that the world moves. He compares it to professional recertification: lawyers retake the bar and doctors recertify, so models should also be tested on the current state of the world rather than on a stale snapshot of it. A legal benchmark built on superseded case law measures the wrong thing even if no model has saturated it. For a founder reading a leaderboard, this means the age of the test matters as much as the score, and a board where the top ten sit within a couple of points of each other and of the ceiling is reporting that the test is finished, not that the models are equivalent.

Build your AI Operating System

A practical course to grow with AI, build internal tools, and operate safely. Join the waitlist and you'll be first in when the course opens.