Arena has built a $100 million run-rate business from AI evaluation, but the same leaderboard that made it trusted is now carrying a conflict the company can't wave away.
Eight months is a short time to build a $100 million business. Arena, the company behind the Chatbot Arena leaderboard, only launched its commercial service in September 2025, and TechCrunch reported that it has already crossed that annualized run-rate mark. Back in January, when the company raised a $150 million Series A at a $1.7 billion valuation, its annualized revenue was $30 million. You don't have to admire the AI boom to see the speed here. The number is real enough to matter.
The question is what kind of business that number is buying.
Arena grew out of a UC Berkeley research project led by Wei-Lin Chiang, Anastasios Angelopoulos and Ion Stoica, the Berkeley professor and Databricks co-founder. It started as a free, crowdsourced way to compare AI models through blind, head-to-head votes. Users asked a question, two unnamed models answered, and the user picked the better response. That simple format became one of the AI industry's most watched scoreboards because it tested something synthetic benchmarks often miss: what real people prefer when they don't know whose model they're using.
That community is now the product. Arena's commercial service, AI Evaluations, sells deeper analysis from those user comparisons to model labs and enterprises. The company says its users have submitted more than 10 million head-to-head model comparisons. That's not a small sample, and it explains why Andreessen Horowitz, Kleiner Perkins, Felicis and Lightspeed have all backed the company across funding rounds.
Here's the thing. A leaderboard's value depends on trust before it depends on revenue. If you use Arena to judge whether one model is better than another, you need to believe the contest is clean. If you're building a model, you want the same thing, unless the rules let you test privately, tune around the test, and publish only the result that flatters you.
That is where Arena's business gets uncomfortable. Its paying customers include AI companies whose models appear on the leaderboard. OpenAI, Google, Anthropic and Meta are not just names in this market, they are the market many readers are trying to understand. When companies like that depend on a ranking and can also pay for evaluation products connected to the same ecosystem, Arena has to do more than say the teams are independent. It has to show you why the independence should be believed.
The strongest criticism came in April 2025, when researchers from Cohere, Stanford, MIT and AI2 argued that Arena had allowed some companies to privately test multiple model variants before public release, then surface the best-performing version. Meta reportedly tested 27 variants of what became Llama 4 between January and March before publicizing one score that ranked near the top. Arena disputed the way the episode was characterized, and it later tightened its policies. Fine. But the criticism stuck because it described a mechanism, not just a mood.
If private pre-release testing is available to some model makers, and if the public later sees only the winner, the leaderboard starts to look less like a neutral contest and more like a launch tool. That doesn't require corruption. It just requires incentives.
Arena says model submissions are processed blindly and that its crowdsourced voting pool is too large and distributed to manipulate easily. Those points count. A public, blind, user-driven benchmark is still more useful than a vendor's own launch chart with a dozen cherry-picked tasks. But independence isn't only about whether someone can stuff the ballot box. It's also about whether the public can see enough of the rules to know who got to practice before the match.
The revenue figure needs the same plain treatment. TechCrunch described Arena's $100 million figure as annualized run-rate revenue, not contracted recurring revenue. That distinction matters because Arena charges on consumption. A model lab can spend heavily while it is evaluating a new release, then pull back once that work is done. That doesn't make the revenue fake, but it does make it less predictable than a fixed subscription business. Investors know the difference. Readers should too.
None of this means Arena is a bad company or that its growth is some accounting trick. The opposite is more likely. AI evaluation has become important because the old benchmark culture is visibly strained. A model can ace a static test and still be irritating, evasive or weak in the tasks people actually care about. Arena's original insight was that human preference, gathered at scale, could tell you something useful. That insight is still valuable.
But the company is now selling access to the trust it built as a public referee. Frankly, that is the whole story. If Arena wants to remain the AI industry's scoreboard, it needs a clearer wall between paid private evaluation and public leaderboard placement. It should disclose more about how many private variants are tested, how public submissions are selected, and what restrictions apply before launch. A $100 million run-rate business is impressive. A trusted benchmark is rarer.
Also read: Cursor's mobile app signals that coding has become a job you supervise, not a desk you sit at • Copper has broken records this year and AI data centers are the reason the rally isn't done • The semiconductor layer is where the real AI money is being made in 2026