A company that doesn't build AI generates $100 million in annual revenue!
The business miracle was created by Arena, the grand arena where Silicon Valley’s AI giants compete fiercely.
Its predecessor was called Chatbot Arena, initially launched in 2023 as an open-source research project by the UC Berkeley team.

Who could have imagined that, in such a short time, it became the critical battleground for controlling the lifeline of large models?
Today, just eight months after the launch of Arena's commercial services, annual revenue reached $100 million, marking a new milestone.

ChatGPT, Claude dominate rankings; the large model competition
For many people, Arena is not unfamiliar.
Its most celebrated feature is the large model ranking list built entirely through real user blind tests.
The gameplay is incredibly simple yet full of competition—
Enter a prompt, and the system will anonymously send it to two models for simultaneous responses; then select the better one.

The system aggregates millions of such votes into a single Elo-style ranking.
This "tournament-style" mechanism has made it a holy land for AI enthusiasts and developers worldwide.
To date, the platform has accumulated over 10 million user reviews, 700 million conversations, 82 million votes, and more than 10 million monthly visitors from over 150 countries worldwide.
More importantly, approximately 80% of user inquiries each day are entirely new, making it impossible for any model to pre-memorize answers.

How valuable is it?
OpenAI, Google, Anthropic, and Meta—all top companies that normally compete fiercely—are now sending their flagship models to Arena to be tested by the community.
OpenAI secretly tested the model under the codename "Summit" even before the official release of GPT-5.

In other words, the strongest models in all of Silicon Valley are waiting for a project by a Berkeley student to give them the official seal of approval.
How to generate $100 million in revenue?
So the question arises—how did a free leaderboard become a $100 million money machine?
In September last year, Arena launched a commercial service for AI Evaluations:
Model vendors and large enterprises can pay to have Arena mobilize its tens-of-millions-strong community to conduct in-depth evaluations of their models, gaining real-world performance insights that benchmark scores alone cannot provide.
This is a CI/CD system designed for the real world.
Once the model is ready for public release, Arena will freely evaluate it for the community;
Enterprises that want to understand where their model excels, where it falls short, and where it generates misinformation when used by real users must pay.
This is a classic "selling water" business—during a gold rush, those digging for gold may not make money, but those selling water or shovels are guaranteed to profit.
The more fiercely large model vendors compete and the more they strive to extract every last bit of performance, the greater their appetite becomes for this kind of “post-launch optimization” service.
And Arena happened to be positioned exactly where everyone had to pass.
Three people from Berkeley
Make the most profitable trade
Arena was previously called Chatbot Arena.
Previously, it belonged to the renowned LMSYS research group at Berkeley.
Two Berkeley roommates simply wanted to build a neutral arena for large language models, allowing everyone to compare them fairly.
No one expected this student project to rapidly grow into a unicorn.
The timeline is breathtaking: in spring 2025, the project spun off from the university and officially incorporated as a company, securing a $100 million seed round within weeks at a $600 million valuation;
The commercial product launched months later, and within just four months, annualized revenue reached $30 million.
Shortly after, in January of this year, a $150 million Series A round led by Felicis and UC Investments closed, setting the post-money valuation at $1.7 billion.

The three individuals in charge are also no ordinary people.
CEO Anastasios Angelopoulos is, at heart, a mathematician.
While studying for his undergraduate degree in electrical engineering at Stanford, he was taught by the legendary figure in convex optimization, Stephen Boyd.
When I arrived at Berkeley for my Ph.D., my advisors were two legendary figures—machine learning pioneer Michael I. Jordan and computer vision pioneer Jitendra Malik.
Over the years, his research has primarily focused on how to make mathematically rigorous judgments about a black-box model.

CTO Wei-Lin Chiang is a well-known figure in the open-source community—he is the creator of Vicuna, the open-source chatbot that went viral online.
He earned his PhD at Berkeley under Ion Stoica, specializing in distributed systems, and previously worked at Google, Amazon, and Microsoft.
At the moment ChatGPT entered public beta at the end of 2022, he paused all his prior research and dove straight into Arena.

His obsession with the project was described by his partner Angelopoulos as "a labor of love."
For this project, the two worked so long that they moved in together. Two roommates built a $1.7 billion company.
The third co-founder is the renowned UC Berkeley professor and Databricks co-founder, Ion Stoica, who served as an advisor until the project was incorporated in April 2025.
Being a referee is more important than being a player.
Arena's latest move is the launch of Agent Mode.
It no longer evaluates just “who chats better,” but the real work millions of users are doing with agents: writing code, debugging, conducting research, analyzing documents—long tasks involving hundreds of tool calls and multiple rounds of interaction.
It now scores using objective metrics such as task completion rate and hallucination rate, far exceeding the original scope of "human preference voting."

AI is evolving from "chatbots" into autonomous "agents" capable of handling increasingly complex tasks with higher stakes.
Evaluation is humanity's final probe into the inner workings of AI.
The business of Arena could be worth $100 million or $1.7 billion—it ultimately hinges on the bet that it will become increasingly important and more valuable.
But everyone will eventually have to answer the same question—who is still qualified to grade when machines start creating their own questions?
Reference materials:
https://techcrunch.com/2026/06/29/arena-the-ai-leaderboard-everyone-uses-is-now-a-100m-business/
https://x.com/ml_angelopoulos/status/2071629882057228680?s=20
This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation, edited by Peach.
