Arena was initially launched in 2023 at the University of California, Berkeley as a research project. The project ranked the AI model through a crowdsourcing approach. The company announced on Thursday that it has completed a $200 million Series B financing round, with a valuation of $3.1 billion.
Previously, Arena stated that its annualized revenue operating rate reached $100 million in June.
This round of financing was led by Lightspeed Venture Partners and Khosla Ventures, with participation from institutions such as Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis, etc. Arena previously announced the completion of a $150 million Series A financing in January, with a post-investment valuation of $1.7 billion at that time. The company stated that its annual revenue was $30 million at that time. Based on this, its valuation has almost doubled in about 10 months.
Arena provides a crowdsourcing platform for consumers to use for free. Users enter prompts or propose project requirements for “vibe-coded”, and then vote on which model performs better. Arena claims that its platform attracts tens of millions of visits per month.
Last September, the company launched a commercial product AI Evaluations, which is a service that provides detailed performance analysis for model laboratories and enterprises based on community feedback. The timing of its launch proved to be very appropriate. This year, the AI laboratory realized that their models were "brushing" benchmark tests in order to obtain high scores without actually winning. At the same time, enterprises also hoped to get assistance in determining which model best suits their internal needs, rather than relying solely on standardized benchmarks.
"The development speed of AI exceeds our ability to assess it, and once the model realizes it is being tested, the static benchmarks become ineffective," the company stated in its financing announcement. "The world needs a neutral third party to measure how secure AI is in the hands of real users and how consistent it is with human intentions. Arena is taking on this role today," the company added.
For this reason, Arena has also added a new category to its rankings: alignment (Alignment). This category ranks models based on unauthorized actions (performing operations that were not requested), misattribution of errors (attributing statements or facts to improper sources), and what the company refers to as “deceptive completion” (falsely claiming to have completed tasks that were not actually finished).
Currently, a series of models from OpenAI rank at the top of their preliminary alignment charts, while Claude Opus 5.5 and Claude Fable are ranked sixth and ninth respectively.












