On October 9th, according to local time on the 8th, the Arena Arena, a model evaluation platform owned by AI, announced the completion of a $200 million (Note from IT: Current exchange rate is approximately 1.343 billion RMB) Series B financing, raising the company's valuation to $3.1 billion (Current exchange rate is approximately 20.817 billion RMB). Arena originated as a research project at the University of California, Berkeley in 2023, which involved ranking models through public voting.
As previously disclosed, in June of this year, the annualized revenue reached 100 million US dollars (equivalent to approximately 672 million RMB at current exchange rates).
This round of financing was co-led by Guangsu Venture Capital and Khosla Ventures. Institutions such as Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, and Felicis also participated in the investment.
In January this year, Arena completed a Series A financing of $150 million (approximately 1.007 billion yuan at current exchange rates), with a post-investment valuation of $1.7 billion (about 11.416 billion yuan at current exchange rates). At that time, its annual revenue was $30 million (approximately 201 million yuan at current exchange rates). In just about 10 months, the valuation of Arena has nearly doubled.

The evaluation platform for Arena is open to individual users for free. Users can enter prompts, or ask AI to write programs and develop projects according to their own requirements, and then compare the completion results of different models and rate them. Arena claims that the platform attracts tens of millions of visitors each month.
Last September, Arena launched a commercial service AI Evaluations targeting AI model research and development institutions and enterprises. By utilizing feedback data from community users, it provides detailed model performance analysis.
This year, the AI research institution discovered that their models are able to achieve high scores by conforming to the rules of benchmark tests, but they may not necessarily possess the corresponding actual capabilities. Enterprises are no longer satisfied with standardized test results and wish to understand which models are more suitable for their business needs.
In the financing announcement, Arena stated: 'The development speed of AI has surpassed our ability to assess it. Once the model realizes it is being tested, the fixed benchmark tests become ineffective. The world needs a neutral third party to verify whether AI is safe in actual use and whether its behavior meets human expectations. Arena has begun to take on this role.'
For this reason, Arena has added an alignment capability evaluation to the rankings, which focuses on whether the model will perform operations without user request, whether it will misattribute statements or facts to the wrong sources, and whether it will falsely claim to have completed a task. Arena refers to the last behavior as "deceptive completion."
In the preliminary alignment capability rankings, multiple models of OpenAI ranked at the top, with Claude Opus coming in at sixth place with a score of 5.5, and Claude Fable ranking ninth.












