Tencent built its own Agent benchmark, yet its own CodeBuddy has never used Claude Code.
2026-08-11 17:22:51
According to CoinMeta, Tencent has built its own set of Agent benchmarks to test the performance of its own CodeBuddy and Claude Code. The 260 real-world problems are divided into four categories: coding, web, office work, and security. Each of the 7 models was tested using two sets of harness. The results showed that Claude Code won 17 out of the comparisons, while CodeBuddy only won 11. In the coding category, Claude Code emerged victorious with a score of 7:0. Although CodeBuddy had a slight 4:3 victory in the web and office work categories, it was defeated by Claude Code with a score of 4:3 in the security category. Tencent's ranking system also has its weaknesses, as there are only 50 to 80 questions in each category. The community has raised doubts about this, suggesting that adding more difficult questions could change the rankings.
Source:Internet
This content is for market information only and does not constitute investment advice.
Follow CoinMeta official accounts to stay updated

Hot Articles
Refresh

'No longer a distant place': F2Pool Co-founder Chun Wang joins SpaceX's 2-year mission to Mars
05-22 18:25

Polymarket Targets Japan Approval Despite Gambling Laws
05-22 18:00

ZachXBT flags suspected exploit involving Polymarket's UMA adapter contract on Polygon
05-22 17:57

ZachXBT flags $520K Polymarket exploit on Polygon, team says funds are safe
05-22 17:24

Verus bridge exploiter returns 4,052 ETH, retains $2.8 million bounty: onchain analyst
05-22 17:24



