FreeToken Open-source: Running a DeepSeek-V4-Flash with a maximum of 25 Token /s
2026-09-01 15:37:47
According to CoinMeta, the FreeToken open-source project was launched by researchers from Berkeley, MIT, and others. It aims to create a hybrid expert model inference engine designed for local devices, addressing the issue of insufficient video memory in very large models. This engine stores some of the weights in system memory, allowing CPU and GPU to participate in the inference process together. Tests have shown that a desktop equipped with a RTX 5090 graphics card running a DeepSeek-V4-Flash model with 284B parameters can achieve a speed of 22–25 Token per second, while on a 8GB video memory RTX 4060 laptop, a 35B model can reach a speed of 39.3 Token per second. The project currently supports over 20 different MOE models and has been made open-source under the Apache 2.0 license.
Source:Internet
This content is for market information only and does not constitute investment advice.
Follow CoinMeta official accounts to stay updated

Hot Articles
Refresh

'No longer a distant place': F2Pool Co-founder Chun Wang joins SpaceX's 2-year mission to Mars
05-22 18:25

Polymarket Targets Japan Approval Despite Gambling Laws
05-22 18:00

ZachXBT flags suspected exploit involving Polymarket's UMA adapter contract on Polygon
05-22 17:57

ZachXBT flags $520K Polymarket exploit on Polygon, team says funds are safe
05-22 17:24

Verus bridge exploiter returns 4,052 ETH, retains $2.8 million bounty: onchain analyst
05-22 17:24



