DeepSeek Breaks the "impossible trinity" of large models: stronger, faster, and cheaper
2026-09-10 15:05:20
According to CoinMeta, DeepSeek announced that its v4.1 flash version has broken the "impossible trinity" of large models, achieving stronger, faster, and more affordable results. The main parameters of this model reach 552 billion, with an additional 196 billion parameters from engram conditional memory. Pre-training utilized over 450,000 modal token, while post-training incorporated real agent tasks and failure cases. The performance of DeepSeek v1.1 reached 74.2%, surpassing that of Claude Opus 5 and GPT-5.6 SOL. The new architecture splits the 40-layer model into two halves; during reading, only 8 billion parameters are activated per token, and during generation, 16 billion parameters are activated. Additionally, the context length has been extended from 4k to 1m, with only a quarter increase in computational load. Furthermore, by changing the main cache to fp4 and implementing cross-layer reuse, the global kv for each token has been reduced to 890 bytes, which is about 1/4 of that of v4 flash.
Source:Internet
This content is for market information only and does not constitute investment advice.
Follow CoinMeta official accounts to stay updated

Hot Articles
Refresh

Privacy Coin Sector Outlook Sept 2026: 4 Dark Horses
39m ago

S&P 500 Top 5 Companies: The Leaders Explained at a Glance
09-09 17:57

2026 Crypto Prediction Market: Is Polymarket the Leader or Being Surpassed?
09-08 18:26

SOL Price Prediction September: Can It Break $150?
09-07 17:10

US Stocks vs A-Share Market: 5 Key Mechanisms That Drive Price Movements
09-04 18:47



