Apple opens-source LensVLM-9B: 100-page document compressed into images first, KV cache saves 84%
2026-09-24 18:53:16
According to CoinMeta, Apple has open-sourced a visual language model specifically designed for processing long documents, LensVLM-9B, which is trained based on Qwen3.5-9b-base. When handling long documents, LensVLM first compresses the entire document into a low-resolution version for a quick scan. After locating the relevant pages, it then reads the original text or high-definition images. Tests show that directly reading a 100-page document requires 51,273 tokens, whereas LensVLM only needs 8,090. The cache size of KV was reduced from about 1.6GB to 253MB, a decrease of 84.2%. In seven document Q&A tests, the average accuracy of directly reading the full text was 72.4%, while LensVLM still maintained an accuracy rate of 68.9% even after being compressed by about 4.3 times. Although the accuracy rate does not decrease significantly after compression, the processing speed is slower; in the paper tests, each response took approximately 17 seconds.
Bullish 0
Bearish 1
Source:Internet
This content is for market information only and does not constitute investment advice.