News from IT on October 2nd: NVIDIA officially announced an update to its DGX Spark product today, introducing a 64GB memory version of the DGX Spark desktop AI computer, with a starting price of $4,999 (approximately 33,566 RMB at current exchange rates). It will be officially available for purchase on October 23rd. Manufacturers such as Acer, Dell, ASUS, Gigabyte, MSI, and H3C will also launch complete products based on this platform.
The official previously launched 128GB version has also seen its pricing updated accordingly; the 128GB FE version is now priced at $6,950 (Note from IT: The current exchange rate is approximately 46,666 RMB).


In terms of hardware, DGX Spark is based on GB10 Grace to Blackwell supercomputing chips, and relies on a unified memory architecture to connect the CPU and GPU memory spaces. The newly added 64GB memory version has a total memory of 64GB, of which about 8GB is reserved for the system, and approximately 56GB is available for model weights and KV caching. This device supports NVFP4 quantization format and can smoothly deploy mainstream open-source large models such as Gemma4 26B, Qwen3.8-27B, Meta Muse Glimmer, Nemotron3.5, and Lightning. NVIDIA states that even after the model is loaded, there is still sufficient memory remaining for KV caching to support long-context inference tasks.
One of the major highlights of this device is its native capability for multi-machine cluster expansion, featuring a built-in ConnectX-7 high-speed network card, and it comes paired with a brand-new NVIDIA Sync toolkit. According to the official statement, ordinary developers can set up a multi-node cluster without needing in-depth knowledge of network operations and maintenance. The toolkit's built-in Cluster Assistant cluster assistant can automatically complete network configuration, device verification, and SSH environment deployment, supporting up to 4 DGX Spark devices in networking.
Official tests have shown that when two 64GB DGX Spark devices are used to form a cluster to run the Qwen3.8-27B model, the performance is increased by up to 1.7 times compared to using a single device, significantly improving inference throughput. Developers can also remotely access the DGX Spark cluster from ordinary PC devices, connect to mainstream VS Code, Cursor, etc., remotely monitor the status of the entire system, and one-click start the vLLM inference container, thereby reducing the barriers to distributed AI development.
In terms of software ecosystem, DGX Spark is fully equipped with the NVIDIA CUDA acceleration AI software stack, and is deeply adapted to the vLLM and llama.cpp inference frameworks. NVIDIA claims that the product has been specially optimized for local agent workflows, achieving a maximum speed increase of 1.9 times in local agent inference. Perplexity has completed the adaptation, launching a portable computer agent optimized for DGX Spark; moreover, the Laguna S2 118B large model can also be locally deployed and run on this hardware.
At the same time, mainstream open-source community models and frameworks such as Nemotron, Gemma, Qwen, DeepSeek, Mistral, and Stability.ai have also been adapted, enabling users to run cutting-edge models out of the box. Reports indicate that hardware manufacturers and open-source communities have formed a complete developer ecosystem matrix.
Performance tests show that the Qwen3.8 27B model deployed on DGX Spark has a score that is only 10 points behind the leading closed-source models. The report indicates that desktop hardware now has the capability to run high-quality models. The unified memory architecture has also been recognized by industry developers. Perplexity founder Aravind Srinivas commented that DGX Spark can nearly fully utilize the memory capacity, with stable heat dissipation performance. The unified memory design maximizes the efficiency per watt of power consumption, making it suitable for 24/7 uninterrupted operation of local agent tasks.
However, this platform is still targeted at professional developers. With a pricing starting from $4,999 (which is approximately 33,566 yuan at current exchange rates), its target audience mainly consists of AI researchers, independent developers, and small AI startups, rather than ordinary consumer users. With the launch of the 64GB version, the entry barrier has been lowered compared to the 128GB version, allowing more teams to experience the capabilities of local clusters.












