As AI models become increasingly lightweight, VRAM is becoming more expensive. NVIDIA is addressing this contradiction with a product that cuts memory.
On Friday, October 2, Eastern Time, NVIDIA announced the DGX Spark desktop AI computer with a 64GB unified memory configuration, starting at $4,999, and set for market release on October 23 through OEM partners including Acer, ASUS, Dell, Gigabyte, HP, and MSI. The new system retains the GB10 Grace Blackwell superchip, DGX OS, the full NVIDIA AI software stack, and ConnectX-7 networking capabilities, but reduces unified memory from the existing 128GB to 64GB.
Interestingly, the starting price of this reduced-spec DGX Spark is higher than the initial launch price of $3,999 for last year’s 128GB version; at the same time, NVIDIA has further increased the price of the 128GB Founders Edition to $6,950. Media outlets note that NVIDIA attributes this price hike to limited memory supply and rising costs.
This means that, amid ongoing demand for AI servers squeezing memory supply, "reducing memory" has become a way to lower the price barrier for local AI devices—but it doesn't mean the devices themselves have actually become cheaper.
Is 64GB enough? NVIDIA targets local AI agents.
The core logic behind NVIDIA's recent adjustment is that an increasing number of open-source models can now run on smaller memory capacities.
NVIDIA stated that with the improving capabilities of open-source models and their continuously shrinking sizes, 64GB of unified memory is now sufficient to support a range of local AI applications. The new DGX Spark can run models with up to 100 billion parameters locally and supports workloads such as AI agents, inference, fine-tuning, data science, and edge development.
Tom's Hardware also notes that some of the latest high-performance, dense models can now run on around 32GB of memory, though they remain limited in scenarios requiring larger context windows. This means that the 128GB of memory originally prepared for "running large models locally" is not necessary for all developers.
This is the practical basis for NVIDIA’s launch of the 64GB version: for users primarily engaged in local inference and developing AI agents, rather than large-scale model training or fine-tuning, 64GB can cover a significant portion of their workloads.
NVIDIA even summarized this trend in its announcement as, "As technology continues to evolve, the practicality of local AI is growing stronger," as AI agents move from experimentation to everyday development, increasing the demand for models that run locally.
Meanwhile, on-premises deployment offers advantages in privacy and cost. Developers can process their own data and run AI agents directly on DGX Spark without needing to call cloud-based models for each task.
The performance of the same GB10, 64GB version has not been reduced.
From a hardware architecture perspective, the 64GB version is not a completely new chip product.
NVIDIA stated that the new system continues to use the GB200 Grace Blackwell Superchip and retains the full DGX OS and AI software stack. Media noted that the 64GB version also retains the original 20-core Arm CPU and 273 GB/s shared memory bandwidth, meaning that for models that fit within 64GB of memory, the underlying computational performance remains unchanged despite the memory capacity being halved.
In other words, the focus of this change is not on reducing hashing power, but on eliminating a portion of memory capacity that some users may not need.
The 64GB version still supports mainstream inference frameworks such as llama.cpp, Ollama, vLLM, and LM Studio, and comes pre-installed with NVIDIA Agent Toolkit, CUDA-X AI libraries, and open models like Nemotron.
NVIDIA’s strategy is also clear: if developers currently only need 64GB, they can start by purchasing a machine with lower capacity; if model sizes continue to grow in the future, they can scale up memory and computing power through clustering.
Two 64GB units can be combined to form 128GB, with performance increasing by up to approximately 70%.
Another key focus of this DGX Spark update is NVIDIA's further enhancement of multi-machine collaboration.
The 64GB version also features a built-in ConnectX-7 network interface, allowing the two devices to be directly connected via QSFP cables and automatically configured for networking using NVIDIA Sync Cluster Assistant, forming a local AI cluster. NVIDIA states that two 64GB DGX Spark units can create a 128GB memory pool, supporting models with up to 200 billion parameters.
In NVIDIA's testing of Qwen 3.8 27B, a cluster of two 64GB machines achieved a maximum performance of approximately 1.7 times that of a single machine.
NVIDIA will also launch the NVIDIA Sync Model Launcher at the end of October to further simplify the deployment of models on single machines or clusters. For example, developers can use this tool to run Qwen 3.8 27B and integrate it with programming tools such as OpenCode.
This means that DGX Spark’s product logic is shifting from “a desktop AI supercomputer” to “a locally deployable AI node that can be scaled incrementally.”
Of course, two machines do not equate to a single physical 128GB DGX Spark; whether a specific model can run across nodes still depends on the software and workload. Therefore, for tasks requiring large memory on a single machine, the 64GB version still has significant limitations.
Under conditions of GPU memory shortage, the 64GB version at $4,999 is not truly "cheap."
What truly matters is the price.
NVIDIA officially lists the starting price of the 64GB DGX Spark at $4,999, and this version will not include a NVIDIA-branded Founders Edition; it will be sold exclusively through OEM partners. Partners include Acer, ASUS, Dell, Gigabyte, HP, and MSI, with specific configurations and pricing potentially varying.
For comparison, when the DGX Spark was first launched in 2025, the official price for the 128GB version was $3,999; media reported that NVIDIA subsequently raised the price of the 128GB Founders Edition to $4,699 in February 2026, and this latest adjustment brings it to $6,950.
Therefore, if we compare only the original launch prices, the 64GB model today is not cheaper than the 128GB model—it’s actually $1,000 more expensive, representing a 25% increase.
Compared to the current official NVIDIA price of $6,950 for the 128GB version, the 64GB version at $4,999 is still $1,951 cheaper, but at the cost of halving the memory capacity.
PC Watch, citing relevant information, stated that the price increase for the 128GB version is related to memory supply constraints and rising costs; Tom's Hardware noted that the current market price for a 128GB GB10 system has reached approximately $7,000 to $9,000.
Therefore, another implication of this product adjustment is that AI computing power is increasingly moving to local devices, but memory has become a significant cost bottleneck in this trend.
Transition from "heap memory" to "on-demand scaling"
From a product strategy perspective, NVIDIA is not simply turning the DGX Spark into a "lower-spec" version.
In the past, one of DGX Spark’s key selling points was its 128GB unified memory, enabling developers to run large-parameter models on desktop devices; today, advancements in model compression and quantization technologies have made such large memory capacities unnecessary for some models.
Therefore, NVIDIA's solution is: a single machine with 64GB of memory suffices for mainstream local AI inference, and scaling to a multi-machine cluster is used when memory requirements increase further.
The 64GB version will be officially launched on October 23 with a starting price of $4,999. Meanwhile, the price of the 128GB version has increased to $6,950, making this new product particularly distinctive—it represents NVIDIA’s attempt to lower the barrier to entry for local AI hardware, as well as a reflection of the current tension in the AI industry between surging compute demands and tightening memory supply.
