Every local AI purchase runs into the same question sooner or later: will the model I want actually load?
Every local AI purchase runs into the same question sooner or later: will the model I want actually load? Speed comes second. A fast machine that refuses to load your model is a toolbox missing the one socket size you need on a Sunday night. It looks complete and helps nobody.
NVIDIA’s new DGX Spark 64GB goes after that question with half the memory of the original DGX Spark and the same GB10 chip inside. NVIDIA says it reaches buyers through Acer, ASUS, Dell, Gigabyte, HP and MSI from October 23, 2026, and ProX PC will have it available in India when it launches. We have not run the 64GB unit yet. We did test the 128GB version against an RTX 5090 workstation, so this guide uses those results to show where the 64GB should fit, and says clearly where we are reasoning instead of measuring.
Quick answer
The 64GB keeps the GB10 chip, the 273 GB/s memory bandwidth and the 200GbE ConnectX-7 networking port of the 128GB. Only the memory pool shrinks.
It targets the gap between 32GB and 64GB: models too big for an RTX 5090 that do not need the full 128GB.
Models that fit inside 32GB run faster on an RTX 5090 build. Our tests showed that clearly.
The DGX Spark is a desktop AI computer built around NVIDIA’s GB10 Grace Blackwell Superchip, which puts a 20-core Arm CPU and a Blackwell GPU into one package that shares a single pool of memory. That shared pool is called unified memory, and it lets the GPU use most of the machine’s RAM for a model, where a normal PC limits the GPU to the VRAM on its own card. The 64GB version offers 64GB of that pool with the same headline 1 petaflop AI rating as the 128GB. It runs DGX OS, NVIDIA’s Ubuntu-based Linux, so Windows is off the table.
The surprise is what NVIDIA kept. The 64GB retains the ConnectX-7 200GbE port of the larger model, a networking part usually found on server hardware. That suggests NVIDIA sees the 64GB as a starting block you can add a second unit to later.
Same chip and same 273 GB/s of bandwidth mean the two should generate tokens at similar speeds on any model that fits in both. Treat that as an expectation until we run the 64GB ourselves. Memory size sets the ceiling. On our 128GB unit, Qwen 2.5 72B and Llama 3.2 Vision 90B both ran steadily at about 4.6 tokens per second. A 4-bit 70B model file typically takes 40 to 47GB, so it should load in 64GB with some room left, though long prompts and the operating system draw from the same shared pool.
That shared pool is where the 128GB earns its place. Our GPT-OSS 120B fine-tune peaked at about 59GB of memory. On a 64GB unit that leaves roughly 5GB for everything else, so we would plan that job on the 128GB.
Two 64GB units can link over ConnectX-7 to act like one 128GB system. NVIDIA’s own clustering slide lists 1.7x relative performance for the pair, which falls short of double. Our advice is to scale up before you scale out, meaning one bigger machine before several smaller ones. Starting with one 64GB and adding a second later gives you a growth path. Buying two on day one to reach 128GB is like doing a RAM upgrade in two rounds when the single bigger kit sat on the shelf the first time.
In January 2026 we ran the ASUS Ascent GX10 (a 128GB DGX Spark platform) against our RTX 5090 workstation. The full benchmark is here for anyone who wants every table. The short version:
| Test | RTX 5090 workstation | DGX Spark 128GB |
|---|---|---|
| Qwen 2.5 7B generation | about 220 tokens/s | about 47 to 50 tokens/s |
| Qwen 2.5 72B | could not load | about 4.6 tokens/s |
| Llama 3.2 Vision 90B | could not load | about 4.6 tokens/s |
| Fine-tune GPT-OSS 20B | 3.45 min | 4.62 min |
| Fine-tune GPT-OSS 120B | failed | 17.82 min |
| Flux Dev image generation (first run) | 50 s | 237 s |
| Hunyuan Video 1.5 | 1310 s | 3606 s |
| Power drawn at the wall | 800 to 900W | under 100W |
The 5090 wins whenever the model fits in its 32GB. Memory bandwidth, meaning how fast the chip can read a model’s weights, explains most of it: the 5090’s GDDR7 moves about 1,792 GB/s against the Spark’s 273 GB/s. For the 27B to 32B class, our Spark ran at roughly 10 to 12 tokens per second while the 5090 ran the same size class at about 64 to 66.
Once the model outgrows 32GB, the picture flips. The 5090 could not load the 72B and 90B models at all, and the Spark ran them, with a wait attached. On very long prompts, the first word took 133 seconds on the 90B model and about three minutes on DeepSeek R1 70B. You trade time for the ability to run the model.
Fine-tuning showed a gap smaller than most people expect. The 5090 finished the 20B job about a third faster, a much narrower margin than chat speed would suggest, and the 120B job only completed on the Spark.
Power was the number that surprised us most. The 5090 system pulled 800 to 900W from the wall, with the card alone rated at 575W, while the Spark stayed under 100W. In an Indian office, where a UPS (the battery box that rides out power cuts) is part of any serious setup, that gap decides how many machines one backup can carry.
Software on these machines moves quickly, so read the table as a January 2026 snapshot.
| Pick | Best for | Keep in mind |
|---|---|---|
| RTX 5090 workstation | Models under 32GB, fastest chat speed, image and video generation, Windows software, rendering and gaming alongside AI | A 70B model will not load on a single card |
| DGX Spark 64GB | Models between 32GB and roughly 55GB such as 4-bit 70B class, agent development, fine-tuning mid-size models, a quiet low-power always-on box, room to add a second unit | Small models run slower than on the 5090, and the platform is Arm-based Linux |
| DGX Spark 128GB | 90B and larger models, 120B fine-tuning, long contexts, a single-box plan | Same bandwidth as the 64GB, so extra memory adds capacity rather than speed |
If your daily model sits at 27B to 32B and fits in 32GB, look at the 5090 build first. Teams that cannot send data to a cloud API, such as customer records under India’s DPDP Act, client IP or internal code, get a self-contained box from either Spark.
Yes. ProX PC will have the DGX Spark 64GB available when it launches. We also build RTX 5090 workstations (Pro Maven), so our advice has no reason to lean one way. Tell us which models you plan to run and we will say which option fits, using the numbers from our own DGX Spark vs RTX 5090 testing. Onsite deployment is available across India, and unusual requirements are welcome.
Call 011-40727769 or email sales@proxpc.com to book.
Memory size. Both use the GB10 chip with 273 GB/s of bandwidth and ConnectX-7 networking.
Yes. They link over ConnectX-7 to pool 128GB, and NVIDIA’s slide lists 1.7x relative performance for two units.
A 4-bit 70B model typically needs 40 to 47GB, so it should load with some room left. Our 128GB unit ran 70B and 72B models at about 4.6 tokens per second, and we expect similar speeds on the 64GB because the bandwidth is identical.
No. It runs DGX OS, an Ubuntu-based Linux on Arm.
Resources you may find helpful.

AI is revolutionizing the hardware industry by boosting design, manufacturing, maintenance, supply chains, personalization, autonomy, security, and energy efficiency.

Discover model-inferencing workstations: advanced AI systems revolutionizing industries from retail to smart cities and mastering complex tasks with cutting-edge technology.

Explore how hybrid systems revolutionize transportation, energy, manufacturing, and daily life by integrating diverse technologies for innovative solutions.

Delve into Edge AI's impact on latency, privacy, and autonomy. Uncover its applications in healthcare, manufacturing, retail, and security.