The Honest Guide To Local AI Supercomputers (NVIDIA DGX Spark)

Por Zen van Riel · 21 ago 2026 · 17:21

Visualizaciones
46.5K vistas
Likes
599 likes
Comentarios
110 comentarios

🎁 Get My Free Local AI Config: https://zenvanriel.com/ai-coding ⚡ Work directly with me to become a high-earning AI Engineer: https://aiengineer.community/join On this little machine I run the latest Qwen models, including one with over a hundred billion parameters. This is the NVIDIA DGX Spark with 128GB of unified memory, and all of it was impossible on local hardware a few years ago. I go through the whole box, from the chassis and the 250 watt USB-C power delivery to my full local Claude Code workflow, where a 30B Qwen coder does the day to day work and calls a 122B Qwen model when a question is too hard for it. I also cover the real constraints of this setup: how two models split about 93GB of memory at once, why local sub-agents queue instead of running in parallel, and the one design choice on this machine that annoyed me. What You'll Learn - What the NVIDIA DGX Spark actually is: 128GB unified memory, 250W over a single USB-C cable, and a fan that does kick in - How to use LM Studio and LM Link to run Spark models from a MacBook that only has 48GB of memory - Why local inference slows down as the context window fills, tested live instead of on a toy prompt - How to point Claude Code at a local model using .claude/settings.json and LM Studio on localhost:1234 - Why Claude Code spends over 32,000 tokens of context before you type your first message - How Claude Code auto mode makes slow local models practical for hour-long tasks - How a 30B Qwen coder hands hard questions to a 122B Qwen model with one command - How a 30B coder at 160K context and a 122B model at 130K context share about 93GB of memory - Why local sub-agents get queued and batched instead of running in parallel like they do in the cloud - How to put the machine behind Tailscale and reach Open WebUI from your phone over 5G - DGX Spark vs RTX 5090: 128GB of unified memory against 32GB of VRAM, and which one I travel with Timestamps 0:00 A mini supercomputer that runs 100B+ models locally 0:28 Chassis, cooling, and 250W over one USB-C cable 1:39 Ports, and why the HDMI output stays unused 2:07 Using Spark models on a MacBook with LM Link 3:35 What happens to speed when the context window fills 4:25 My tmux terminal control plane 5:32 Pointing Claude Code at a local model on the Spark 7:14 Auto mode builds a Python server over Tailscale 8:53 Calling a 122B big brain model for hard questions 10:17 How two models share 128GB of memory 12:12 Sovereignty, cost, and the career case for local AI 12:55 Securing the Spark with Tailscale 14:52 DGX Spark vs RTX 5090, and what it costs 16:33 The one thing that annoyed me Video references: - My FREE local AI projects, they run on a normal CPU too: https://zenvanriel.com/open-source - LM Studio: https://lmstudio.ai/ - LM Link (use your local models remotely): https://lmstudio.ai/docs/lmlink - Tailscale: https://tailscale.com/ - Open WebUI: https://openwebui.com/ - My local AI coding masterclass: https://youtu.be/rp5EwOogWEw Why I Made This Video Most DGX Spark coverage stops at benchmark charts. I wanted to show what the machine is like to actually work on, including the local Claude Code setup I use on it every day, the memory trade-off behind running two Qwen models at the same time, and the parts of the box I would change. My free local AI projects in the references above run on a normal CPU too, so you can start without a DGX Spark. NVIDIA provided me with the DGX Spark: no paid sponsorship :) #localai #dgxspark #nvidia #claudecode #lmstudio #qwen #localllm #aicoding #selfhostedai #tailscale #aiengineering #localclaudecode #edgeai Connect LinkedIn: https://www.linkedin.com/in/zen-van-riel Community: https://www.skool.com/ai-engineer Sponsorships & Business Inquiries: business@aiengineer.community