cloudburn.dev

Building a Control Plane for a Single DGX Spark

vLLM and Atlas both worked. Living with them by hand didn't. Notes on building spark-control, the bugs it surfaced, and what a single unified-memory box actually needs to stay reliable.

Previous in this series: DGX Spark: vLLM vs Atlas for Local Inference That post compares the two runtimes head to head. This one is about what happens after you pick: running either of them, plus a third option, day to day without babysitting SSH sessions.


The problem vLLM-vs-Atlas didn’t cover

System: DGX Spark (GB10), Linux 7.0.0-1019-nvidia, NVIDIA Driver 595.91.07

Picking a runtime was the easy part. Running one on a single box, long enough for it to matter, surfaced a different class of problem entirely: nothing about a single DGX Spark scales down cleanly from “server you SSH into and start a process on.”

One GPU. One unified memory pool. One model loaded at a time, realistically. But also: a daily-driver coding model, an occasional deep-reasoning model, a lightweight batch worker, and an agent (Hermes) that needs to know right now which one is actually up and what context window it’s serving at. Doing that by hand : docker ps, tail -f over SSH, manually editing Hermes’s config file every time a model swap happened : was fine for a week. It did not survive contact with actually using the thing daily.

So: a small control plane. spark-control : a FastAPI backend plus a plain-JS frontend, no framework, living on the Spark itself, one page for everything: which model is running, hot-swap between named profiles, live log tail without opening a terminal, a catalog of every model that’s been evaluated with real benchmark numbers attached, and a background loop that keeps Hermes’s own config in sync with whatever’s actually loaded instead of whatever it was last told.

That last part turned out to be the interesting problem. Getting it right meant finding : and fixing : several bugs that only show up when a model server has been left running unattended for more than a few hours.


What it does

  • Profiles, not commands. Each usable model is a named profile (hermes-fast, qwen35-122b-spark, coder-next, batch-worker…). Activating one tears down whatever’s running and boots the target, whether that’s a raw docker run, a docker-compose stack, or a sparkrun recipe : the profile hides which.
  • Live status, not a snapshot. The UI polls a real health probe against the active container, not just “is a container object present.” A container can be up and still not answering : that distinction turned out to matter a lot (see below).
  • Streamed logs, held open. One long-lived SSH subprocess (tail -F or docker logs -f depending on runtime) piped to the browser over SSE, instead of poll-and-recat every few seconds. Cheaper, and doesn’t produce the scroll-jank of re-fetching the whole tail on every refresh.
  • Hermes auto-sync. A background thread reconciles Hermes’s config.yaml against whatever model is actually live : provider URL, model id, context length : on a timer, and once at startup. Hermes no longer needs to be told by hand that the model underneath it changed.
  • A model catalog with real numbers, not vendor marketing: measured VRAM footprint, measured tok/s, which agent scaffolds (NemoClaw, Ironclaw) it’s been verified compatible with, and dated notes on anything that broke while testing it.

Here’s what the control plane looks like in practice:

Spark Control Dashboard

The dashboard showing Qwen3.5-122B running via vLLM with 260K context, Hermes agent synced, and all the real-time status you’d expect from something that’s been through enough 3am incidents to know what matters. Click the image to view full size.


A thought on open-sourcing

I keep circling back to this: should I open-source spark-control?

The DGX Spark ecosystem is still small. There’s no “standard” way to manage models on this hardware yet. Everyone’s cobbling together their own scripts, their own dashboards, their own reconciliation logic. I’ve been there : the first version of this was just a bash script with a case statement and a while true loop.

But here’s the thing: the bugs that matter aren’t the ones you find in a quick test. They’re the ones that show up when you leave the box running unattended for a week. They’re the race conditions that kill a healthy model because a health check timed out. They’re the floating image tags that silently swap your working deployment for a broken one.

If I open-source this, it wouldn’t be the polished, production-ready solution you’d want to deploy at a datacenter. It’d be the “here’s what I learned the hard way” version. The code that knows exactly how many times you need to check if something’s actually running before you’re allowed to kill it. The catalog that tracks not just what works, but what broke and why.

Maybe that’s useful to other DGX tinkerers. Maybe it’s not. Either way, the decision’s still open. Drop me a line if you’re building something similar : curious what problems you’re hitting.


Bugs that only showed up under real, unattended use

None of these were visible in a quick manual test. Each one needed the system left alone for hours, or hit with real traffic, to surface.

1. A restart race that killed a healthy model

check_default_boot_profile() : the background job that makes sure something is always serving : took a single status read, and if it looked like nothing was active, killed whatever container it found and booted the configured default over it. A UI-triggered restart at 22:04 landed mid-read: the read concluded nothing was active (it was mid-restart, not actually down), killed a perfectly healthy 122B model (Exited 137, unattended, overnight), and booted the configured default in its place : which itself failed to load (see bug 3), leaving nothing running for hours.

Fix: two independent status reads, spaced past the status-cache TTL, have to agree before the reconciler is allowed to kill anything.

2. The same race, worse: killed a model mid-conversation

Even after fix #1, a second variant showed up: the health probe itself could time out under real load : a container busy serving an actual request just doesn’t answer a liveness ping fast enough : and the reconciler read that as “not running” and killed it. This one fired while genuinely serving a live chat request. Confirmed after the fact from the container logs: POST /v1/chat/completions from an active client seconds before the kill.

Fix: before killing anything, check docker ps directly for a running container matching the expected name pattern. A slow health check is not the same as no container.

3. A floating recipe tag silently changed the running image

The fastest model in the catalog (75.7 tok/s, an Intel AutoRound int4 build) failed to reload one day with:

ValueError: There is no module or parameter named layers.0.mlp.gate.qweight
in Qwen3NextModel. The available parameters belonging to layers.0.mlp.gate
(GateLinear) are: {layers.0.mlp.gate.weight}

Nothing changed on our side. Root cause: the sparkrun recipe alias it deploys from (@official/eugr) isn’t pinned to a release : it rebuilt the underlying image between the last working run and this one, landing a dev/nightly vLLM snapshot that mis-loads this quant format. Cross-referenced against Intel’s own AutoRound issue tracker (intel/auto-round#1776) and an in-flight vLLM fix (vllm-project/vllm#35261) : this is a known, still-open upstream gap, not a problem with the checkpoint. No CLI flag works around it; the fix has to land in vLLM’s loader. Confirmed deterministic on a second independent attempt, not a timing fluke.

Mitigation: switched the automatic-recovery default away from this profile so an unattended restart doesn’t keep retrying something that’s guaranteed to fail. Real fix is pinning the recipe to a release once one exists past that patch.

4. An SSH process leak from unheld log-tail subprocesses

The old poll-and-recat log viewer spawned a fresh docker logs over SSH on every poll interval without reliably killing the previous one. Left running across a few days of dashboard-open-in-a-tab, that grew to roughly 1,500 orphaned processes, close to exhausting the box’s sshd connection limit. Fixed by holding one long-lived subprocess per stream and explicitly killing any prior one before starting a new one : caught, in the process, a pkill pattern that was matching its own argv and killing itself before it could kill its target (fixed with the bracket-escape trick, [d]ocker logs -f, so the pattern in the process list doesn’t match the pattern in the pkill command’s own command line).

5. Catalog metadata lying about what’s actually downloaded

A speculative-decode variant showed as not-downloaded in the UI despite being fully deployed and in daily use : because the catalog’s id field, meant to point at the real HuggingFace repo for cache-path lookups, had been set to a placeholder. Fixed by splitting one overloaded field into three: the real HF repo id (cache lookups), what the running server actually reports itself as (served_as, for “is this the active model” comparisons), and the profile name the hotswap button targets (profile_model) : three different identities for the same checkpoint that were never interchangeable to begin with.


Benchmarking, honestly

The catalog page shows measured numbers, not claimed ones. Methodology: hermes -z "<prompt>" --usage-file <path> : one-shot mode, no interactive session, no log noise : run against each profile after it’s confirmed live. Tool-calling correctness checked with unpredictable random values baked into the prompt, so a pass can’t be explained by the model guessing or hallucinating a plausible-looking answer.

ModelEngineVRAMtok/sNotes
Qwen3.6-35B-A3B-NVFP4Atlas~22GB~112 (solo)Daily driver. MoE, ~3B active. NemoClaw + Ironclaw compatible.
Qwen3.5-122B-A10B-NVFP4 (“Spark”)vLLM82GB~24 (MTP on)MoE, community single-Spark coding pick. Best-quality single model that fits.
Qwen3.5-122B-A10B-DFlashvLLM82GB~28-31 steady-stateSame checkpoint, block-speculative decode instead of MTP. Reliable, just slower than Coder-Next.
Qwen3-Coder-Next-int4 (AutoRound)sparkrun53GB~75.7Fastest verified. Currently blocked by bug #3 above.
Nemotron-3-Nano-30B-A3B-NVFP4Atlas18GB~60Fast. Clean with Ironclaw; NemoClaw’s own tool-parser chokes on it (malformed XML → loop : parser bug, not a model-quality problem).
Nemotron-3-Super-120B-A12B-NVFP4Atlas::Doesn’t fit. Weights alone ~108GB against a 121.7GB usable pool. Documented and closed, not attempted again.

MoE confirmed the hard way where it mattered, not from the model name: vLLM logs its own resolved architecture class at load (Resolved architecture: Qwen3_5MoeForConditionalGeneration) and picks MoE-specific kernels (CUTLASS NvFp4 MoE backend, FlashInfer MoE backend) : that’s not something you get by accident on a dense model, whatever a model card summary claims.


What a single unified-memory box actually teaches you

  • Unified memory doesn’t fail cleanly. GPU and host RAM are one 121.7GB pool on GB10. Overcommit it and you don’t get a tidy CUDA OOM : you get the kernel OOM-killer picking off whatever process it wants, which on this box has included wpa_supplicant and the telemetry daemon, not just the model server. Tested this directly by pushing gpu_memory_utilization past its validated ceiling twice, on purpose, to confirm : both times, real processes died. Don’t raise that number past what’s been tested for a given profile.
  • “Container running” and “container healthy” are different claims, and the gap between them is exactly where automated reconciliation logic tends to do the most damage : see bugs #1 and #2. Anything that’s allowed to kill a container on your behalf needs to check both, not one.
  • A floating alias/tag is silent risk, not convenience. sparkrun’s @registry/recipe shorthand and any :latest image tag both let the thing underneath your deployment change without your config changing at all. Bug #3 is what that looks like when it goes wrong.
  • MoE naming conventions (A10B = 10B active) are a reasonable first signal, but they’re not proof. Trust the runtime’s own resolved architecture and kernel selection over a model card summary : see the confirmation above.

Open items

  • The AutoRound MoE-gate bug (#3) needs an upstream vLLM build past PR #35261 before Coder-Next is safe to run unattended again.
  • End-to-end validation of Ironclaw and NemoClaw against the current profile set is still outstanding : most of the work so far has verified the model layer, not the full agent scaffold on top of it.
  • A GSP firmware wedge earlier in the box’s life hasn’t recurred since a firmware update and reboot, but it’s too soon to call that closed with confidence.

This is the control-plane layer sitting on top of the runtime choice from the previous post. Same box, same 128GB pool : the difference is whether you find out about a problem from the dashboard or from Hermes going quiet at 3am.