Tuning Local AI for Hermes on DGX Spark

[Update 9/15/2026 - added the September 13 recipe change and updated the tables. Flash-Next is 17-23% faster.]

Back in August I wrote about why we keep a local LLM standing by for incident response (here), after Hugging Face couldn't get the American frontier models to read their own attacker logs. This is the follow-up: what it actually took to keep that box running for three weeks, and where the configuration landed. The first half is for CISOs and business owners. The second half is the recipe for whoever has to build it.

TL;DR, the tuning was the easy part. The same $4,000 box went from 11 tok/s to 84 tok/s with no hardware changes, all from settings. Keeping it running was the hard part. The model server locked up about thirty times in three weeks while its health check said "healthy," and the root cause turned out to be a bug in a vendor library that was already documented on GitHub. I spent a week chasing the wrong thing before I read the bug tracker.

What happened (for CISOs)

The setup. One NVIDIA DGX Spark (an ASUS Ascent GX10, about $4,000, 128 GB of unified memory) running Hermes, an open-source agent framework that talks to a locally hosted model over the OpenAI API format. Nothing leaves the network. We use it for SOC work that the frontier models refuse to help with.

The performance. This is the part everyone asks about, so here it is:

Date Model Speed (tok/s) What changed
Aug 15 Qwen3.8-27B (NVFP4) 11.3 Out of the box
Aug 16 Qwen3.8-27B (NVFP4) ~20 One community tuning guide
Aug 27 Ornith-1.5-35B 75-79 Different model, tuned
Sep 5 Ornith-1.5-35B 84 (production median, n=214) Attention backend + draft depth
Sep 9 Qwen3.8-Flash-Next 45 idle, 31 production median Bigger, slower, much smarter model
Sep 13 Qwen3.8-Flash-Next 56 idle, 36.5 production median Smaller draft vocabulary (one line in .env)

Every number moved because of settings, not hardware. I said that in August and it held up.

The problem. Starting August 19, the model server would stop producing output while still answering the health check with HTTP 200. Agents just sat there waiting. Sometimes it crashed outright with a CUDA "illegal memory access" error instead. Between August 19 and September 8 my incident log shows about 30 restarts.

What I tried. I gave myself three attempts, each changing one setting, judged on how long the server stayed up under real traffic:

  1. Change the attention backend. Went from 4 failures in 26 hours to 43 hours, then 24 hours between failures. Better, not fixed.
  2. Turn off speculative decoding. I was sure this was the cause. It was worse: 4 restarts in 72 hours, and the exact same crash happened with that feature completely disabled. So it wasn't the cause.
  3. Turn off prefix caching. Never got a verdict, because the GPU itself faulted (Xid 119) and took down the test. Twice.

What it actually was. Two different problems that looked identical from the outside:

  • The GPU dropping off the bus (an Xid 119 GSP timeout, a driver/hardware fault). Only a reboot fixes that. My auto-restart logic restarted the container 490 times into a dead GPU one morning before I noticed. We unplugged the HDMI cable, which is the leading theory for the trigger, and it has not happened since.
  • A memory corruption bug in FlashInfer, a math library that vLLM uses, when CUDA graphs are enabled. That is vllm issue #52540. Another DGX Spark owner had already posted the same crash signature (Xid 31 MMU fault at address 0) on the same library version I was running. The workaround is one environment variable.

I found it by reading four GitHub issues against my own logs, not by changing settings. In my opinion that is the lesson: before you tune anything, search the vLLM and FlashInfer issue trackers for your GPU model and your error string. Somebody has probably been there.

My monitor made things worse. Twice, my watchdog restarted a perfectly healthy server. The new model takes about 36 seconds to answer because it "thinks" before it responds, and when the box was busy my test question timed out and the watchdog decided the server was dead. It wasn't. It was producing tokens the whole time. I had to teach the watchdog to check the server's own token counter before restarting anything.

What this costs. On paper the hardware pays for itself in under eight months compared to API tokens. What is not on paper is three weeks of my time, one of them on the wrong problem. As a business owner, if you plan to replace an API bill with a box under someone's desk, the engineer who keeps it running is the real cost. That is why we put local AI at Level 6 on our AI Maturity Ladder. It is not a beginner activity.

Should you build one? Yes, one, for the emergency I wrote about in August. Then treat it like production: monitor it with something that does not depend on the model being up, save the logs before you restart it, and change one thing at a time.

The recipe (for whoever builds it)

The model

I started on Ornith-1.5-35B and switched to Qwen3.8-Flash-Next on September 8. In August I said Flash-Next was too big for this box (99 GB checkpoint). I was wrong. About 27 GB of that is a lookup table that only gets read a few hundred bytes at a time, and this recipe from MiaAI Lab memory-maps it from the SSD instead of loading it into the GPU. Resident weights drop to 72 GB. It fits, with room for a 1M token KV cache.

Ornith-1.5-35B Flash-Next, Sep 9 Flash-Next, Sep 15
Weights on GPU 22 GB 72 GB (+27 GB memory-mapped) same
Speed, single user, idle (code / prose) 78-84 tok/s 45 / 31 tok/s 56 / 35 tok/s
Speed, production, 1 request (median) 84 tok/s 31 tok/s 36.5 tok/s
Speed, production, 4 requests (total) not measured 78 tok/s 84 tok/s
Time to first token (production) 0.9 s 2.5-5 s median 40-80 s median (see below)
Free host memory ~60 GB 15-17 GB ~16 GB median, 11 GB lowest
Startup time ~4 min ~12 min ~12 min
Uptime 18-43 h between failures 21 h when I published 5.7 days, no hangs or crashes

Flash-Next is about half the speed per token and it needs a minute to think. It is also a much stronger model with a 5x larger context cache and working image input. For an agent that spends most of its time running tools, that trade has been fine. For a chatbot a human is watching, I would stay on the faster model.

The one bad number is time to first token, and that's capacity. For 22 hours on September 14 the agents kept all 4 request slots full with more waiting in line. Today, on the same settings, the queue is empty. If you run a team of agents in parallel, plan for one Spark to back up.

The recipe

I run the MiaAI Lab repo as-is: ./download.sh, then ./start.sh. Every setting in .env is at the author's default except one line (below). The important thing the recipe does that I did not do in August: it caps GPU memory from the host side, runs a memory watchdog next to the container, and refuses to start if the box is already too full. That is exactly what put my box into swap in August. Read the README before changing anything, the author measured every default.

The one line I changed

EXTRA_DOCKER_ARGS="-e VLLM_USE_V2_MODEL_RUNNER=1 \
  -e VLLM_USE_FLASHINFER_SAMPLER=0 \
  -e VLLM_DISABLED_KERNELS=FlashInferFP8ScaledMMLinearKernel \
  -e CUDA_ENABLE_COREDUMP_ON_EXCEPTION=1 \
  -e CUDA_COREDUMP_GENERATION_FLAGS=skip_global_memory \
  -e CUDA_COREDUMP_FILE=/root/.cache/vllm/coredumps/core_%h_%p_%t.nvcudmp"

Why each one:

  1. VLLM_DISABLED_KERNELS=FlashInferFP8ScaledMMLinearKernel is the fix from vllm#52540. It keeps vLLM off the cuDNN FP8 code path that corrupts memory under CUDA graphs. The real fix (flashinfer#4553) is still not merged as of this writing, so the workaround is what you get. Flash-Next does not actually use this kernel (its attention layers are a different format that goes through CUTLASS), so on this model it is insurance. On Ornith it was the bug.
  2. VLLM_USE_FLASHINFER_SAMPLER=0 comes from vllm#49203, a different GB10 hang with the same symptoms (96% GPU, 17 watts, zero output). I ruled it out for my case because my library already had that fix and hung anyway. Costs nothing, so I left it on.
  3. The three CUDA_COREDUMP lines make the next crash write a dump that names the kernel that actually faulted, instead of the Python frame that happened to report it. Every one of my crash logs blamed the wrong frame. The dump path is on a mounted folder because the launch script deletes the container on restart.

To check if you are affected on any model: docker logs <container> | grep FlashInferFP8ScaledMMLinearKernel. If that comes back with a "Selected" line and you have CUDA graphs on with FlashInfer 0.6.18 or older, disable the kernel.

The update (September 13)

Flash-Next guesses three tokens ahead with a small built-in draft layer, and the full model checks every guess. The recipe's author has since made a smaller draft vocabulary the default (here): the draft scores 47,149 common English and code tokens instead of all 248,320. Each guess is cheaper and quality doesn't change. The freed memory went to the KV cache (1.17M to 1.28M tokens).

If you clone today you already have it. On an older commit, copy the vocab file in and add one line to .env, with the full path (older start.sh doesn't resolve relative ones):

MTP_DRAFT_VOCAB=/home/<you>/Qwen3.8-Flash-Next-Single-DGX-Spark/files/draft_vocab_en_code_47k.txt

Pause your watchdog before a planned restart, or its next check lands mid-startup and launches a second copy.

The monitor

This is what my watchdog does now, every 5 minutes. It is a shell script in the Hermes cron.

  1. Send a real chat completion with max_tokens: 512 and require finish_reason: "stop". Do not trust /health, it returns 200 on a frozen engine.
  2. If that fails, read vllm:generation_tokens_total from /metrics, wait 30 seconds, read it again. If it went up, the server is busy, not dead. Leave it alone. This one change would have prevented both of my false restarts.
  3. If it is really stuck, check the GPU first: if cuInit() returns NO_DEVICE, the card is gone. Send one alert and stop. Do not restart into a dead GPU.
  4. Save the container log, the watchdog log, free -g, and nvidia-smi to an incident file BEFORE restarting. I lost the definitive evidence to an early restart once.
  5. Restart through the launch script, not docker restart. Docker's restart policy reuses the container's original settings and never re-reads your .env. I ran a whole day of "fixed" restarts that were actually the old config. The incident file now records the environment from the running container so I can tell the difference.

Status

As of September 15: 5.7 days with no hangs, crashes, GPU faults or core dumps, well past Ornith's 18-43 hours. The memory watchdog stopped the server once, the afternoon I published, when free memory dropped under its floor (that's its job), and my monitor brought it back. The only other restart was mine, for the change above. Flash-Next doesn't use the kernel behind the Ornith bug, so this doesn't prove the fix, but the box is stable.

Disclaimer: this is what worked on one box. Test in a lab before you put it in front of a SOC.

Need help building or securing a local AI capability? Email us at Hello At PatriotConsultingTech.com

Sources and references:

About the author. Joe Stocker is the Founder and Chief Technical Officer of Patriot Consulting, a former Microsoft Security MVP (2020-2026), and author of the book "Securing Microsoft 365." In his spare time Joe volunteers with several Microsoft programs including Microsoft's Defending Democracy (AccountGuard) program, Microsoft Tech for Social Impact, and the Microsoft Software and Systems Academy (MSSA), which serves military service members desiring to transition into the civilian workspace.

Note: The author created this article with assistance from AI. Learn more