Tuning Local AI for Hermes on DGX Spark
[Update 9/15/2026 - added the September 13 recipe change and updated the tables. Flash-Next is 17-23% faster.]
Back in August I wrote about why we keep a local LLM standing by for incident response (here), after Hugging Face couldn't get the American frontier models to read their own attacker logs. This is the follow-up: what it actually took to keep that box running for three weeks, and where the configuration landed. The first half is for CISOs and business owners. The second half is the recipe for whoever has to build it.
TL;DR, the tuning was the easy part. The same $4,000 box went from 11 tok/s to 84 tok/s with no hardware changes, all from settings. Keeping it running was the hard part. The model server locked up about thirty times in three weeks while its health check said "healthy," and the root cause turned out to be a bug in a vendor library that was already documented on GitHub. I spent a week chasing the wrong thing before I read the bug tracker.
What happened (for CISOs)
The setup. One NVIDIA DGX Spark (an ASUS Ascent GX10, about $4,000, 128 GB of unified memory) running Hermes, an open-source agent framework that talks to a locally hosted model over the OpenAI API format. Nothing leaves the network. We use it for SOC work that the frontier models refuse to help with.
The performance. This is the part everyone asks about, so here it is:
| Date | Model | Speed (tok/s) | What changed |
|---|---|---|---|
| Aug 15 | Qwen3.8-27B (NVFP4) | 11.3 | Out of the box |
| Aug 16 | Qwen3.8-27B (NVFP4) | ~20 | One community tuning guide |
| Aug 27 | Ornith-1.5-35B | 75-79 | Different model, tuned |
| Sep 5 | Ornith-1.5-35B | 84 (production median, n=214) | Attention backend + draft depth |
| Sep 9 | Qwen3.8-Flash-Next | 45 idle, 31 production median | Bigger, slower, much smarter model |
| Sep 13 | Qwen3.8-Flash-Next | 56 idle, 36.5 production median | Smaller draft vocabulary (one line in .env) |
Every number moved because of settings, not hardware. I said that in August and it held up.
The problem. Starting August 19, the model server would stop producing output while still answering the health check with HTTP 200. Agents just sat there waiting. Sometimes it crashed outright with a CUDA "illegal memory access" error instead. Between August 19 and September 8 my incident log shows about 30 restarts.
What I tried. I gave myself three attempts, each changing one setting, judged on how long the server stayed up under real traffic:
- Change the attention backend. Went from 4 failures in 26 hours to 43 hours, then 24 hours between failures. Better, not fixed.
- Turn off speculative decoding. I was sure this was the cause. It was worse: 4 restarts in 72 hours, and the exact same crash happened with that feature completely disabled. So it wasn't the cause.
- Turn off prefix caching. Never got a verdict, because the GPU itself faulted (Xid 119) and took down the test. Twice.
What it actually was. Two different problems that looked identical from the outside:
- The GPU dropping off the bus (an Xid 119 GSP timeout, a driver/hardware fault). Only a reboot fixes that. My auto-restart logic restarted the container 490 times into a dead GPU one morning before I noticed. We unplugged the HDMI cable, which is the leading theory for the trigger, and it has not happened since.
- A memory corruption bug in FlashInfer, a math library that vLLM uses, when CUDA graphs are enabled. That is vllm issue #52540. Another DGX Spark owner had already posted the same crash signature (
Xid 31 MMU fault at address 0) on the same library version I was running. The workaround is one environment variable.
I found it by reading four GitHub issues against my own logs, not by changing settings. In my opinion that is the lesson: before you tune anything, search the vLLM and FlashInfer issue trackers for your GPU model and your error string. Somebody has probably been there.
My monitor made things worse. Twice, my watchdog restarted a perfectly healthy server. The new model takes about 36 seconds to answer because it "thinks" before it responds, and when the box was busy my test question timed out and the watchdog decided the server was dead. It wasn't. It was producing tokens the whole time. I had to teach the watchdog to check the server's own token counter before restarting anything.
What this costs. On paper the hardware pays for itself in under eight months compared to API tokens. What is not on paper is three weeks of my time, one of them on the wrong problem. As a business owner, if you plan to replace an API bill with a box under someone's desk, the engineer who keeps it running is the real cost. That is why we put local AI at Level 6 on our AI Maturity Ladder. It is not a beginner activity.
Should you build one? Yes, one, for the emergency I wrote about in August. Then treat it like production: monitor it with something that does not depend on the model being up, save the logs before you restart it, and change one thing at a time.
The recipe (for whoever builds it)
The model
I started on Ornith-1.5-35B and switched to Qwen3.8-Flash-Next on September 8. In August I said Flash-Next was too big for this box (99 GB checkpoint). I was wrong. About 27 GB of that is a lookup table that only gets read a few hundred bytes at a time, and this recipe from MiaAI Lab memory-maps it from the SSD instead of loading it into the GPU. Resident weights drop to 72 GB. It fits, with room for a 1M token KV cache.
| Ornith-1.5-35B | Flash-Next, Sep 9 | Flash-Next, Sep 15 | |
|---|---|---|---|
| Weights on GPU | 22 GB | 72 GB (+27 GB memory-mapped) | same |
| Speed, single user, idle (code / prose) | 78-84 tok/s | 45 / 31 tok/s | 56 / 35 tok/s |
| Speed, production, 1 request (median) | 84 tok/s | 31 tok/s | 36.5 tok/s |
| Speed, production, 4 requests (total) | not measured | 78 tok/s | 84 tok/s |
| Time to first token (production) | 0.9 s | 2.5-5 s median | 40-80 s median (see below) |
| Free host memory | ~60 GB | 15-17 GB | ~16 GB median, 11 GB lowest |
| Startup time | ~4 min | ~12 min | ~12 min |
| Uptime | 18-43 h between failures | 21 h when I published | 5.7 days, no hangs or crashes |
Flash-Next is about half the speed per token and it needs a minute to think. It is also a much stronger model with a 5x larger context cache and working image input. For an agent that spends most of its time running tools, that trade has been fine. For a chatbot a human is watching, I would stay on the faster model.
The one bad number is time to first token, and that's capacity. For 22 hours on September 14 the agents kept all 4 request slots full with more waiting in line. Today, on the same settings, the queue is empty. If you run a team of agents in parallel, plan for one Spark to back up.
The recipe
I run the MiaAI Lab repo as-is: ./download.sh, then ./start.sh. Every setting in .env is at the author's default except one line (below). The important thing the recipe does that I did not do in August: it caps GPU memory from the host side, runs a memory watchdog next to the container, and refuses to start if the box is already too full. That is exactly what put my box into swap in August. Read the README before changing anything, the author measured every default.
The one line I changed
EXTRA_DOCKER_ARGS="-e VLLM_USE_V2_MODEL_RUNNER=1 \
-e VLLM_USE_FLASHINFER_SAMPLER=0 \
-e VLLM_DISABLED_KERNELS=FlashInferFP8ScaledMMLinearKernel \
-e CUDA_ENABLE_COREDUMP_ON_EXCEPTION=1 \
-e CUDA_COREDUMP_GENERATION_FLAGS=skip_global_memory \
-e CUDA_COREDUMP_FILE=/root/.cache/vllm/coredumps/core_%h_%p_%t.nvcudmp"
Why each one:
VLLM_DISABLED_KERNELS=FlashInferFP8ScaledMMLinearKernelis the fix from vllm#52540. It keeps vLLM off the cuDNN FP8 code path that corrupts memory under CUDA graphs. The real fix (flashinfer#4553) is still not merged as of this writing, so the workaround is what you get. Flash-Next does not actually use this kernel (its attention layers are a different format that goes through CUTLASS), so on this model it is insurance. On Ornith it was the bug.VLLM_USE_FLASHINFER_SAMPLER=0comes from vllm#49203, a different GB10 hang with the same symptoms (96% GPU, 17 watts, zero output). I ruled it out for my case because my library already had that fix and hung anyway. Costs nothing, so I left it on.- The three
CUDA_COREDUMPlines make the next crash write a dump that names the kernel that actually faulted, instead of the Python frame that happened to report it. Every one of my crash logs blamed the wrong frame. The dump path is on a mounted folder because the launch script deletes the container on restart.
To check if you are affected on any model: docker logs <container> | grep FlashInferFP8ScaledMMLinearKernel. If that comes back with a "Selected" line and you have CUDA graphs on with FlashInfer 0.6.18 or older, disable the kernel.
The update (September 13)
Flash-Next guesses three tokens ahead with a small built-in draft layer, and the full model checks every guess. The recipe's author has since made a smaller draft vocabulary the default (here): the draft scores 47,149 common English and code tokens instead of all 248,320. Each guess is cheaper and quality doesn't change. The freed memory went to the KV cache (1.17M to 1.28M tokens).
If you clone today you already have it. On an older commit, copy the vocab file in and add one line to .env, with the full path (older start.sh doesn't resolve relative ones):
MTP_DRAFT_VOCAB=/home/<you>/Qwen3.8-Flash-Next-Single-DGX-Spark/files/draft_vocab_en_code_47k.txt
Pause your watchdog before a planned restart, or its next check lands mid-startup and launches a second copy.
The monitor
This is what my watchdog does now, every 5 minutes. It is a shell script in the Hermes cron.
- Send a real chat completion with
max_tokens: 512and requirefinish_reason: "stop". Do not trust/health, it returns 200 on a frozen engine. - If that fails, read
vllm:generation_tokens_totalfrom/metrics, wait 30 seconds, read it again. If it went up, the server is busy, not dead. Leave it alone. This one change would have prevented both of my false restarts. - If it is really stuck, check the GPU first: if
cuInit()returnsNO_DEVICE, the card is gone. Send one alert and stop. Do not restart into a dead GPU. - Save the container log, the watchdog log,
free -g, andnvidia-smito an incident file BEFORE restarting. I lost the definitive evidence to an early restart once. - Restart through the launch script, not
docker restart. Docker's restart policy reuses the container's original settings and never re-reads your.env. I ran a whole day of "fixed" restarts that were actually the old config. The incident file now records the environment from the running container so I can tell the difference.
Status
As of September 15: 5.7 days with no hangs, crashes, GPU faults or core dumps, well past Ornith's 18-43 hours. The memory watchdog stopped the server once, the afternoon I published, when free memory dropped under its floor (that's its job), and my monitor brought it back. The only other restart was mine, for the change above. Flash-Next doesn't use the kernel behind the Ornith bug, so this doesn't prove the fix, but the box is stable.
Disclaimer: this is what worked on one box. Test in a lab before you put it in front of a SOC.
Need help building or securing a local AI capability? Email us at Hello At PatriotConsultingTech.com
Sources and references:
- FlashInfer cuDNN FP8 memory corruption under CUDA graphs (root cause): https://github.com/vllm-project/vllm/issues/52540; DGX Spark confirmation: https://github.com/vllm-project/vllm/issues/50934#issuecomment-5344337727
- GB10 sampler livelock (same symptoms, different bug): https://github.com/vllm-project/vllm/issues/49203
- Qwen3.8-Flash-Next on a single DGX Spark: https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark; smaller draft vocabulary: https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark/pull/39
- Measured on one DGX Spark (GB10, 119.7 GB usable unified memory), August 19 to September 15, 2026. Flash-Next "idle" is a fixed prompt with nothing else running; "production" is vLLM's own logs and metrics over 8,095 requests before the change and 3,439 after.
About the author. Joe Stocker is the Founder and Chief Technical Officer of Patriot Consulting, a former Microsoft Security MVP (2020-2026), and author of the book "Securing Microsoft 365." In his spare time Joe volunteers with several Microsoft programs including Microsoft's Defending Democracy (AccountGuard) program, Microsoft Tech for Social Impact, and the Microsoft Software and Systems Academy (MSSA), which serves military service members desiring to transition into the civilian workspace.
Note: The author created this article with assistance from AI. Learn more