By MiniPC.guru · MiniPC.guru
- Source published
- Sep 25, 2026
- Test performed
- Sep 25, 2026
- Added to MiniPC.guru
- Sep 25, 2026
- Source reviewed
- Sep 25, 2026
- Tested hardware
- NVIDIA GB10, 20 Arm cores (10 × Cortex-X925 + 10 × Cortex-A725) · NVIDIA GB10 integrated Blackwell GPU · 128 GB LPDDR5X unified memory, 121.6 GiB visible to Linux · not applicable: the model is fully resident in memory
- Software / version / OS
- vLLM OpenAI-compatible server, streaming; Python client on the same machine · vLLM 0.28.0; NVIDIA driver 580.178.04; CUDA 13.0 · DGX OS 7.6.0 (Ubuntu 24.04.5 LTS), kernel 7.0.0-1019-nvidia
- Power profile
- default DGX OS power settings, no manual limit; GPU power logged with nvidia-smi in the Our test section
- Test conditions
- nvidia/Qwen3.6-35B-A3B-NVFP4 (35B parameters, about 3B active, 4-bit NVFP4 weights); context window 131,072 tokens; thinking disabled; no speculative decoding; temperature 0.3; batch 1 (one request); 33-token prompt; 301, 305, 301 tokens generated in three runs (77.94, 77.95, 77.95 tokens/s, median shown); median time to first token 0.07 s (0.16 s on the cold first run)
Our own measurement on one unit. Other units, firmware or software versions can give different results.



