Changelog
New white papers and new LLM, VLM and CV benchmark results, newest first. Each entry links to what it describes.
October 2026
- 6 OctoberNew LLM benchmarkQwen3.8-27B — RTX PRO 6000Qwen3.8-27B on a single RTX PRO 6000: NVFP4 weights served with vLLM, with MTP speculative decoding (k=3), measured up to 128 concurrent requests.
- 6 OctoberNew LLM benchmarkQwen3.8-27B — RTX PRO 6000Qwen3.8-27B on a single RTX PRO 6000: BF16 weights served with vLLM, with MTP speculative decoding (k=3), measured up to 128 concurrent requests.
- 6 OctoberNew white paperThe First Open-Weight Alternatives to JevSix open-weight typed-decision models compared on one NVIDIA DGX Spark (GB10) through a shared state + questions interface and one 25-question set: accuracy, calibration, option-order sensitivity and latency — and which model fits which job.
- 6 OctoberNew white paperVLM Inference Benchmark ExplorerExplore vision-language model inference benchmarks on RTX PRO 6000 Blackwell, Jetson AGX Orin and Jetson Orin NX. Pick the image size, set how long a camera may wait for its answer, and see how many cameras each configuration keeps up with.
- 1 OctoberNew LLM benchmarkGemma-4-31B-it — ThorGemma-4-31B-it on Jetson AGX Thor: NVFP4 weights served with vLLM, with MTP speculative decoding (k=3) and a 256K context.
- 1 OctoberNew LLM benchmarkQwen3-VL-30B-A3B-Instruct — ThorQwen3-VL-30B-A3B-Instruct, a vision-language model, on Jetson AGX Thor: FP8 weights served with vLLM, with speculative decoding and a 256K context.
- 1 OctoberNew LLM benchmarkQwen3.8-27B — RTX PRO 6000Qwen3.8-27B on a single RTX PRO 6000: FP8 weights served with vLLM, with MTP speculative decoding (k=3), measured up to 128 concurrent requests.
- 1 OctoberNew LLM benchmarkMiMo-V2.6-Pro-MOPD — DGX B300MiMo-V2.6-Pro-MOPD (1.02T) on DGX B300: FP4 weights served with vLLM on four GPUs (TP4), with DFlash speculative decoding (k=7) and a 1M-token context, measured up to 256 concurrent requests.
- 1 OctoberNew LLM benchmarkMiMo-V2.6-Flash-RL — 2× DGX SparkMiMo-V2.6-Flash-RL (310B) on 2× DGX Spark: FP4 weights served with vLLM across both nodes (TP2), with DFlash speculative decoding (k=7) and a 256K context.
- 1 OctoberNew LLM benchmarkMiMo-V2.6-Flash-RL — 4× DGX SparkMiMo-V2.6-Flash-RL (310B) on 4× DGX Spark: FP4 weights served with vLLM across four nodes (TP4), with DFlash speculative decoding (k=7) and a 500K context.
September 2026
- 29 SeptemberNew white paperEngram Offloading in LLM InferenceEngram tables are large but read only a few rows per token, so they can move from GPU memory to host RAM or NVMe. What that means on DGX Spark, RTX PRO 6000 and DGX B300, and measured results with DeepSeek-V4.1-Flash and Qwen3.8-Flash-Next.
- 28 SeptemberNew LLM benchmarkQwen3.8-Flash-Next — 1× DGX SparkQwen3.8-Flash-Next (176B) on a single DGX Spark: NVFP4 weights served with SGLang, with MTP speculative decoding (k=3) and its 47.7 GiB n-gram embedding table read from NVMe.
- 28 SeptemberNew LLM benchmarkQwen3.8-Flash-Next — RTX PRO 6000Qwen3.8-Flash-Next (176B) on a single RTX PRO 6000: NVFP4 weights served with SGLang, with MTP speculative decoding (k=3) and its 47.7 GiB n-gram embedding table held in host RAM.
- 28 SeptemberNew LLM benchmarkMiMo-V2.6-Pro-RL — DGX B300MiMo-V2.6-Pro-RL (1.02T) on DGX B300: FP4 weights served with vLLM on four GPUs (TP4), with DFlash speculative decoding (k=7) and a 1M-token context.
- 18 SeptemberNew CV benchmarkPeopleNet INT8 — Jetson AGX ThorPeopleNet person detection (INT8, 960×544 input) on Jetson AGX Thor: frames per second per camera, measured from 1 to 32 simultaneous cameras.
- 18 SeptemberNew CV benchmarkYOLO11-N — Jetson AGX ThorYOLO11-N object detection (640×640 input) on Jetson AGX Thor: frames per second per camera, measured from 1 to 30 simultaneous cameras.
- 18 SeptemberNew CV benchmarkYOLO11S — Jetson AGX ThorYOLO11S object detection (640×640 input) on Jetson AGX Thor: frames per second per camera, measured from 1 to 32 simultaneous cameras.
- 18 SeptemberNew CV benchmarkYOLO11M — Jetson AGX ThorYOLO11M object detection (640×640 input) on Jetson AGX Thor: frames per second per camera, measured from 1 to 32 simultaneous cameras.
- 18 SeptemberNew CV benchmarkYOLO11L — Jetson AGX ThorYOLO11L object detection (640×640 input) on Jetson AGX Thor: frames per second per camera, measured from 1 to 32 simultaneous cameras.
- 17 SeptemberNew white paperDeepSeek-V4.1-Flash 8× DGX Spark TP8 DeploymentDeployment of DeepSeek-V4.1-Flash (763B MoE, FP8, DSpark k=5) on 8× NVIDIA DGX Spark (GB10) with TP8: two configurations (300K Engram-in-memory, 1M Engram-on-disk), NCCL optimization, benchmark results and TP4 comparison.
- 17 SeptemberNew LLM benchmarkDeepSeek-V4.1-Flash — 8× DGX SparkDeepSeek-V4.1-Flash (763B) on 8× DGX Spark: FP8 weights served with vLLM across eight nodes (TP8), with DSpark speculative decoding (k=5) and a 300K context, the Engram table kept in memory.
- 17 SeptemberNew LLM benchmarkDeepSeek-V4.1-Flash — 8× DGX SparkDeepSeek-V4.1-Flash (763B) on 8× DGX Spark: the same TP8 setup with a 1M-token context, the Engram table moved to disk.
- 16 SeptemberNew white paperDeepSeek-V4.1-Flash 4× DGX Spark DeploymentDeployment of DeepSeek-V4.1-Flash (763B MoE, FP8, DSpark k=5) on 4× NVIDIA DGX Spark (GB10) with tensor parallelism: vLLM build chain, 7 SM 12.1a patches, Engram-on-disk, benchmark results and B300 comparison.
- 16 SeptemberNew LLM benchmarkDeepSeek-V4.1-Flash — 4× DGX SparkDeepSeek-V4.1-Flash (763B) on 4× DGX Spark: FP8 weights served with vLLM across four nodes (TP4), with DSpark speculative decoding (k=5) and a 300K context.
- 16 SeptemberNew CV benchmarkPeopleNet INT8 — GB10PeopleNet person detection (INT8, 960×544 input) on GB10 (DGX Spark): frames per second per camera, measured from 1 to 32 simultaneous cameras.
- 16 SeptemberNew CV benchmarkPeopleNet INT8 — Jetson Orin NanoPeopleNet person detection (INT8, 960×544 input) on Jetson Orin Nano: frames per second per camera, measured from 1 to 8 simultaneous cameras.
- 16 SeptemberNew CV benchmarkPeopleNet INT8 — RTX 3060PeopleNet person detection (INT8, 960×544 input) on RTX 3060: frames per second per camera, measured from 1 to 16 simultaneous cameras.
- 16 SeptemberNew CV benchmarkPeopleNet INT8 — RTX 3090PeopleNet person detection (INT8, 960×544 input) on RTX 3090: frames per second per camera, measured from 1 to 30 simultaneous cameras.
- 16 SeptemberNew CV benchmarkYOLO11-N — RTX 3090YOLO11-N object detection (640×640 input) on RTX 3090: frames per second per camera, measured from 1 to 28 simultaneous cameras.
- 14 SeptemberNew LLM benchmarkDeepSeek-V4.1-Flash — DGX B300DeepSeek-V4.1-Flash (763B) on DGX B300: FP8 weights served with vLLM on four GPUs (TP4), with DSpark speculative decoding (k=3) and a 1M-token context.
- 14 SeptemberNew LLM benchmarkTencent-Hy4-preview — DGX B300Tencent-Hy4-preview (780B) on DGX B300: BF16 weights served with vLLM on all eight GPUs (TP8), with MTP speculative decoding (k=3) and a 512K context.
- 14 SeptemberNew LLM benchmarkDeepSeek-V4-Flash-Vision-Exp — 2× DGX SparkDeepSeek-V4-Flash-Vision-Exp (305B) on 2× DGX Spark: FP8 weights served with vLLM across both nodes (TP2), with DSpark speculative decoding (k=3) and a 131K context.
- 4 SeptemberNew LLM benchmarkGLM-5.3 — 8× DGX SparkGLM-5.3 (753B) on 8× DGX Spark: NVFP4 weights served with vLLM across eight nodes (TP8), with the model's built-in MTP speculative decoding (k=4) and a 256K context.
- 4 SeptemberNew LLM benchmarkMuse-Glimmer-30B — RTX PRO 6000Muse-Glimmer-30B on a single RTX PRO 6000: BF16 weights served with vLLM, with DFlash speculative decoding through the model's native MTP (k=15).
- 1 SeptemberNew LLM benchmarkGLM-5.3 — DGX B300GLM-5.3 (753B) on DGX B300: FP8 weights served with SGLang on four GPUs (TP4), with a 1M-token context.
August 2026
- 3 AugustNew white paperNVIDIA DGX B300 vs GB300 NVL72 Cluster Architecture ComparisonTechnical comparison of two NVIDIA Blackwell Ultra architectures, DGX B300 and GB300 NVL72: system design, scaling approach, network fabric, power and cooling, and which workloads suit which platform.
July 2026
- 30 JulyNew white paperKimi K3 Inference Benchmark on DGX-B300Performance evaluation of Moonshot AI Kimi K3 (2.8T MoE, MXFP4) on NVIDIA DGX-B300 (8x Blackwell Ultra, TP=8): vLLM vs SGLang, direct vs DSpark speculative decoding, with SLO-driven capacity planning.
- 30 JulyNew white paperDGX Spark 3-Node AI Cluster Setup GuideRing (mesh) topology AI cluster setup with 3 NVIDIA DGX Spark nodes: management and compute networks, RoCEv2/RDMA, sparkrun configuration.
- 30 JulyNew white paperDGX Spark 2-Node AI Cluster Setup GuidePoint-to-point topology AI cluster setup with 2 NVIDIA DGX Spark nodes: management and compute networks, RoCEv2/RDMA, sparkrun configuration.
- 24 JulyNew white paperDGX Spark 8-Node AI Cluster Setup GuideSwitch-based AI cluster setup with 8 NVIDIA DGX Spark nodes over a MikroTik CRS804 with 200G breakout: management and compute networks, RoCEv2/RDMA, sparkrun and NAS.
- 24 JulyNew white paperDGX Spark 4-Node AI Cluster Setup GuideSwitch-based AI cluster setup with 4 NVIDIA DGX Spark nodes over a MikroTik CRS812: management and compute networks, RoCEv2/RDMA, sparkrun and NAS.
- 6 JulyNew white paperQwen3.6-27B DGX Spark Cluster ScalingMulti-node scaling study of Qwen3.6-27B-NVFP4 on 1x, 2x, and 4x NVIDIA DGX Spark (GB10): tensor parallelism over 200GbE, SLO-driven capacity planning, and TP-vs-replication deployment guidance.
- 3 JulyNew white paperQwen3.6-27B DGX Spark BenchmarkPerformance evaluation of the Qwen3.6-27B model on the NVIDIA DGX Spark (GB10) platform with FP8, FP8-MTP, AWQ-MTP, NVFP4, and NVFP4-MTP quantization variants.
June 2026
- 30 JuneNew white paperLocal LLM Usage GuideEnd-to-end decision guide for local LLM usage: hardware (NVIDIA Jetson, RTX PRO, DGX Spark, DGX/HGX), model selection, software stack and scenario mapping.