Aetherix Technical White Papers
In-depth technical analysis of AI inference on NVIDIA hardware. From Jetson-based edge AI to DGX Spark clusters and RTX PRO workstations — deployment guides, cluster setups and measured LLM and computer-vision benchmarks.
Latest White Papers
How model architecture decides what must remain in GPU memory and what can move to host RAM or NVMe: conditional memory, memory hierarchies of DGX Spark, RTX PRO 6000 and DGX B300, and measured results with DeepSeek-V4.1-Flash and Qwen3.8-Flash-Next.
September 2026 LLM DeploymentDeepSeek-V4.1-Flash 8× DGX Spark TP8 DeploymentDeployment of DeepSeek-V4.1-Flash (763B MoE, FP8, DSpark k=5) on 8× NVIDIA DGX Spark (GB10) with TP8: two configurations (300K Engram-in-memory, 1M Engram-on-disk), NCCL optimization, benchmark results and TP4 comparison.
September 2026 CV BenchmarkCV Inference Benchmark ExplorerExplore computer-vision inference benchmarks on NVIDIA GPUs and Jetson devices. Pick a GPU and a detection model to see the frame rate it sustains, how it falls as cameras are added, and how many cameras it carries at your target FPS.
September 2026 LLM DeploymentDeepSeek-V4.1-Flash 4× DGX Spark DeploymentDeployment of DeepSeek-V4.1-Flash (763B MoE, FP8, DSpark k=5) on 4× NVIDIA DGX Spark (GB10) with tensor parallelism: vLLM build chain, 7 SM 12.1a patches, Engram-on-disk, benchmark results and B300 comparison.
September 2026 LLM BenchmarkLLM Inference Benchmark ExplorerExplore LLM inference benchmarks on NVIDIA DGX Spark, DGX B300, RTX PRO 6000 Blackwell and Jetson Thor. Filter by model, parameter count, device, quantization and concurrency, and set your own performance targets.
August 2026 Architecture ComparisonNVIDIA DGX B300 vs GB300 NVL72 Cluster Architecture ComparisonTechnical comparison of two NVIDIA Blackwell Ultra architectures, DGX B300 and GB300 NVL72: system design, scaling approach, network fabric, power and cooling, and which workloads suit which platform.
August 2026 LLM BenchmarkKimi K3 Inference Benchmark on DGX-B300Performance evaluation of Moonshot AI Kimi K3 (2.8T MoE, MXFP4) on NVIDIA DGX-B300 (8x Blackwell Ultra, TP=8): vLLM vs SGLang, direct vs DSpark speculative decoding, with SLO-driven capacity planning.
July 2026 Cluster SetupDGX Spark 3-Node AI Cluster Setup GuideRing (mesh) topology AI cluster setup with 3 NVIDIA DGX Spark nodes: management and compute networks, RoCEv2/RDMA, sparkrun configuration.
July 2026 Cluster SetupDGX Spark 2-Node AI Cluster Setup GuidePoint-to-point topology AI cluster setup with 2 NVIDIA DGX Spark nodes: management and compute networks, RoCEv2/RDMA, sparkrun configuration.
July 2026 Cluster SetupDGX Spark 8-Node AI Cluster Setup GuideSwitch-based AI cluster setup with 8 NVIDIA DGX Spark nodes over a MikroTik CRS804 with 200G breakout: management and compute networks, RoCEv2/RDMA, sparkrun and NAS.
July 2026 Cluster SetupDGX Spark 4-Node AI Cluster Setup GuideSwitch-based AI cluster setup with 4 NVIDIA DGX Spark nodes over a MikroTik CRS812: management and compute networks, RoCEv2/RDMA, sparkrun and NAS.
July 2026 LLM ScalingQwen3.6-27B DGX Spark Cluster ScalingMulti-node scaling study of Qwen3.6-27B-NVFP4 on 1x, 2x, and 4x NVIDIA DGX Spark (GB10): tensor parallelism over 200GbE, SLO-driven capacity planning, and TP-vs-replication deployment guidance.
July 2026 LLM BenchmarkQwen3.6-27B DGX Spark BenchmarkPerformance evaluation of the Qwen3.6-27B model on the NVIDIA DGX Spark (GB10) platform with FP8, FP8-MTP, AWQ-MTP, NVFP4, and NVFP4-MTP quantization variants.
July 2026 Decision GuideLocal LLM Usage GuideEnd-to-end decision guide for local LLM usage: hardware (NVIDIA Jetson, RTX PRO, DGX Spark, DGX/HGX), model selection, software stack and scenario mapping.
June 2026