LLM Inference Benchmark Explorer
Compare the inference performance of different LLM, hardware and serving configurations in one place. Use the filters to select the configurations you want to compare; the table updates automatically when you change the concurrency level or your performance targets. Expand any row to view the detailed results for that configuration.
How to use the benchmark explorer
What a row is
A row is one deployment configuration: a model, on specific hardware, in a specific number format, served by a specific engine with a specific parallelism and speculative-decoding setting, measured end to end. The same model therefore appears in several rows, and two rows are directly comparable only once you know which of those settings differ between them.
Filters and targets do different jobs
Filters decide which rows you see. The model, parameter-count, device, quantization and speculative-decoding filters, and the TPS, TTFT and capacity sliders, only show or hide rows.
Targets and assumptions decide what the numbers say. They sit under Performance Targets and Capacity Assumptions. Changing one leaves the rows in place but recalculates Max C, both capacity columns and the green and red colouring of every row.
1. Choose a concurrency level
Concurrency (C) is the number of requests the system is working on at the same moment. The selector picks which measurement the TPS and TTFT columns show: C=1 is the best case a single request sees, and higher values show the system under load. Rows not measured at the selected level are hidden, so the list shortens as C rises.
2. Narrow and sort
Filters combine, and every column sorts. The most instructive comparisons change a single setting: the same model at FP8 and at NVFP4, the same configuration with and without speculative decoding, or vLLM against SGLang on the same hardware.
3. Set your targets
Two targets define acceptable speed: a maximum TTFT (default 1000 ms) and a minimum TPS per request (default 20 tok/s). A measured level counts as supported only when it meets both; a value exactly on a target passes.
TTFT matters most where someone waits for the reply to start, as in chat. TPS matters most for long answers, where the wait is spread over the whole reply. To see what a TPS value feels like, the preview under the targets streams sample text at that rate.
4. Read Max C and the capacity columns
Max C is the highest measured concurrency that meets both targets. It counts simultaneous requests. A plus — 64+ — means the run met both targets even at the highest level it was tested at, so its real maximum was not reached.
Chat Capacity and Agentic Capacity turn that into a number of people and check it against memory. Each shows the smaller of two estimates, and the icon beside the figure says which one set it:
- a lightning bolt — the speed targets;
- a memory chip — the KV cache memory;
- a warning triangle — the chosen context length is longer than the model can hold.
A plus on a capacity figure — 256+ — follows from a Max C with a plus: speed set the figure, but the speed limit was never reached, so the figure is a minimum. Sorting and the minimum-capacity filters use the number itself.
Hover over a figure for a one-line reason. What the numbers mean and how they are calculated, below, explains both estimates in full.
5. Open a row
Click a row to see:
- the full concurrency sweep, marked PASS or FAIL against your targets at every level;
- a speed preview for any measured level, so C=1 and C=32 can be compared by eye;
- a chart of per-request TPS, the speed each user sees, against aggregate TPS, what the system produces in total — the trade-off capacity planning sits between;
- a short capacity summary: what each limit allows and how much memory is left for the KV cache;
- the run’s notes: cache precision, kernels, memory settings, speculation depth.
The chart downloads as a PNG for reports and presentations.
Where to start
“We want to run this model for an agentic workload used by 20 people.” Select the model and set the minimum Agentic Capacity to 20. Keep the default targets for a first estimate, or set them to your application. What remains are the candidate hardware and serving configurations.
“How does this model do on DGX B300 versus DGX Spark?” Select the model and both device families, then compare rows whose quantization, engine, speculative decoding and parallelism match. Changing the concurrency shows how the gap develops under load.
“What does quantization, speculative decoding or the engine actually change?” Hold the model and hardware fixed and compare the rows that differ in that one setting, so the effect is not mixed with a hardware change.
“We already own this hardware; what can we run on it?” Start with the device filter, sort the remaining models by capability or size, and narrow with the TPS, TTFT or capacity you need.
“Which model is capable enough without exceeding our limits?” Sort by Intelligence Index, Agentic Index or parameter count to shortlist models, then apply your device and performance requirements. This keeps choosing a model separate from sizing the infrastructure, instead of assuming the fastest model is the most suitable.
What the numbers mean and how they are calculated
How the numbers were measured
Every TPS and TTFT figure was measured with the open-source CordatusAI LLM Benchmark Tool on NVIDIA DGX B300, one to eight DGX Spark nodes, RTX PRO 6000 Blackwell and Jetson AGX Thor.
Each configuration is run with 128 input tokens and 128 output tokens, ten
rounds per concurrency level with prompts on different topics, at
C = 1, 2, 4, 8, 16, 32, 64. The table shows the mean of the ten rounds. A
token is the unit a model reads and writes, about three quarters of an
English word.
The measured columns
TPS (tokens per second) — how fast one request’s reply is produced at the selected concurrency. It is a per-request figure: at C=8, each of the eight requests receives this rate. The system’s total output is aggregate TPS = C × TPS, shown in the expanded row. Higher is better.
TTFT (time to first token) — how long a request waits before its first token arrives. It includes time spent queueing and the prefill, the pass in which the model reads the whole prompt. Lower is better.
The columns that describe the setup
Parameters — the total number of weights in the model. For a mixture-of-experts model this is the total, not the part active for each token, because all of it has to be held in memory.
Intelligence Index and Agentic Index — capability scores published by Artificial Analysis; higher is better. The first combines evaluations of reasoning, coding, science and long-context work; the second measures multi-step work with tool calls. Both describe the model, so every row of one model carries the same pair, whatever the hardware. Where several reasoning-effort settings are scored, the highest is shown, and a dash means no score has been published. The values are from Intelligence Index v4.3, retrieved on 28 September 2026. Scores from different index versions are not comparable — v4.2 and v4.3 added harder tasks, so every model scores lower than it did under v4.1 — and some older models have no v4.3 Agentic Index yet. The Agentic Index measures what the model can do; Agentic Capacity, further along the row, estimates how many people the hardware can serve.
Quantization — the number format the weights are stored in. Fewer bits per weight means less memory and usually more speed, at some risk to output quality. BF16 and FP16 are full precision; FP8 and MXFP8 use 8 bits per weight; NVFP4, MXFP4, FP4, INT4 and AWQ use 4.
Inference engine — the server software that loads the model and schedules requests, such as vLLM or SGLang. The same model on the same hardware can perform measurably differently under two engines.
Speculative decoding — the model drafts several tokens ahead and checks them in one pass; the accepted tokens are kept, so the same output arrives sooner. The column shows whether a run used it; the mechanism and how far ahead it drafted are in the row’s notes.
TP / DP / PP — how the model is split across GPUs or machines. Tensor parallelism (TP) divides the work inside each layer, pipeline parallelism (PP) places different layers on different devices, and data parallelism (DP) runs several complete copies, each serving its own requests.
Max C
Max C is the highest measured concurrency at which mean TTFT and mean TPS both meet your targets. Only measured levels count, nothing is interpolated, and if no level passes Max C is 0. It depends on your targets and moves with them: the same row may reach C=16 at 20 tok/s and only C=8 at 30 tok/s.
From requests to people: the capacity columns
Max C counts requests running at the same moment. The capacity columns estimate how many people a configuration can serve. A system can run out of speed or of memory first, so each column takes the smaller of two limits:
capacity = min(speed limit, KV cache memory limit)
The speed limit
speed limit = floor(Max C × usage multiplier)
A person does not have a request running all the time. A chat user sends a message, waits for the reply, then reads and types for a while, and uses no capacity in between. The usage multiplier is the number of people who, on average, share one request slot. The chat default of 4 assumes a chat user has a request running about a quarter of the time. The agentic default of 1.5 assumes an agent has one running about two thirds of the time, because it chains calls while it plans, runs tools and checks results.
With Max C = 8, that is 8 × 4 = 32 chat users or 8 × 1.5 = 12 agentic users.
The KV cache memory limit
While it works through a conversation, a model keeps a KV cache: for every token so far, the intermediate results each layer needs so that earlier tokens do not have to be processed again. The cache grows with the length of the conversation.
The speed figures come from 128-token prompts, and a real turn is that quick only if the user’s conversation is still in the cache, so that just the new message has to be processed. The memory limit therefore counts how many users’ whole sessions fit in the cache at once. A session is as long as the context length set under the assumptions — every token it holds, history and replies included: 16K by default for chat, 64K for agentic work, whose sessions also carry tool calls and their results. Beyond that number the system still runs, but a returning user waits while their conversation is processed again.
It is worked out per device — per GPU, or per node on DGX Spark:
- Memory for the engine = device memory × Engine Memory Allocation: 95% on a discrete GPU (DGX B300, RTX PRO 6000), 85% on unified memory (DGX Spark, Jetson Thor), where the operating system shares the same pool.
- Room for weights and cache = that × Weights and KV Cache Share (90%). The other 10% is working memory for activations and runtime buffers.
- Free for the KV cache = that − the weights held on this device: the size of the checkpoint the run served, divided over the TP × PP devices it is split across.
- One session = context length × the model’s cache per token, stored at FP8 (one byte per value), plus any fixed per-session part the model has (see below), divided over the devices that share it.
- Users = floor(free memory ÷ one session) × the number of DP copies.
The cache per token comes from the model’s published configuration. For a standard transformer it is 2 (a key and a value) × layers × KV heads × head size.
Worked example. A hypothetical 32-layer model with 8 KV heads of size 128, served as a 32 GB FP8 checkpoint on one 96 GB GPU:
- memory for the engine: 96 × 95% = 91.2 GB; room for weights and cache: 91.2 × 90% = 82.08 GB
- free for the KV cache: 82.08 − 32 = 50.08 GB
- cache per token: 2 × 32 × 8 × 128 = 65,536 bytes, so 50.08 GB holds 764,160 tokens
- chat: 764,160 ÷ 16,384 = 46 sessions; agentic: 764,160 ÷ 65,536 = 11
With Max C = 8 the speed limit is 32 chat and 12 agentic users, so the table shows 32, set by speed, and 11, set by memory. Tighten TTFT until Max C falls to 4 and the speed limit becomes 16 and 6: both are then set by speed.
How different model designs are handled
Models differ in what they keep for each token, so every row uses its own model’s figures. For example:
- Sliding-window layers keep only their most recent tokens — the last 128 or 1,024, say — so they cost a fixed amount per session instead of growing with it.
- Linear-attention and Mamba layers keep a fixed-size state instead of a per-token cache, counted once for each session.
- Compressed caches, such as the MLA of DeepSeek, Kimi and GLM-5, store far less per token, but every GPU holds a full copy, so adding GPUs does not divide them.
- A standard cache is split across GPUs by its KV heads; with more GPUs than heads, the heads are copied rather than split further.
Some rows show only the speed limit, and the expanded row says why: either the served checkpoint alone is larger than the memory assumed to be available — the run gave the engine more memory than the standard allocation, or kept part of the model or its cache in CPU memory or on disk — or the model stores its cache in a way this estimate does not model. Choose a context length longer than a model’s own context window, the most it can hold, and the model cannot serve sessions that long at all: its capacity shows a dash with a warning triangle.
Reading the results at longer prompts
The TTFT target applies directly to 128-token prompts. With longer prompts, TTFT rises roughly in proportion to their length, because the prefill work grows with the prompt. TPS falls more slowly: the main cost of producing each token, the weight-matrix multiplications, does not depend on the prompt’s length, although attention and cache reads do grow with the context.
What the numbers do not tell you
The table uses one fixed workload, mean values, simple usage multipliers and a standard memory calculation, so that configurations can be compared without first building a traffic model. It is not a substitute for a production load test: tail latency, varied prompt and output lengths, arrival patterns, agent call chains and prefix sharing all change real capacity.
The memory limit is calculated, not read from the engine. Engine-specific storage details — block rounding, scale factors, separate cache pools, experts spread across data-parallel devices — can move it either way. A session holding 16K or 64K tokens will also see a longer TTFT and a somewhat lower TPS than these 128-token measurements. Capacity is most meaningful where a full concurrency sweep stands behind it; a row measured only at C=1 rests on a single data point.
Planned refinements include workload filters, speed measured at longer contexts, and capacity estimates based on measured latency, user think time and Little’s Law.
Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution. All other columns are measured with the CordatusAI LLM Benchmark Tool.
In this measurement, GPT-OSS 120B (120B parameters), served in MXFP4 format with vLLM on 1× DGX Spark, reached a generation speed of 55.6 tok/s per request and a time to first token (TTFT) of 220 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 16 users for chat use and 6 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GPT-OSS 120B (120B parameters), served in MXFP4 format with vLLM on 2× DGX Spark (TP=2), reached a generation speed of 69.9 tok/s per request and a time to first token (TTFT) of 168 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 32 users for chat use and 12 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GPT-OSS 120B (120B parameters), served in MXFP4 format with vLLM on RTX PRO 6000, reached a generation speed of 169.6 tok/s per request and a time to first token (TTFT) of 52 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 55 users for chat use and 13 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3-4B-Instruct-2507 (4B parameters), served in NVFP4 format with vLLM on 1× DGX Spark, reached a generation speed of 49.6 tok/s per request and a time to first token (TTFT) of 42 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 78 users for chat use and 19 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3-Coder-30B-A3B-Instruct (30B parameters), served in FP16 format with vLLM on 1× DGX Spark, reached a generation speed of 29.9 tok/s per request and a time to first token (TTFT) of 171 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3-Coder-30B-A3B-Instruct (30B parameters), served in NVFP4 format with vLLM on 1× DGX Spark, reached a generation speed of 65 tok/s per request and a time to first token (TTFT) of 66 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 99 users for chat use and 24 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.6-27B (27B parameters), served in FP8 format with vLLM on 1× DGX Spark, reached a generation speed of 8.1 tok/s per request and a time to first token (TTFT) of 271 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Qwen3.6-27B (27B parameters), served in FP8 format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 19.5 tok/s per request and a time to first token (TTFT) of 323 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Qwen3.6-27B (27B parameters), served in AWQ format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 25.4 tok/s per request and a time to first token (TTFT) of 267 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 32 users for chat use and 12 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.6-27B (27B parameters), served in NVFP4 format with vLLM on 1× DGX Spark, reached a generation speed of 9.9 tok/s per request and a time to first token (TTFT) of 228 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Qwen3.6-27B (27B parameters), served in NVFP4 format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 21 tok/s per request and a time to first token (TTFT) of 519 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.6-27B (27B parameters), served in NVFP4 format with vLLM on 2× DGX Spark (TP=2), reached a generation speed of 22.6 tok/s per request and a time to first token (TTFT) of 149 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.6-27B (27B parameters), served in NVFP4 format with vLLM on 4× DGX Spark (TP=4), reached a generation speed of 33.1 tok/s per request and a time to first token (TTFT) of 164 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 32 users for chat use and 12 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.5-27B (27B parameters), served in FP8 format with vLLM on RTX PRO 6000, reached a generation speed of 28.6 tok/s per request and a time to first token (TTFT) of 66 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 60 users for chat use and 20 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.6-27B (27B parameters), served in BF16 format with vLLM on DGX B300, reached a generation speed of 82.6 tok/s per request and a time to first token (TTFT) of 69 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 225 users for chat use and 77 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.6-35B-A3B (35B parameters), served in FP8 format with vLLM on 1× DGX Spark, reached a generation speed of 52.2 tok/s per request and a time to first token (TTFT) of 115 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 32 users for chat use and 12 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.6-35B-A3B (35B parameters), served in FP8 format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 64.8 tok/s per request and a time to first token (TTFT) of 140 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 64 users for chat use and 24 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.6-35B-A3B (35B parameters), served in NVFP4 format with vLLM on 4× DGX Spark (TP=4), reached a generation speed of 91.2 tok/s per request and a time to first token (TTFT) of 231 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 128 users for chat use and 48 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.6-35B-A3B (35B parameters), served in BF16 format with vLLM on DGX B300, reached a generation speed of 241.4 tok/s per request and a time to first token (TTFT) of 54 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 256 users for chat use and at least 96 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, Qwen3.5-35B-A3B (35B parameters), served in FP8 format with vLLM on RTX PRO 6000, reached a generation speed of 113.8 tok/s per request and a time to first token (TTFT) of 47 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 150 users for chat use and 55 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.8-27B (27B parameters), served in NVFP4 format with SGLang and speculative decoding on 1× DGX Spark, reached a generation speed of 47.9 tok/s per request and a time to first token (TTFT) of 232 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 32 users for chat use and 12 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.8-27B (27B parameters), served in NVFP4 format with SGLang and speculative decoding on 4× DGX Spark (TP=4), reached a generation speed of 77.3 tok/s per request and a time to first token (TTFT) of 227 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 32 users for chat use and 12 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.8-27B (27B parameters), served in BF16 format with vLLM on DGX B300, reached a generation speed of 85.3 tok/s per request and a time to first token (TTFT) of 63 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 225 users for chat use and 77 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.8-27B (27B parameters), served in BF16 format with vLLM and speculative decoding on DGX B300, reached a generation speed of 172.9 tok/s per request and a time to first token (TTFT) of 77 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 225 users for chat use and 77 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.8-Flash-Next (176B parameters), served in NVFP4 format with SGLang on 2× DGX Spark (TP=2), reached a generation speed of 35.8 tok/s per request and a time to first token (TTFT) of 200 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 16 users for chat use and 6 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.8-Flash-Next (176B parameters), served in BF16 format with SGLang on DGX B300 (TP=2), reached a generation speed of 250.1 tok/s per request and a time to first token (TTFT) of 108 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 256 users for chat use and at least 96 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, Qwen3.8-Flash-Next (176B parameters), served in NVFP4 format with SGLang and speculative decoding on 1× DGX Spark, reached a generation speed of 28.5 tok/s per request and a time to first token (TTFT) of 302 ms with a single request.
Considering the speed targets alone (the KV cache limit was not calculated for this configuration), the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.8-Flash-Next (176B parameters), served in NVFP4 format with SGLang and speculative decoding on RTX PRO 6000, reached a generation speed of 155.8 tok/s per request and a time to first token (TTFT) of 139 ms with a single request.
Considering the speed targets alone (the KV cache limit was not calculated for this configuration), the estimated capacity is 64 users for chat use and 24 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.5-397B-A17B (397B parameters), served in INT4 format with vLLM on 3× DGX Spark (PP=3), reached a generation speed of 17 tok/s per request and a time to first token (TTFT) of 500 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Qwen3.5-397B-A17B (397B parameters), served in INT4 format with vLLM on 4× DGX Spark (TP=4), reached a generation speed of 38.1 tok/s per request and a time to first token (TTFT) of 325 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 16 users for chat use and 6 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.2 (753B parameters), served in INT4 format with vLLM and speculative decoding on 4× DGX Spark (TP=4), reached a generation speed of 27.4 tok/s per request and a time to first token (TTFT) of 490 ms with a single request.
Considering the speed targets alone (the KV cache limit was not calculated for this configuration), the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.2 (753B parameters), served in NVFP4 format with vLLM and speculative decoding on 8× DGX Spark (TP=8), reached a generation speed of 20.7 tok/s per request and a time to first token (TTFT) of 466 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 4 users for chat use and 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.2 (753B parameters), served in NVFP4 format with vLLM on 8× DGX Spark (TP=8), reached a generation speed of 24.9 tok/s per request and a time to first token (TTFT) of 451 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 4 users for chat use and 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.2 (753B parameters), served in FP8 format with vLLM on DGX B300 (TP=4), reached a generation speed of 160.1 tok/s per request and a time to first token (TTFT) of 307 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 73 users for chat use and 18 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.2 (753B parameters), served in NVFP4 format with vLLM on DGX B300 (TP=4), reached a generation speed of 213.2 tok/s per request and a time to first token (TTFT) of 89 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 166 users for chat use and 41 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.3-Flash (321B parameters), served in NVFP4 format with vLLM on 2× DGX Spark (TP=2), reached a generation speed of 24.3 tok/s per request and a time to first token (TTFT) of 395 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 2 users for chat use and 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.3-Flash (321B parameters), served in FP8 format with vLLM on 4× DGX Spark (TP=4), reached a generation speed of 29.6 tok/s per request and a time to first token (TTFT) of 383 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.3-Flash (321B parameters), served in FP8 format with SGLang on DGX B300 (TP=2), reached a generation speed of 148.7 tok/s per request and a time to first token (TTFT) of 262 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 32 users for chat use and 12 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.3-Flash (321B parameters), served in FP8 format with SGLang on DGX B300 (TP=4), reached a generation speed of 161 tok/s per request and a time to first token (TTFT) of 292 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 128 users for chat use and 48 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.3 (753B parameters), served in FP8 format with SGLang on DGX B300 (TP=4), reached a generation speed of 155.9 tok/s per request and a time to first token (TTFT) of 286 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 73 users for chat use and 18 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.3 (753B parameters), served in NVFP4 format with vLLM and speculative decoding on 8× DGX Spark (TP=8), reached a generation speed of 22.6 tok/s per request and a time to first token (TTFT) of 630 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 4 users for chat use and 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Inkling (952B parameters), served in NVFP4 format with vLLM and speculative decoding on 8× DGX Spark (TP=8), reached a generation speed of 27.3 tok/s per request and a time to first token (TTFT) of 412 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Inkling (952B parameters), served in NVFP4 format with vLLM on DGX B300 (TP=4), reached a generation speed of 134.2 tok/s per request and a time to first token (TTFT) of 51 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 256 users for chat use and at least 96 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, DeepSeek V4 Flash 0731 (304B parameters), served in NVFP4 format with vLLM and speculative decoding on 2× DGX Spark (TP=2), reached a generation speed of 47.3 tok/s per request and a time to first token (TTFT) of 308 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 16 users for chat use and 6 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, DeepSeek V4 Flash 0731 (304B parameters), served in NVFP4 format with vLLM on 8× DGX Spark (TP=2, DP=4), reached a generation speed of 48.9 tok/s per request and a time to first token (TTFT) of 289 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 64 users for chat use and 24 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, MiniMax-M2.7 (229B parameters), served in NVFP4 format with vLLM on 2× DGX Spark (TP=2), reached a generation speed of 26.6 tok/s per request and a time to first token (TTFT) of 289 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, MiniMax-M2.7 (229B parameters), served in NVFP4 format with vLLM on 2× DGX Spark (PP=2), reached a generation speed of 17.2 tok/s per request and a time to first token (TTFT) of 396 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, MiniMax-M2.7 (229B parameters), served in NVFP4 format with vLLM on 3× DGX Spark (PP=3), reached a generation speed of 18 tok/s per request and a time to first token (TTFT) of 211 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, MiniMax-M2.7 (229B parameters), served in NVFP4 format with vLLM on 4× DGX Spark (TP=4), reached a generation speed of 30.1 tok/s per request and a time to first token (TTFT) of 235 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 16 users for chat use and 6 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, MiniMax-M3 (427B parameters), served in NVFP4 format with vLLM and speculative decoding on 4× DGX Spark (TP=4), reached a generation speed of 34.5 tok/s per request and a time to first token (TTFT) of 453 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, MiniMax-M3 (427B parameters), served in MXFP8 format with vLLM on DGX B300 (TP=4), reached a generation speed of 114.4 tok/s per request and a time to first token (TTFT) of 57 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 256 users for chat use and 91 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, Kimi K3 (2.78T parameters), served in MXFP4 format with vLLM on DGX B300 (TP=8), reached a generation speed of 100.8 tok/s per request and a time to first token (TTFT) of 67 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 150 users for chat use and 50 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Kimi K3 (2.78T parameters), served in MXFP4 format with vLLM and speculative decoding on DGX B300 (TP=8), reached a generation speed of 187.7 tok/s per request and a time to first token (TTFT) of 71 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 128 users for chat use and 48 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Kimi K3 (2.78T parameters), served in MXFP4 format with SGLang on DGX B300 (TP=8), reached a generation speed of 83.4 tok/s per request and a time to first token (TTFT) of 400 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 150 users for chat use and 50 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Gemma-4-26B-A4B-it (26B parameters), served in FP16 format with vLLM on 1× DGX Spark, reached a generation speed of 21.6 tok/s per request and a time to first token (TTFT) of 247 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 4 users for chat use and 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Gemma-4-26B-A4B-it (26B parameters), served in FP16 format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 30.6 tok/s per request and a time to first token (TTFT) of 256 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Gemma-4-31B-it (31B parameters), served in NVFP4 format with vLLM on 1× DGX Spark, reached a generation speed of 10.8 tok/s per request and a time to first token (TTFT) of 202 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Gemma-4-31B-it (31B parameters), served in NVFP4 format with vLLM on 2× DGX Spark (TP=2), reached a generation speed of 19.4 tok/s per request and a time to first token (TTFT) of 136 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Gemma-4-31B-it (31B parameters), served in NVFP4 format with vLLM on 4× DGX Spark (TP=4), reached a generation speed of 30.5 tok/s per request and a time to first token (TTFT) of 280 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 64 users for chat use and 24 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Gemma-4-31B-it (31B parameters), served in BF16 format with vLLM on DGX B300, reached a generation speed of 74.8 tok/s per request and a time to first token (TTFT) of 35 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 168 users for chat use and 59 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Gemma-4-31B-it (31B parameters), served in NVFP4 format with vLLM on RTX PRO 6000, reached a generation speed of 40 tok/s per request and a time to first token (TTFT) of 38 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 53 users for chat use and 18 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Laguna-S-2.1 (118B parameters), served in NVFP4 format with vLLM on Thor, reached a generation speed of 12.9 tok/s per request and a time to first token (TTFT) of 578 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Muse-Glimmer-30B (30B parameters), served in BF16 format with SGLang on 1× DGX Spark, reached a generation speed of 25 tok/s per request and a time to first token (TTFT) of 765 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 4 users for chat use and 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Muse-Glimmer-30B (30B parameters), served in BF16 format with vLLM and speculative decoding on RTX PRO 6000, reached a generation speed of 165.4 tok/s per request and a time to first token (TTFT) of 118 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 128 users for chat use and 47 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, DeepSeek-V3.1 (685B parameters), served in NVFP4 format with vLLM on DGX B300 (TP=2), reached a generation speed of 108.6 tok/s per request and a time to first token (TTFT) of 140 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 68 users for chat use and 17 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, DeepSeek-V4-Pro (1.60T parameters), served in FP4 format with vLLM on DGX B300 (TP=4), reached a generation speed of 86.8 tok/s per request and a time to first token (TTFT) of 303 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 128 users for chat use and 48 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, DeepSeek-V4-Pro-0813 (1.65T parameters), served in FP8 format with vLLM on DGX B300 (TP=4), reached a generation speed of 117.8 tok/s per request and a time to first token (TTFT) of 279 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 128 users for chat use and 48 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Gemma-4-26B-A4B-it (26B parameters), served in BF16 format with vLLM on DGX B300, reached a generation speed of 168.2 tok/s per request and a time to first token (TTFT) of 60 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 256 users for chat use and at least 96 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, GLM-4.5-Air (110B parameters), served in BF16 format with vLLM on DGX B300, reached a generation speed of 107.9 tok/s per request and a time to first token (TTFT) of 62 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 16 users for chat use and 4 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-4.6 (357B parameters), served in BF16 format with vLLM on DGX B300 (TP=4), reached a generation speed of 62.1 tok/s per request and a time to first token (TTFT) of 275 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 87 users for chat use and 21 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-4.7 (358B parameters), served in BF16 format with vLLM on DGX B300 (TP=4), reached a generation speed of 85.9 tok/s per request and a time to first token (TTFT) of 53 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 86 users for chat use and 21 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-4.7 (358B parameters), served in FP8 format with vLLM on DGX B300 (TP=2), reached a generation speed of 80.4 tok/s per request and a time to first token (TTFT) of 53 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 42 users for chat use and 10 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-4.7 (358B parameters), served in NVFP4 format with vLLM on DGX B300, reached a generation speed of 63.9 tok/s per request and a time to first token (TTFT) of 157 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 21 users for chat use and 5 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5 (754B parameters), served in FP8 format with vLLM on DGX B300 (TP=4), reached a generation speed of 94 tok/s per request and a time to first token (TTFT) of 66 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 63 users for chat use and 15 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, GLM-5.1 (754B parameters), served in FP8 format with vLLM on DGX B300 (TP=4), reached a generation speed of 91.5 tok/s per request and a time to first token (TTFT) of 453 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 63 users for chat use and 15 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Kimi-K2.5 (1.03T parameters), served in NVFP4 format with vLLM on DGX B300 (TP=4), reached a generation speed of 128.8 tok/s per request and a time to first token (TTFT) of 64 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 171 users for chat use and 42 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Kimi-K2.6 (1.03T parameters), served in INT4 format with vLLM on DGX B300 (TP=4), reached a generation speed of 203.9 tok/s per request and a time to first token (TTFT) of 56 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 128 users for chat use and 42 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, Kimi-K2-Thinking (1.03T parameters), served in NVFP4 format with vLLM on DGX B300 (TP=4), reached a generation speed of 53.6 tok/s per request and a time to first token (TTFT) of 48 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 128 users for chat use and 42 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, MiMo-V2.5-Pro (1.02T parameters), served in FP8 format with vLLM on DGX B300 (TP=8), reached a generation speed of 137.1 tok/s per request and a time to first token (TTFT) of 64 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 128 users for chat use and at least 48 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, MiniMax-M2 (229B parameters), served in FP8 format with vLLM on DGX B300 (TP=2), reached a generation speed of 60.1 tok/s per request and a time to first token (TTFT) of 452 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 126 users for chat use and 31 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, MiniMax-M2.5 (229B parameters), served in FP8 format with vLLM on DGX B300 (TP=2), reached a generation speed of 66.8 tok/s per request and a time to first token (TTFT) of 169 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 126 users for chat use and 31 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, MiniMax-M2.7 (229B parameters), served in FP8 format with vLLM on DGX B300 (TP=2), reached a generation speed of 129.2 tok/s per request and a time to first token (TTFT) of 46 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 126 users for chat use and 31 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Nemotron-3-Ultra (550B parameters), served in BF16 format with vLLM on DGX B300 (TP=8), reached a generation speed of 177.3 tok/s per request and a time to first token (TTFT) of 132 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 64 users for chat use and 24 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3-235B-A22B (235B parameters), served in BF16 format with vLLM on DGX B300 (TP=4), reached a generation speed of 83.9 tok/s per request and a time to first token (TTFT) of 132 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 128 users for chat use and at least 48 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, Qwen3.5-397B-A17B (397B parameters), served in FP8 format with vLLM on DGX B300 (TP=2), reached a generation speed of 130.1 tok/s per request and a time to first token (TTFT) of 74 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 128 users for chat use and at least 48 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, Qwen3.8-2.4T-A95B (2.4T parameters), served in NVFP4 format with vLLM on DGX B300 (TP=8), reached a generation speed of 115.9 tok/s per request and a time to first token (TTFT) of 113 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 192 users for chat use and 71 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Tencent-Hy3 (299B parameters), served in BF16 format with vLLM on DGX B300 (TP=4), reached a generation speed of 149.6 tok/s per request and a time to first token (TTFT) of 37 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 144 users for chat use and 36 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Qwen3.8-27B (27B parameters), served in BF16 format with vLLM on 1× DGX Spark, reached a generation speed of 4.5 tok/s per request and a time to first token (TTFT) of 335 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Qwen3.8-27B (27B parameters), served in BF16 format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 9.9 tok/s per request and a time to first token (TTFT) of 611 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Qwen3.8-27B (27B parameters), served in FP8 format with vLLM on 1× DGX Spark, reached a generation speed of 7.9 tok/s per request and a time to first token (TTFT) of 172 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Qwen3.8-27B (27B parameters), served in FP8 format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 18.4 tok/s per request and a time to first token (TTFT) of 330 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Qwen3.8-27B (27B parameters), served in NVFP4 format with vLLM on 1× DGX Spark, reached a generation speed of 9.7 tok/s per request and a time to first token (TTFT) of 155 ms with a single request.
At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.
In this measurement, Qwen3.8-27B (27B parameters), served in NVFP4 format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 30.6 tok/s per request and a time to first token (TTFT) of 281 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 4 users for chat use and at least 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
This configuration was measured with a single request only, so the estimate rests on one measurement.
In this measurement, DeepSeek V4 Flash 0731 (304B parameters), served in NVFP4 format with vLLM on 4× DGX Spark (TP=2, DP=2), reached a generation speed of 48.5 tok/s per request and a time to first token (TTFT) of 277 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 4 users for chat use and at least 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
This configuration was measured with a single request only, so the estimate rests on one measurement.
In this measurement, DeepSeek V4 Flash 0731 (304B parameters), served in NVFP4 format with vLLM on 4× DGX Spark (TP=4), reached a generation speed of 69.5 tok/s per request and a time to first token (TTFT) of 256 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 4 users for chat use and at least 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
This configuration was measured with a single request only, so the estimate rests on one measurement.
In this measurement, DeepSeek V4 Flash 0731 (304B parameters), served in NVFP4 format with SGLang on 8× DGX Spark (TP=4, DP=2), reached a generation speed of 63.6 tok/s per request and a time to first token (TTFT) of 238 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 4 users for chat use and at least 1 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
This configuration was measured with a single request only, so the estimate rests on one measurement.
In this measurement, DeepSeek-V4.1-Flash (763B parameters), served in FP8 format with vLLM and speculative decoding on DGX B300 (TP=4), reached a generation speed of 284.5 tok/s per request and a time to first token (TTFT) of 64 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 256 users for chat use and at least 96 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.
In this measurement, DeepSeek-V4.1-Flash (763B parameters), served in FP8 format with vLLM and speculative decoding on 4× DGX Spark (TP=4), reached a generation speed of 29.5 tok/s per request and a time to first token (TTFT) of 272 ms with a single request.
Considering the speed targets alone (the KV cache limit was not calculated for this configuration), the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, DeepSeek-V4.1-Flash (763B parameters), served in FP8 format with vLLM and speculative decoding on 8× DGX Spark (TP=8), reached a generation speed of 36 tok/s per request and a time to first token (TTFT) of 213 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 16 users for chat use and 6 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, DeepSeek-V4.1-Flash (763B parameters), served in FP8 format with vLLM and speculative decoding on 8× DGX Spark (TP=8), reached a generation speed of 33.3 tok/s per request and a time to first token (TTFT) of 199 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, Tencent-Hy4-preview (780B parameters), served in BF16 format with vLLM and speculative decoding on DGX B300 (TP=8), reached a generation speed of 65.7 tok/s per request and a time to first token (TTFT) of 110 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 65 users for chat use and 16 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, DeepSeek-V4-Flash-Vision-Exp (305B parameters), served in FP8 format with vLLM and speculative decoding on 2× DGX Spark (TP=2), reached a generation speed of 36 tok/s per request and a time to first token (TTFT) of 300 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
In this measurement, MiMo-V2.6-Pro-RL (1.02T parameters), served in FP4 format with vLLM and speculative decoding on DGX B300 (TP=4), reached a generation speed of 333.5 tok/s per request and a time to first token (TTFT) of 46 ms with a single request.
Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 256 users for chat use and at least 96 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.
“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.