← Roofline Labs

Benchmark table

Measured LLM decode speed, model × hardware × engine × quantization, with every row scored against the hardware's memory-bandwidth roofline. Serving runs under stated configs, run by us. Nothing third-party, nothing unreproduced.

45measured cells
9model architectures
3serving engines
125best single-stream tok/s · Qwen3.6-35B-A3B
2026-05-22to 2026-09-10

NVIDIA DGX Spark (GB10, Grace Blackwell)

memory
122 GiB unified LPDDR5X, no separate VRAM, swap off
spec bandwidth
273 GB/s
STREAM copy
245.6 GB/s
decode roofline divisor
187 GB/s

% roofline = measured ÷ (nodes × 187 ÷ active GB per token). 100% is the anchor: a clean single-stream, no-speculation run. Rows above 100% have speculation on or a better kernel path. Rows well below it are dequant-compute-bound or misconfigured. — means active bytes per token is not derived yet.

under 70%70 to 105%over 105%, speculation paying · click a header to sort · click a row for its notes

Single-stream decode concurrency 1

modelenginequantspecctxdecode tok/sroofline% rooflineprefill tok/sdate
OpenAI gpt-oss-120bgpt-oss-120b MXFP4 GGUFllama.cpp2026-05 buildMXFP4off40.62026-05-22
Poolside Laguna-S 2.1 118B-A8BLaguna-S-2.1 Q4_K_M GGUF (90 GB — model card said 68)llama.cppPoolside fork, branch laguna (2026-08)Q4_K_Moff22.530.7
73%at ceiling
2026-08-19
Poolside Laguna-S 2.1 118B-A8BLaguna-S-2.1 Q4_K_M GGUF + DFlash draftllama.cppPoolside fork, branch laguna (2026-08)Q4_K_Mdflash27.330.7
89%at ceiling
2026-08-19
Poolside Laguna-S 2.1 118B-A8Bpoolside/Laguna-S-2.1-NVFP4 @ 07614121b318 (15 shards, 71,916,714,120 B) + DFlash draftvllm0.25.1NVFP4dflash50.641.2
123%spec on
2026-08-23
Poolside Laguna-S 2.1 118B-A8Bpoolside/Laguna-S-2.1-NVFP4 @ 07614121b318 (71,916,714,120 B), no DFlashsglangv0.5.18-cu130NVFP4 + fp8 KVoff18.341.2
44%under
2447@ 290102026-09-06
Poolside Laguna-S 2.1 118B-A8Bpoolside/Laguna-S-2.1-NVFP4 @ 07614121b318, no DFlashvllmv0.28.0 (upstream image)NVFP4off19.641.2
48%under
2859@ 290102026-09-06
MiniMax M2.7 230B-A10Bcyankiwi/MiniMax-M2.7-AWQ-4bit @ 70ed6a57 (130,495,823,656 B)vllmv0.28.0 (upstream image) · 2× TP=2 over 200G RoCEAWQ-int4-g32 + fp8 KVoff41.963.2
66%under
2486@ 286002026-09-07
Mistral Small 4 119B (2603)mistralai/Mistral-Small-4-119B-2603-NVFP4 @ b1a90485 (70,808,028,368 B), MLA + expert Linears NVFP4, rest BF16vllmv0.28.0 (upstream image)NVFP4off31.62494@ 290102026-09-06
NVIDIA Nemotron 3 Nano 30B-A3BNemotron-3-Nano-30B-A3B Q8_0 GGUFllama.cpp2026-05 buildQ8_0off38.82026-05-22
NVIDIA Nemotron 3.5 Lightning 30B-A3Bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 @ cc84af2f (21,561,882,284 B), ModelOpt NVFP4sglangv0.5.18-cu130NVFP4 + fp8 KVnextn106.793.0
115%spec on
4713@ 290102026-09-06
NVIDIA Nemotron 3.5 Lightning 30B-A3Bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 @ cc84af2fsglangv0.5.18-cu130NVFP4 + fp8 KVnextn152505105.393.0
113%spec on
2026-09-10
NVIDIA Nemotron 3.5 Lightning 30B-A3Bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 @ cc84af2fsglangv0.5.18-cu130NVFP4 + fp8 KVnextn1013116.693.0
125%spec on
2026-09-10
NVIDIA Nemotron 3.5 Lightning 30B-A3Bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 @ cc84af2fsglangv0.5.18-cu130NVFP4 + fp8 KVnextn25414599.793.0
107%spec on
2026-09-10
NVIDIA Nemotron 3.5 Lightning 30B-A3Bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 @ cc84af2fsglangv0.5.18-cu130NVFP4 + fp8 KVnextn39934582.093.0
88%at ceiling
2026-09-10
Ornith-1.0-35BOrnith-1.0-35B, experts NVFP4 via llama-quantize RTN (data-free), attention Q8_0, no MTP headllama.cpp2026-08 buildNVFP4-RTN-expertsoff83.068.0
122%spec on
2026-08-08
Ornith-1.0-35BOrnith NVFP4 (AEON build, 23 GB) + DFlash drafter (905 MB)vllm0.27.1+aeon.sm121a (community container)NVFP4dflash95.368.0
140%spec on
5197@ 40962026-08-19
Ornith-1.0-35BOrnith-35B heretic APEX mixed-IQ quant (SC117), has MTP headllama.cpp2026-08 buildAPEX-mixeddraft-mtp70.878.1
91%at ceiling
2026-08-19
Ornith-1.0-35BOrnith-1.0-35B Q8_0 GGUF (no MTP head)llama.cpp2026-08 buildQ8_0off58.558.5
100%at ceiling
1608@ 41382026-08-19
Qwen3-Coder-Next 80B-A3BRedHatAI/Qwen3-Coder-Next-NVFP4 @ 27a8f16f (47,566,688,792 B), llm-compressor requant, no MTP headsglangv0.5.18-cu130NVFP4 + fp8 KVoff35.8112.7
32%under
4750@ 290102026-09-06
Qwen3.6-35B-A3BQwen3.6-35B-A3B heretic, experts NVFP4 via llama-quantize RTN (data-free), attention Q8_0llama.cpp2026-08 buildNVFP4-RTN-expertsdraft-mtp80.568.0
118%spec on
1823@ 41382026-08-19
Qwen3.6-35B-A3BQwen3.6-35B-A3B heretic (decensored) Q8_0 GGUF, native MTP head preservedllama.cpp2026-08 buildQ8_0draft-mtp69.158.5
118%spec on
2026-08-19
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1e, repacked to GGUF (MTP head left BF16)llama.cppv0.3.0 nativeNVFP4off55.168.0
81%at ceiling
2263@ 290102026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1e, repacked to GGUF (MTP head left BF16)llama.cppv0.3.0 nativeNVFP4draft-mtp89.068.0
131%spec on
2026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1esglangv0.5.18-cu130NVFP4 + fp8 KVoff75.868.0
112%spec on
5792@ 290102026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1esglangv0.5.18-cu130NVFP4 + fp8 KVnextn123.268.0
181%spec on
5571@ 290102026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1evllmv0.28.0 (upstream image)NVFP4off75.468.0
111%spec on
4967@ 290102026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1evllmv0.28.0 (upstream image)NVFP4 + fp8 KVoff76.668.0
113%spec on
4740@ 290102026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1evllmv0.28.0 (upstream image)NVFP4mtp124.968.0
184%spec on
4623@ 290102026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1esglangv0.5.18-cu130NVFP4 + fp8 KVnextn122.268.0
180%spec on
5584@ 290102026-09-06
Qwen3.6-35B-A3BQwen3.6-35B-A3B heretic Q8_0 GGUF, native MTP head preservedllama.cppv0.3.0 native (llama-swap)Q8_0draft-mtp68.158.5
116%spec on
2026-09-08

Fan-out decode concurrency > 1, aggregate tok/s

modelenginequantspecconcaggregate tok/sper stream× single-stream rooflinedate
OpenAI gpt-oss-120bgpt-oss-120b MXFP4 GGUFllama.cpp2026-05 buildMXFP4off1684.02026-05-22
MiniMax M2.7 230B-A10Bcyankiwi/MiniMax-M2.7-AWQ-4bit @ 70ed6a57vllmv0.28.0 (upstream image) · 2× TP=2 over 200G RoCEAWQ-int4-g32 + fp8 KVoff16184.011.6
2.91×
2026-09-07
NVIDIA Nemotron 3 Nano 30B-A3BNemotron-3-Nano-30B-A3B Q8_0 GGUFllama.cpp2026-05 buildQ8_0off1671.32026-05-22
NVIDIA Nemotron 3.5 Lightning 30B-A3Bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 @ cc84af2fsglangv0.5.18-cu130NVFP4 + fp8 KVnextn4210.356.1
2.26×
2026-09-06
NVIDIA Nemotron 3.5 Lightning 30B-A3Bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 @ cc84af2fsglangv0.5.18-cu130NVFP4 + fp8 KVnextn16405.526.4
4.36×
2026-09-07
NVIDIA Nemotron 3.5 Lightning 30B-A3Bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 @ cc84af2fsglangv0.5.18-cu130NVFP4 + fp8 KVnextn8290.0
3.12×
2026-09-07
NVIDIA Nemotron 3.5 Lightning 30B-A3Bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 @ cc84af2fsglangv0.5.18-cu130 · 2× DP=2 (independent replicas, client-side split)NVFP4 + fp8 KVnextn16419.0
2.25×
2026-09-07
Ornith-1.0-35BOrnith-1.0-35B, experts NVFP4 via llama-quantize RTN (data-free), attention Q8_0, no MTP headllama.cpp2026-08 buildNVFP4-RTN-expertsoff8250.9
3.69×
2026-08-08
Ornith-1.0-35BOrnith-1.0-35B Q8_0 GGUF (no MTP head)llama.cpp2026-08 buildQ8_0off8185.6
3.17×
2026-08-08
Qwen3-Coder-Next 80B-A3BRedHatAI/Qwen3-Coder-Next-NVFP4 @ 27a8f16fsglangv0.5.18-cu130NVFP4 + fp8 KVoff4119.931.6
1.06×
2026-09-06
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1e, repacked to GGUFllama.cppv0.3.0 nativeNVFP4off4126.732.6
1.86×
2026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1esglangv0.5.18-cu130NVFP4 + fp8 KVoff4189.248.1
2.78×
2026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1esglangv0.5.18-cu130NVFP4 + fp8 KVnextn4256.8
3.78×
2026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1evllmv0.28.0 (upstream image)NVFP4off4193.849.5
2.85×
2026-09-04
Qwen3.6-35B-A3Bnvidia/Qwen3.6-35B-A3B-NVFP4 @ 491c2f1evllmv0.28.0 (upstream image)NVFP4mtp4285.5
4.20×
2026-09-04