Appearance

Qwen on a GB10: 197 to 779 tok/s With Receipts, Not Headlines

A controlled inference-tuning matrix for Qwen3.6-35B on one GB10: 3.93x baseline throughput with raw receipts, prompt-length boundaries, and a five-point integrity checklist.

Listen to this post
00:00
Browser TTS · Ada voice

We ran a controlled inference-tuning matrix for Qwen3.6-35B-A3B-NVFP4 on one NVIDIA GB10. Baseline: 197.5517 tok/s. After tuning, three consecutive c64x512 runs returned 774.5402, 779.0010, and 778.4611 tok/s. Mean: 777.3341 tok/s, 3.93x baseline.

That number applies to 99 to 114 prompt tokens. At 1,335 to 1,505 prompt tokens, the same configuration returns 521.4124 tok/s.

The benchmark contract

  • Workload: c64x512 (64 concurrent, 512 completion tokens, 99 to 114 prompt tokens).
  • Numerator: server-reported completion tokens from raw SSE usage.
  • Denominator: full client wall time. No trimming.
  • Distinctness: 64 distinct request hashes and content hashes per run.
  • Repeatability: three consecutive runs within a 0.24s band.

Winning configuration

FP8 KV cache, FlashInfer attention, Marlin MoE backend
GPU util 0.6, seq64, batch32768
Prefix caching: disabled
MTP-1 with Triton drafter MoE

Three qualifying runs

RunCompletion tokensWall timeThroughput
132,76842.306s774.5402 tok/s
232,76842.064s779.0010 tok/s
332,76842.093s778.4611 tok/s

Upstream recipe and setup: https://github.com/Weschera/spark-bench at 45491e80a22f4701d8f3a82ad0b5ce7fa9597089.

← Back to Ship Log

Keep reading

All posts
agents

GLM-5.2 on Two DGX Sparks: A Seven-Gate Bring-Up

Loading a 753B parameter model across two DGX Sparks exposed seven gates: artifact fit, tensor audit, compiler canaries, rejected optimization branches, content canaries, and throughput measurement.

local models

Local Models Are Finally Useful. Just Not in the Way Most People Think.

A tiny 1.7B local model beating tested Gemma variants on an operator benchmark is not a frontier-model story. It is a routing story.

Dispatches from the Edge

Dispatches from the Edge #16

OpenAI pulled GPT-6.1 Astra for failing its own safety bar, a nonprofit sued over the Hugging Face attack, and Gemini 4 Argon ships through a cyber-defender gate before anyone else.

Ship signal

Get the next one when it ships.

Subscribe