Model Detail
Gemma 4 31B
The better local model when quality matters more than speed and concurrency.
Premium local escalation
- Rank
- #6 of 13
- Source
- Enterprise Ollama
- Run date
- Avg latency
- 22.59s
01 · Benchmark score 92/100
02 · Average latency 22.59s
03 · Role Escalation local model
Per-suite breakdown
Scorecard
Operator board #6
92/100 operator score
Messaging Quick execution pack
Messaging Tool Planning v2
External canon Fast local quality leader in current Gemma run
Public benchmark references
Cost Enterprise local
Best local quality of the Gemma pair, but significantly slower.
Strengths
- Best quality in the Gemma local pair
- Better routing judgment on the benchmark pack
- Cleaner concise answers overall
Weaknesses
- Much slower average latency
- Worse fit for high-concurrency local loops
- Still wrapped strict JSON in code fences
Stronger output quality and better routing judgment than 26B, but roughly 3x slower in this quick benchmark pack.
Comparison
OperatorIndex Where Gemma 4 31B lands
- 01 Gemini 3.1 ProMessaging benchmark + external canon 100
- 01 Claude Sonnet 4.6Messaging benchmark + external canon 100
- 01 Gemini FlashMessaging benchmark canon 100
- 04 Claude Opus 4.6Operator Suite v2 95.3
- 05 GLM-5-TurboOperator Suite v2 95
- 06 Gemma 4 31BEnterprise Ollama 92
- 07 MiniMax M2.7Operator Suite v2 90.8
- 08 GPT-5.4External benchmark canon 88
- 09 Gemma 4 26BEnterprise Ollama 80
- 10 PrismML Bonsai 1.7BPrismML local benchmark 56
- 11 Qwen 3.6 27B NVFP4Enterprise oMLX 40
- 11 Qwen 3.6 27B MXFP4Enterprise oMLX 40
- 13 Qwen 3.6 27B 4bitEnterprise oMLX 20
Source artifacts
Raw machine-readable files for anyone who wants to dig deeper or run their own analysis.
- internal artifact output/benchmarks/2026-04-12-gemma-enterprise-benchmark/summary.md
- internal artifact output/benchmarks/2026-04-12-gemma-enterprise-benchmark/results-raw.json
- internal artifact output/benchmarks/2026-04-12-gemma-enterprise-benchmark/results-scored.json