Model Detail
Gemini 3.1 Pro
Monster external benchmark profile and perfect messaging score, but still missing full operator-suite verification in our canon.
Top external contender
- Rank
- #1 of 13
- Source
- Messaging benchmark + external canon
- Run date
01 · Benchmark score 100/100
02 · Source Messaging benchmark + external canon
03 · Role Top external contender
Per-suite breakdown
Scorecard
Operator board #2
100/100 operator score
Messaging 100/100
Messaging Tool Planning v2
External canon 80.6 SWE-bench · 1492 Arena · 91.9 GPQA
Public benchmark references
Cost $7.00/M blended
Looks elite, still needs full operator-suite validation.
Strengths
- Excellent SWE-bench/GPQA profile
- Perfect messaging benchmark score
- Likely top-tier general capability
Weaknesses
- Missing full Suite v2 operator run
- Less grounded in our internal operator data than Opus/GLM
Monster external benchmark profile and perfect messaging score, but still missing full operator-suite verification in our canon.
Comparison
OperatorIndex Where Gemini 3.1 Pro lands
- 01 Gemini 3.1 ProMessaging benchmark + external canon 100
- 01 Claude Sonnet 4.6Messaging benchmark + external canon 100
- 01 Gemini FlashMessaging benchmark canon 100
- 04 Claude Opus 4.6Operator Suite v2 95.3
- 05 GLM-5-TurboOperator Suite v2 95
- 06 Gemma 4 31BEnterprise Ollama 92
- 07 MiniMax M2.7Operator Suite v2 90.8
- 08 GPT-5.4External benchmark canon 88
- 09 Gemma 4 26BEnterprise Ollama 80
- 10 PrismML Bonsai 1.7BPrismML local benchmark 56
- 11 Qwen 3.6 27B NVFP4Enterprise oMLX 40
- 11 Qwen 3.6 27B MXFP4Enterprise oMLX 40
- 13 Qwen 3.6 27B 4bitEnterprise oMLX 20
Source artifacts
Raw machine-readable files for anyone who wants to dig deeper or run their own analysis.
- internal artifact memory/model-benchmark-reference.md