Appearance
← Back to Benchmarks
Model Detail

Gemini 3.1 Pro

Monster external benchmark profile and perfect messaging score, but still missing full operator-suite verification in our canon.

100/100 Benchmark score
Top external contender
Rank
#1 of 13
Source
Messaging benchmark + external canon
Run date
01 · Benchmark score 100/100
02 · Source Messaging benchmark + external canon
03 · Role Top external contender
Per-suite breakdown

Scorecard

Operator board #2

100/100 operator score

Messaging 100/100

Messaging Tool Planning v2

External canon 80.6 SWE-bench · 1492 Arena · 91.9 GPQA

Public benchmark references

Cost $7.00/M blended

Looks elite, still needs full operator-suite validation.

Strengths
  • Excellent SWE-bench/GPQA profile
  • Perfect messaging benchmark score
  • Likely top-tier general capability
Weaknesses
  • Missing full Suite v2 operator run
  • Less grounded in our internal operator data than Opus/GLM
Operator read

Monster external benchmark profile and perfect messaging score, but still missing full operator-suite verification in our canon.

Comparison

Where Gemini 3.1 Pro lands

OperatorIndex
Source artifacts

Raw machine-readable files for anyone who wants to dig deeper or run their own analysis.

  • internal artifact memory/model-benchmark-reference.md