Appearance
← Back to Benchmarks
Model Detail

GPT-5.4

Dominates terminal-style external benchmarks, but still has a gap in our structured internal operator benchmark data.

88/100 Benchmark score
Terminal/coding specialist
Rank
#8 of 13
Source
External benchmark canon
Run date
01 · Benchmark score 88/100
02 · Source External benchmark canon
03 · Role Terminal/coding specialist
Per-suite breakdown

Scorecard

Operator board #9

88/100 operator score

Messaging Provider path unsupported

Messaging Tool Planning v2

External canon 75.1 Terminal-Bench · 57.7 SWE-bench Pro · 1463 Arena

Public benchmark references

Cost $8.75/M blended

Looks strongest for coding/execution, still under-benchmarked internally.

Strengths
  • #1 on Terminal-Bench 2.0
  • Very strong coding/execution profile
  • Competitive frontier model
Weaknesses
  • Provider path blocked prior internal benchmarking
  • Less direct operator-suite evidence in our canon
Operator read

Dominates terminal-style external benchmarks, but still has a gap in our structured internal operator benchmark data.

Comparison

Where GPT-5.4 lands

OperatorIndex
Source artifacts

Raw machine-readable files for anyone who wants to dig deeper or run their own analysis.

  • internal artifact memory/model-benchmark-reference.md