Appearance
← Back to Benchmarks
Model Detail

Claude Opus 4.6

Highest raw benchmark score in the canon, best when judgment quality is mission-critical.

95.3/100 Benchmark score
Benchmark king
Rank
#4 of 13
Source
Operator Suite v2
Run date
01 · Benchmark score 95.3/100
02 · Source Operator Suite v2
03 · Role Benchmark king
Per-suite breakdown

Scorecard

Operator board #4

95.3/100 operator score

Messaging 100/100

Messaging Tool Planning v2

External canon #1 SWE-bench · #1 Arena · #1 HLE

Public benchmark references

Cost $15/M blended

Best raw benchmark performer overall.

Strengths
  • Highest operator-suite score
  • Strongest recovery/config/delegation profile
  • Consistent external benchmark leader
Weaknesses
  • Expensive
  • Overkill for many routine tasks
  • Not the rational default when cost matters
Operator read

Highest raw benchmark score in the canon, best when judgment quality is mission-critical.

Comparison

Where Claude Opus 4.6 lands

OperatorIndex
Source artifacts

Raw machine-readable files for anyone who wants to dig deeper or run their own analysis.

  • internal artifact memory/model-benchmark-reference.md