Community Discussion · Policy
Claude Opus 5: Shifting Evaluation Philosophy from "Strongest Model" to "Fewer Omissions"
At last week’s lab meeting, a second-year master’s student rushed in excitedly, saying that Anthropic’s Opus 5 had surpassed GPT-4o and Gemini Ultra on multiple benchmarks. He pulled up screenshots of the leaderboard and asked me, “Professor, should we switch the baseline model in our paper to Opus 5?” I didn’t answer directly; instead, I asked him to clarify exactly which metrics demonstrated Opus 5’s “strength.” He spent ages digging through the documentation but only found one vague statement: “Fewer missed steps on complex tasks.” This detail is precisely what makes the whole situation most worth pondering.
Physix Frontier