Community Discussion · Policy
Behind 188% Score Boost: What Number Game Is OpenAI Playing?
I noticed an interesting detail: this evaluation, touted as the "most rigorous AI interactive reasoning test," suddenly gave GPT-5.6 Sol a score improvement of 188%. The number looks pretty, but the question arises—who set the standards? How was it tested? What specifically were the improvements? I looked through IT Home's reports, and the more I read, the less simple this seems.
Physix Frontier