Astra Acing Tests Doesn't Mean Taking Over Jobs
I noticed an interesting detail... The launch of GPT-6 Astra had the loudest noise coming from several near-saturated scores. FrontierMath Tier 4 hit 97.6%, ARC-AGI-3 jumped from 7.8% in the previous generation to 99.9%, and ExploitBench hit 100% straight. At the same time, OpenAI shifted the narrative from "can chat" to "can operate computers," filling forms, modifying CRMs, writing code, running tests. It looks like a breakthrough. But I'm more worried this is another round of hyping benchmarks as product maturity.
These tests have boundaries. Math has standard answers, exploit development has success/failure, abstract reasoning has fixed feedback. Real work isn't like this. Contract review, customer tickets, data cleaning, online troubleshooting—the inputs are often dirty, fields are often wrong, and responsibilities are often unclear. Recently, I used open-source models for inference services and handled tables and contracts with WorkBuddy. I was deeply impressed by the gap between benchmark scores and production runs. In the screenshots, once merged cells get messed up, no matter how smart OCR is, it mixes jurisdiction courts and dispute resolution clauses into one column. You ask it to "summarize with one click," and it summarizes convincingly, but convincing doesn't mean it can be handed to legal.
Astra emphasizes symbolic world models, sounding like it abstracts environments into operable objects. This direction has value. But enterprise environments aren't symbolic worlds; they're mostly hybrids of semi-symbolic elements and human habits. A button position, an approver's verbal rule—these will interrupt automation. Actual implementation is still early.
After costs drop, bottlenecks become even more obvious. The report also mentions some numbers: burning $360 per question, GPQA Diamond still hits 94.9% under low compute settings, API estimated costs drop by about 37%. But in my tests, complex tasks are truly expensive due to long chains. Calling tools, retries, validation, manual acceptance, permission isolation, rollback plans—all these drive up costs. OSWorld 2.0 score of 72.6% already indicates desktop environments aren't saturated yet. A model being able to open a browser doesn't mean it can stably complete a cross-system operation.
Cybersecurity being labeled Critical shouldn't just be written as proof of capability. It's more of a warning. Being able to find vulnerabilities autonomously and operate real systems means permission models must be redone. I just started with FSD and am not qualified to speak at length, but issues with black-box decision-making and real-time takeover are already obvious. If Agents operating computers autonomously make errors, how is liability defined, how are logs traced, do users dare give them write permissions—these determine whether they can enter production more than benchmark scores.
I previously wrote a cold splash of water regarding data centers, targeting the narrative of packaging compute, models, and test scores as universal benefits. Astra is the same this time. Scores are pretty, narratives are bigger, but productization is still separated by acceptance lists, responsibility boundaries, and dirty data.
If it can truly connect symbolic world models to continuous tasks, I'll watch closely. For now, don't treat near-perfect scores as humans exiting the loop. See if it can run continuously for a week at my client's site without crashing, then we can talk about AGI.
Physix Frontier