Physix Frontier · News Briefing Card (Hacker News · Oct 1, 2026)
AI agent training data gap: open-ended tasks missing
KEY FACTS
- The article points out that current AI agent training data is mostly structured, single-objective tasks, lacking the open-ended data found in real work scenarios.
- Opus 5.5 and GPT-6 Astra claim scores of 72.6% and 81.8% respectively on OSWorld 2.0.
- Computer-use agents rely on trajectory data for training, where trajectories record the detailed steps of a human or agent operating a computer.
- Data quality scales along two dimensions, observability and distribution, with observability improving faster than distribution.
- Real work involves interleaved multitasking, interruption recovery, and context switching, and existing data does not capture these signals.
KEY DATA
72.6%Opus 5.5 on OSWorld 2.0
81.8%GPT-6 Astra on OSWorld
PHYSIX OBSERVATION
Models post flashy scores on structured benchmarks, but real knowledge work is full of interruptions, multithreading, and ambiguous goals, and existing trajectory data barely covers any of it. This means agents will still stumble frequently when deployed in office scenarios. Whoever cracks the collection and annotation of open-ended work data first will secure the ticket into the next stage of enterprise-grade agents.
Source: Hacker News report
Physix Frontier