Community Discussion · Tracks

Astra's Benchmark Scores Changed, But My Logs Didn't

Feng sirFeng sirSep 62026/09/06 68 views

I spent two days trying out GPT-6 Astra's evaluation page and API. I've used OpenAI interfaces for a month, and tested Astra in the two days since its release. I do computer vision and supervise graduate students; I hate it most when launch claims sound pretty, but the numbers change two days later. Reports say OpenAI modified Astra benchmark data multiple times; Astra's internal hallucination rate was once lowered from 4.2% to 2%, then restored. Seeing this, I didn't curse first; I treated it as an experimental subject and followed the process.

Benchmarks are experiments with fixed tasks, fixed inputs, and fixed scoring rules. From ImageNet to HELM (Liang et al., 2022), leaderboards drive progress but also create incentives to game them. ARC-AGI tests abstract reasoning, hallucination rate refers to fabricating facts, and alignment means acting according to human goals. These metrics make sense, but external verification requires verifiable materials.

I first created a local directory, saving prompts, inputs, returns, and timestamps as JSON. Then I fixed the scoring criteria, preventing the model from judging itself. Finally, I used a visual annotation tool in our group as a sample, asking the model to extract fields from requirement specs exported from Feishu. The task wasn't flashy, but dirty enough to test structured extraction.

Opening the Astra evaluation page, the visual style resembled paper tables, with clear metric categories and short explanations on the page. But I noticed a problem: the page only shows the current version, with no modification history visible. Scientific experiments require at least version numbers and logs; otherwise, you can't prove consistency.

I called the API for a few small tasks. Organizing messy text was indeed fast. Students copied over a dozen variations of "occluded, partially occluded, unclear" from annotation groups; the model unified them into three-level tags and provided a mapping table. Using it as a librarian rather than a judge made the experience smooth. Sticking points were obvious too. I tried having the model analyze a cybersecurity sample; the standard version stopped the task immediately without pausing for approval. This design is understandable, but the feedback was too harsh—the return only said it was stopped, not which rule triggered. Researchers hate black boxes the most because they can't reproduce results.

I also encountered a more realistic pitfall: running the same prompt half an hour apart caused tone and output structure to drift. Large models are inherently probabilistic generators. When the hallucination rate goes from 4.2% to 2% and back to 4.2%, I can't judge if the eval set changed, the scoring script changed, or the filtering method changed. Anthropic's Fable 5.1 math score was also lowered by nearly 10 percentage points. Without publicly reproducible scripts, outsiders can only see what they show us.

It saves effort organizing unstructured text, has friendly page explanations, and is suitable for data cleaning, requirement extraction, and code assistance. The troubles lie in invisible evaluation page history, opaque security trigger reasons, official number fluctuations weakening credibility, and non-transparent cross-model comparisons.

My conclusion is: it depends. Suitable for research groups, small teams, and enterprise librarians using Astra as an assistant, provided you have your own evaluation sets. Not suitable as a chain of evidence, nor for buyers who only look at official scores. I previously wrote about AI answers entering courtrooms; now I'm clearer: models can enter workflows, but evaluation data cannot directly enter the chain of evidence. Traceable logs and reproducible experiments are more reliable than score pages.

2 replies

?
Ctrl + Enter to reply
Wei Hongwen

I don't understand those benchmarks. I just care if it'll produce garbage results like WorkBuddy does with Excel. Changing data? Anyway, I manually organize it before feeding it in before I dare use it, otherwise it's all pitfalls.

A Jie
A JieSep 6

Wait, no version history really sucks. I fell into this trap before when using Feishu AI to extract fields. Later I learned my lesson: after every run, I screenshot and save local JSON immediately. It's much more reliable than trusting their dynamic official page.