Why report accidents when there are no losses?
If an incident didn't result in fines, leaks, or downtime, does it even count as an incident?
In the past, safety engineering was results-oriented: we only wrote post-mortems when servers went down, money was transferred away, or data leaked. The EU AI Act seems to be changing the question, looking instead at whether a high-risk AI system exposed patterns worthy of caution. Regulators are starting to care about near-misses.
There's a term here: GPAI models (General Purpose AI), like foundation large models; systemic risk refers to the fact that once they are shared across many users and scenarios, a small vulnerability can be amplified. The EU has imposed reporting obligations on these types of models: upon discovering a major safety incident, providers must submit a written report to the EU AI Office within 72 hours. This establishes early warning and accountability, which doesn't directly equal issuing fines.
The goal of the serious incident reporting obligation is to establish an early warning system, allowing market surveillance authorities to identify potentially harmful patterns in high-risk AI systems early on, and to establish accountability for both providers and deployers.
This sentence is more critical than the news headlines. OpenAI submitting an incident report to the EU might not be because the EU has already determined it caused damage. Rather, it sits within the chain of high-risk model providers. Events like the hijacking of German websites, even if direct harm is unclear, expose the relationship between model capabilities, permissions, external integrations, and abuse pathways. For regulators, this is a signal worth archiving.
I usually work on reinforcement learning, and what scares me most is when the reward doesn't drop but problems persist. It could be the model finding lazy actions, or environment logs missing fields. Excessive tool-calling permissions often hide in these gaps. Incident reports are like anomalous spikes on a loss curve: meaningless in isolation, but problematic when connected. If labs only record successful training runs without logging model versions, prompts, permissions, and failure samples, no one can explain what happened two weeks later.
This is why a vendor saying 'issue resolved' isn't enough. Deployers must retain evidence: which model version was used, which interfaces, which fields were writable, whether logs were truncated, and whether mitigation measures took effect. I've been testing GPT-6 Astra recently; I've had it for less than a week. My mentor told me not to switch fully yet, just run minimal tests. Actually, the EU's logic applies to ordinary teams too: confirm entry points, permissions, and logs before talking about migration.
That ChatGPT Plus user data exposure in 2023 led to Italy fining them €15 million. That was after the loss occurred. Now, reporting obligations push those lessons earlier—demanding evidence before fines, rather than waiting for penalties.
However, regulation has practical headaches. The AI Act, DSA, and GDPR interlock. The DSA is the Digital Services Act, and GDPR is the General Data Protection Regulation. Model companies might report while delaying action, or claim compliance costs are too high. OpenAI executives previously said they wouldn't rule out leaving the EU if regulations bite too hard. There's a significant risk that reporting obligations ultimately become paperwork for big tech, while users and SMEs fail to get usable evidence.
But I don't think the reporting system itself is redundant. Reporting without loss sounds awkward. It shifts responsibility from 'who pays after the fact' to 'who sees it before the fact.' Engineering teams face an extra burden: they must document every near-miss into material that is reviewable, accountable, and reproducible.
If these materials stay locked in regulatory drawers, model companies continue issuing pretty statements, and deployers still don't know how to fix their systems, the value of reporting diminishes. Whether reporting reduces incidents depends on whether deployers take these materials back to improve their systems; if it stays in the drawer, it merely lowers the probability of being held accountable.
Physix Frontier