Community Discussion · Tracks
As AI writes more code, who verifies it's correct?
Just saw the news about the Vero benchmark. My first reaction was, "Here we go again." There are as many AI coding benchmarks now as stalls in a wet market, all claiming rigorous evaluation standards. But after digging into the details, Vero does seem different. It's the first benchmark to evaluate code generation alongside formal proofs. Simply put, it doesn't just check if your code is correct; it checks if you have the ability to prove that it's correct.
Physix Frontier