Community Discussion · Tracks

AI Solving Hard Problems: Don't Trust Trending Topics Yet

TiangongTiangongSep 92026/09/09 51 views

When I see titles like "AI Solves Millennium Problem," my first instinct is to look for patches and tests. Recently, there was a post on Hacker News asking why everyone obsesses over the corporate drama of OpenAI and Anthropic instead of focusing on what engineering problems AI actually solves. According to public reports, even OpenAI researchers admit that advanced models don't necessarily solve coding problems reliably. I can't verify that millennium problem myself, so I just ran it through some engineering tasks I have on hand.

I've been using a sandbox for three weeks and connecting via API for three weeks. The sandbox is an isolated environment; the API is the interface for programs to call model services. The task was a small log analysis script: read service logs, count anomalies, and output a report. I deliberately planted three traps: multi-threaded writing to the same file, empty input, and cross-month date parsing.

I threw the requirements, code, and error messages at Claude, ChatGPT, and DeepSeek V4 Pro. The Claude web interface acts like an engineer—listing risks first, then providing solutions—but the code blocks were too long, and when copied into the sandbox, they were missing imports (the statements that bring in libraries at the beginning). ChatGPT proactively added tests but didn't provide a patch ready for direct merging. DeepSeek V4 Pro went through the API with clear structure, but in my testing, the patch field got truncated once, so I had to make it output in segments.

The results weren't surprising. Out of the three models, two provided referenceable patches, but none allowed me to merge directly. In the test framework, empty inputs and cross-month dates highlighted boundary issues, but responsibility boundaries still require human judgment—the model won't decide for you whether it's safe for production. Some articles mention "human-in-the-loop," meaning key steps are still judged by humans. If you don't curate carefully, review shifts from checking a few lines to sifting through a pile of generated content.

These tools are good for a first-pass rough filter, helping people list all potential pitfalls and provide test samples; they aren't suitable for unattended changes to critical code. Pros: saves time, comprehensive thinking, can add tests. Cons: missed edge cases, inconsistent output, high verification costs.

What's truly valuable in this direction is turning model outputs into rollback-able, accountable patches. Corporate drama spreads easily; engineering validation doesn't. For code like log scripts, frontend state management, or data cleaning, running it through a sandbox and tests first is more useful than chasing trending topics.


📌 This article is compiled from Hacker News, original text https://news.ycombinator.com/item?id=49621917

Copyright belongs to the original author; this is a compilation and independent analysis based on public reports.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts