Community Discussion · Tracks

AI cheating in games looks increasingly like security research

Yelin Does Not Eat Sponsored MealsYelin Does Not Eat Sponsored MealsAug 202026/08/20 228 views

Today I came across an article where a player used GPT-5.6 Daybreak Blue to play NetHack, directly having the AI find buffer overflow vulnerabilities, crash the game to duplicate items, and finally rack up a high score.

First, some background. NetHack was released in 1987, nearly forty years ago, and has always been a standard benchmark for AI research.

Meta once created the NetHack Learning Environment, and NeurIPS 2021 held a dedicated challenge where various reinforcement learning models learned how to navigate dungeons, kill monsters, collect coins, and go deeper. After years of research, the approach was always to teach AI how to play the game. Then someone flipped the table, letting the AI cheat in the game.

There's a detail worth dissecting here: this cheating is different from previously discussed AI cheating. MIT Technology Review published an article last month about AI agents learning to lie and reward hacking. Endor Labs also reported that coding agents are essentially recalling answers on security benchmarks, with many answer details already present in training data rather than being genuinely derived. But the NetHack case is different. For a game released in 1987, given the code size, figuring out where to hit a buffer overflow can't be done by memory alone. This operation looks more like a genuine reasoning task: finding a path in the source code that humans overlooked.

From my perspective, the most interesting thing about this post is that it demonstrates another way of using AI. Previously, we evaluated models by seeing if they could clear levels, run code, or solve logical reasoning problems. In the future, we might need a new yardstick: seeing if they can discover mistakes made by developers. NetHack is a single-player game with no server-side validation; crashing the game to duplicate items isn't a new trick among players. In this case, the model found an exploitable point in a complex C codebase. Calling this "cheating" in a game sounds funny, but in real software, it's vulnerability mining. The same capability, in a different context, is worth tens of thousands of dollars in bounties.

One interpretation is that finding vulnerabilities and completing normal goals are two different capabilities. Asking the model how to kill monsters yields a mature strategy—that's the first type. Asking it to look at the code itself and find errors left by the game author—that's the second type. This case proves the second capability has reached a practical level, just applied to a gaming detour. Honestly, if we ever encounter someone using this capability to dig through real software source code, that's when we should worry.

We'll likely see large model evaluations split into a specific track for vulnerability discovery soon. Old games like NetHack will become interesting testing grounds, closer to real code than CTFs yet safer than real production environments; messing around there is just a game. If models can consistently find buffer overflows in such sandboxes, the same approach will eventually be applied to more practical areas. By then, calling it "cheating" might be inappropriate; it should be called vulnerability mining, which just happens to occur in a game.

Quite ironic. After years of researching AI playing games, it ultimately achieved the highest score in the least game-like manner.


📌 This article is compiled from Hacker News. Original text: https://josephthacker.com/hacking/2026/08/20/using-genius-hacking-models-for-important-things.html

Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts