
LLMs Playing Chess: Don't Ask About Skill Level First
A friend recommended Chess5.ai, so I tried to see if it's actually good.
Conclusion first: Recommended as an LLM behavior observer, not recommended as a chess skill tester. LLM stands for Large Language Model; a behavior observer watches how it outputs and makes mistakes. It's suitable for people researching agents, rule validation, and multi-model comparisons; an agent can be simply understood as a program letting the model read state, choose actions, and write back. It's not suitable for seriously evaluating chess skills. Depends on the situation.
I clicked around a few models on the website over the last two days. The homepage features Chess, Go, Xiangqi, Gomoku, and Othello, with opponents listed including GPT, Claude, Qwen, DeepSeek, etc. I only tested Gomoku and Xiangqi with Qwen and DeepSeek, and tried a small board game of Go.
The interface is plainer than I expected. Left side is the board, right side is for selecting moves. The model can only pick from legal actions; it can't freely output "I play here." Illegal moves are blocked, which actually made me feel it was usable.
Pitfalls were more obvious than imagined. In Gomoku, Qwen sets up appearances in the opening, but in the mid-game, it tends to focus only on connecting its own lines without blocking the opponent. Occasionally, it gives coordinates that aren't in legal empty spots, prompting the frontend to ask for a re-selection. Xiangqi is more obvious: DeepSeek forgets the "blocked horse leg" when moving horses and misses the "mountain crossing" rule for cannons. Small Go boards are manageable, but once the board widens, it starts losing context.
I did a rough comparison using local llama.cpp and Qwen-7B, converting the board to text state, having the model output coordinates, and then validating with Python. Bare models aren't good at maintaining the board themselves; they forget previous moves as the state gets long. The value of Chess5.ai is packaging state, action space, legality checks, and UI together. The bottleneck is mostly context representation and action constraints.
| Game | Common Issues | What to Observe |
|---|---|---|
| Gomoku | Unstable defensive awareness | Whether the action space is small enough |
| Xiangqi | Horse legs, cannon mountain crossings often wrong | Whether rules can block hallucinations |
| Go | Misses boundaries and ko fights | Long-context state management |
| Othello | Forgets flip logic | Whether state updates are explicit |
From an engineering perspective, it doesn't do end-to-end chess fusion; it's more like bolting on a runtime safety gate. This idea is similar to the local agent safety gate I wrote about before: the model can propose actions, but the system must be able to reject them. Operator fusion for board games hasn't been done yet, nor is there a rush. Maintaining the Board IR separately—IR being the intermediate representation of board state—and having the model only output candidate actions is more stable than letting the model count pieces, remember state, and judge wins/losses itself.
Disadvantages are also clear. It's not suitable for serious chess skill evaluation; model versions, prompts, and action selection methods all affect results. It's not friendly to beginners either, who might be fooled by the "LLM vs LLM" hype into thinking the model truly understands the game. My tests showed waiting times, especially with many models listed; the experience depends on the backend API, i.e., the model service interface. Someone tested with 13 models, and the conclusion was cold: they are mostly just moving pieces; winning games is still off the table.
Advantages are practical. It puts multiple games, multiple models, and legal move validation into one entry point, saving you from building your own test shell. For those wanting to research agents, this is more intuitive than looking at benchmarks; you can clearly see when the model hallucinates, when it obeys constraints, and when it needs to roll back.
Suitable for people working on AI toolchains, agent evaluation, local inference, and prompt engineering. Chess enthusiasts can treat it as fun, but don't take it seriously. Not suitable for those wanting to systematically practice chess, nor for formal model capability benchmarking.
Its most useful aspect is putting large models inside board game rules, revealing exactly which layer of constraints they lack.
📌 This article is compiled from Hacker News, original text https://chess5.ai/en
Copyright belongs to the original author. This article is a compilation and independent analysis based on public reports.
Physix Frontier