DEV Community

Lucian (LKB)
Lucian (LKB)

Posted on Edited on Originally published at lkforge.com

Real game AI, not a chatbot: why these opponents don't use an LLM

Syndicated from the original on lkforge.com. The engines are playable in your browser at lkforge.com/games; the harness that produced these numbers is public and seeded.

Every "AI" in a product now seems to mean a large language model. The AI that plays against you on my site doesn't — it's classical game-tree search: minimax, expectimax, breadth-first search. That's a deliberate engineering choice, and it's the difference between an opponent that's provably correct and instant and one that's plausible and slow.

The core point

My tic-tac-toe engine returns a provably-optimal move in about 0.3 ms, on your device, with zero network calls — and it has lost 0 of 1,200 test games. Those are properties a language model, by construction, cannot offer: determinism, a correctness proof, and sub-frame latency without a server.

"Why not just use an LLM?"

Fair question in 2026 — you could prompt a model with the board and ask for a move. The reason I don't: a language model is trained to predict the next token of text, not to search a game tree. It can explain tic-tac-toe strategy fluently and still play a losing move, because fluent text and optimal play are different objectives. Winning a solved game is a search problem, and we already have exact, fast algorithms for it.

The three engines — minimax + alpha-beta for tic-tac-toe, expectimax for 2048, and BFS for Color Lines — are textbook, deterministic, and return a move in a couple of milliseconds or less in a browser tab.

Search vs. a language model, point by point

Game-tree search (mine) A language model
Decides a move by searching the tree of legal positions predicting likely next tokens
Correctness provable at full depth none — fluent ≠ optimal
Same board → same move (deterministic) varies with sampling/phrasing
Latency a few milliseconds or less, on-device a network round-trip
Needs a server no yes

Every row is an architectural difference — how each system decides — not a quoted benchmark. The only measured numbers here are mine.

The payoff: a strength number you can actually pin down

Because the engines are deterministic, I can put an exact figure on how strong they are — run the shipped code headlessly, hundreds of times, and count. That's far harder for a model whose output shifts with sampling and phrasing.

2048 solver, 250 self-play games: 92.8% of games reach the 2048 tile, 64% reach 4096, and 10.8% of the 250 reached 8192 — a corner-snake expectimax search with probability-threshold pruning at ~2 ms/move. A number, with error bars you could compute, precisely because the same board always drives the same search.

Tic-tac-toe is the cleaner case: full-depth minimax is provably optimal, so "unbeatable" is a theorem, not a vibe. Across 1,200 self-play games (1,000 vs random, 200 vs a perfect copy) it lost none. Alpha-beta keeps full depth cheap: 36,528 nodes instead of 549,945 at the opening move — a 93% cut — in about 0.3 ms.

The right tool, not the trendy one

None of this is anti-LLM. Language models are extraordinary at language — and a couple of the tools on my site that are genuinely language tasks could use one. But a board game with fixed rules and a finite tree is exactly the problem classical search was invented for.


Full write-up with charts: *lkforge.com/blog/game-ai-not-llms*. Related: Six Games, Three Classic Algorithms · Minimax & Alpha-Beta, Visualized · and the companion experiment, We Asked ChatGPT and Grok to Benchmark Our Game AI.

Top comments (2)

Collapse
 
seohyun0903 profile image
Seohyun Lee •

I completely agree that using LLMs for real‑time opponents can feel like overkill—latency and unpredictable behavior often break immersion. In my own projects, we’ve found hybrid approaches, where a lightweight decision tree handles core tactics while an LLM enriches dialogue, striking a good balance. Have you experimented with any modular pipelines that keep the gameplay tight but still add narrative depth?

Collapse
 
lucian_lkb_1f009d profile image
Lucian (LKB) •

Exactly the right split. In the LK Forge games the engine is all there is — expectimax / MCTS / minimax deciding moves — and I deliberately skipped any narrative layer, so the opponents don't talk and there's no LLM in the loop at all.

Your hybrid is where I'd add one though: an async dialogue layer that only reads game state and never votes on the move — keeps latency and determinism in the core. Have you had the commentary drift out of sync with the actual position?