Syndicated from the original on lkforge.com. The engines are playable in your browser at lkforge.com/games; the harness that produced these numbers is public and seeded.
Every "AI" in a product now seems to mean a large language model. The AI that plays against you on my site doesn't — it's classical game-tree search: minimax, expectimax, breadth-first search. That's a deliberate engineering choice, and it's the difference between an opponent that's provably correct and instant and one that's plausible and slow.
The core point
My tic-tac-toe engine returns a provably-optimal move in about 0.3 ms, on your device, with zero network calls — and it has lost 0 of 1,200 test games. Those are properties a language model, by construction, cannot offer: determinism, a correctness proof, and sub-frame latency without a server.
"Why not just use an LLM?"
Fair question in 2026 — you could prompt a model with the board and ask for a move. The reason I don't: a language model is trained to predict the next token of text, not to search a game tree. It can explain tic-tac-toe strategy fluently and still play a losing move, because fluent text and optimal play are different objectives. Winning a solved game is a search problem, and we already have exact, fast algorithms for it.
The three engines — minimax + alpha-beta for tic-tac-toe, expectimax for 2048, and BFS for Color Lines — are textbook, deterministic, and return a move in a couple of milliseconds or less in a browser tab.
Search vs. a language model, point by point
| Game-tree search (mine) | A language model | |
|---|---|---|
| Decides a move by | searching the tree of legal positions | predicting likely next tokens |
| Correctness | provable at full depth | none — fluent ≠ optimal |
| Same board → | same move (deterministic) | varies with sampling/phrasing |
| Latency | a few milliseconds or less, on-device | a network round-trip |
| Needs a server | no | yes |
Every row is an architectural difference — how each system decides — not a quoted benchmark. The only measured numbers here are mine.
The payoff: a strength number you can actually pin down
Because the engines are deterministic, I can put an exact figure on how strong they are — run the shipped code headlessly, hundreds of times, and count. That's far harder for a model whose output shifts with sampling and phrasing.
2048 solver, 250 self-play games: 92.8% of games reach the 2048 tile, 64% reach 4096, and 10.8% of the 250 reached 8192 — a corner-snake expectimax search with probability-threshold pruning at ~2 ms/move. A number, with error bars you could compute, precisely because the same board always drives the same search.
Tic-tac-toe is the cleaner case: full-depth minimax is provably optimal, so "unbeatable" is a theorem, not a vibe. Across 1,200 self-play games (1,000 vs random, 200 vs a perfect copy) it lost none. Alpha-beta keeps full depth cheap: 36,528 nodes instead of 549,945 at the opening move — a 93% cut — in about 0.3 ms.
The right tool, not the trendy one
None of this is anti-LLM. Language models are extraordinary at language — and a couple of the tools on my site that are genuinely language tasks could use one. But a board game with fixed rules and a finite tree is exactly the problem classical search was invented for.
Full write-up with charts: *lkforge.com/blog/game-ai-not-llms*. Related: Six Games, Three Classic Algorithms · Minimax & Alpha-Beta, Visualized · and the companion experiment, We Asked ChatGPT and Grok to Benchmark Our Game AI.
Top comments (2)
I completely agree that using LLMs for real‑time opponents can feel like overkill—latency and unpredictable behavior often break immersion. In my own projects, we’ve found hybrid approaches, where a lightweight decision tree handles core tactics while an LLM enriches dialogue, striking a good balance. Have you experimented with any modular pipelines that keep the gameplay tight but still add narrative depth?
Exactly the right split. In the LK Forge games the engine is all there is — expectimax / MCTS / minimax deciding moves — and I deliberately skipped any narrative layer, so the opponents don't talk and there's no LLM in the loop at all.
Your hybrid is where I'd add one though: an async dialogue layer that only reads game state and never votes on the move — keeps latency and determinism in the core. Have you had the commentary drift out of sync with the actual position?