
What happened
Engineer Takashiba, who works on business systems, built a shogi AI on a single RTX 3060 using the dlshogi approach, documenting code and results publicly. His v2 net hit a 52.3% move-match rate and beat v1 31-9.
Why it matters
He abandoned AlphaZero-style self-play because it needs massive compute, and instead used supervised learning on strong-AI game records, showing a viable path for individuals without specialist training to build a competitive game AI.
What to watch
Strength still hinges on search speed, not just the net: after roughly quadrupling positions read per second to about 5,500-7,000, the same net won 18 of 20 games. Watch whether he applies the same recipe to poker next.
WHO IT HITSIndividual hobbyist and non-ML software engineers can replicate this pipeline, since the whole method, the cshogi library, and the code are public, lowering the barrier to entry for building game AIs at home.
Summaries like this, in your inbox every morning.
The author, a software engineer who normally builds business systems, openly states that machine learning is not his specialty, and his aim was to see how far an individual could take a shogi AI on a single consumer GPU. He began with AlphaZero in mind but found it infeasible at home, since zero-to-hero self-play requires thousands of TPUs and millions of games. That led him to dlshogi, whose supervised approach uses game records from strong AIs and looks like image classification plus a Monte Carlo tree search on top.
His first network, v1, reached a 50.3% move-match rate but its value loss ran 0.30 on training versus 0.59 on tests, a clear sign of overfitting. The cause was in the labels: every position in a single game carried the same final result, so with far more positions than games, the network learned to recognize the game itself rather than judge the position. Mixing each move's recorded evaluation score, converted into a win rate, into the value target, along with a larger 15-block, 192-channel net, brought the test and training losses much closer at 0.475 and 0.44.
The search side is where the biggest jump came. Reading about 1,450 positions per second at first, the author found the bottleneck was not GPU compute but the per-iteration overhead of dispatching commands, selecting moves, and writing results back. Batching positions with a virtual loss, CUDA graphs, fp16, and an overlap of GPU compute with CPU gathering pushed throughput to about 5,500-7,000 positions per second, and the same net then beat its old self 18 games to 2. The outcome appears to hinge less on the network than on how efficiently search uses it, a reading that matters most to individual builders trying to squeeze strength from modest hardware.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
A developer published Rai, a small Rust engine that runs language models on an ordinary PC's CPU, using only t…

The developer published fk2000/minecraft-ai-bot, which uses Jev to assemble instructions, Gemini to pick from…

A developer released cameo, a roughly 400-line Go program that reads a TOML file, starts a local HTTP proxy, a…

Meta announced Muse Gadgets, an open-source project with ESP32 firmware and a Linux SDK (a toolkit for develop…

A Reddit user applying to academic jobs asked whether HuggingFace download counts for custom models count as "…

The AI Security Lab outlined its standard experiment setup: Python programs call Ollama's API directly to run…
