
What happened
Writing on Qiita, a practitioner whose demand model broke in production — cross-validation had leaked future time-series data — named four books to read in order: Andriy Burkov's "Machine Learning 100+ Pages Essence", the third edition of "Python Machine Learning" by Sebastian Raschka and Vahid Mirjalili, the PyTorch & scikit-learn edition by Raschka, Yuxi Liu and Mirjalili, and Hastie, Tibshirani and Friedman's "The Elements of Statistical Learning".
Why it matters
He argues the order itself matters, and that a thin overview followed by real code beats a thick theory text first, so a reader who can already write scikit-learn code but treats it as a black box is the intended audience.
WHO IT HITSThis lands on working data scientists and analysts who can already run scikit-learn or PyTorch code but cannot explain why a model behaves as it does on live data — the post frames the gap as conceptual, not a coding-skill gap.
Summaries like this, in your inbox every morning.
The post opens with a specific failure rather than a book list. The author was solely responsible for an internal demand-forecasting model; its score on validation data was fine, but predictions went wrong in production, and re-reading the code turned up no bug. The cause was the way the cross-validation was split: information from the future of the time series had leaked into the training side. He notes that this kind of mistake is hard to find in fragmented online tutorials, which describe the steps but not why those steps are the steps — what he had been missing, he says, was not implementation skill but the thinking underneath it.
That diagnosis shapes the reading order he recommends. Open a thick theory book first and you give up, he writes; read only light introductory books and you cannot explain the subtle problems that show up in practice. So the sequence runs from a short map of terms and methods, to how algorithms and model evaluation actually work, to writing deep-learning training loops yourself, and finally to statistics. For each title he names the audience: the short overview is not for someone new to programming, the Raschka and Mirjalili book is not meant to be read cover to cover at roughly 1,000 pages, and "The Elements of Statistical Learning" is past 1,000 pages and should be picked through chapter by chapter rather than read whole. He also flags that the two Raschka books overlap on the basics, so if you buy both, the PyTorch edition should come first.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.