
A widely-shared essay argues that modern large language models combine two distinct jobs—storing facts and reasoning about them—into one system, limiting both capabilities. Rather than scaling up a single "know-it-all" model further, the author suggests building specialized subsystems: one optimized for accurate fact retrieval, another for learned rules and generalization, and a coordinator to route questions intelligently. While complete separation is unlikely given how knowledge and reasoning are intertwined in current training, even a partial move toward specialization could unlock the next real advance in AI.
Summaries like this, in your inbox every morning.
Sign up free →What happened
An essay argues that current large language models conflate two fundamentally different jobs—storing facts (like a library) and reasoning over those facts (like a librarian)—into a single system, forcing each capability to underperform. The author uses the example of a model confidently producing nonsensical answers by assembling individually correct pieces in illogical ways, despite having all the needed knowledge.
Why it matters
Today's AI is trained to memorize vast amounts of text and simultaneously reason about novel problems, much like asking a single book to both sit on a shelf and personally interpret every other book. This design may be holding back progress; the essay suggests we've over-invested in scaling the 'library' (more parameters, more training data) while under-investing in the 'librarian' (genuine generalization and rule-learning). Separating these functions could unlock better performance than forcing one model to do both.
What to watch
The essay proposes an alternative architecture: a cluster of specialized subsystems—a knowledge store for accurate fact retrieval, a rule store for learned principles, and an orchestrator that routes questions to the right specialist. While the author acknowledges that knowledge and intelligence are tangled in current models and complete separation is unlikely, even partial separation might point toward the next meaningful leap in AI capability.
The essay opens with a concrete failure: a model confidently produced a nonsensical answer by assembling individually correct pieces in an illogical order, not because it lacked knowledge but because it could not apply a rule it had learned to a situation outside its memorized experience. This anecdote sets up the central claim: large language models are being asked to perform two fundamentally different jobs—knowledge storage and reasoning—mashed into a single system, compromising both.
The author develops this insight through the metaphor of a library and a librarian. A library stores facts faithfully; a librarian navigates questions, breaks them into parts, connects disparate information, and answers novel questions without having memorized every book. Current large language models attempt both roles simultaneously: they memorize "enormous swaths of the internet" into their parameters while also being expected to reason over that material to answer questions they have never seen verbatim. The essay notes this is an odd design choice—we would never build a physical library where every book must also personally reason about every other book.
The piece examines why the scaling-focused strategy, which has dominated AI development for years, may be reaching its limit. The internet already functions as an enormous, constantly-updating library. Models' internal memories, by contrast, are frozen at their training cutoff date. Retrieval-augmented systems—which search external sources rather than relying purely on memorized weights—have become central to real-world AI deployment, tacitly admitting that the internet is already a better library than any model's internal memory could be. The author asks: if the library problem is essentially solved by pointing at something that already exists, where should the next leap in intelligence come from? The answer: not from a bigger library, but from a better librarian.
The distinction crystallizes through the comparison of two children learning multiplication. One memorizes the times table; asked "what is 47 times 83," they fail because that combination was never in their memorized set. The other understands multiplication's underlying rule—repeated addition, scaling—and can work out the unfamiliar problem. Large language models, the essay argues, often resemble the first child: they excel at producing correct answers when questions resemble their training data, but crack when pushed outside that shape. What looked like understanding was, underneath, "an extremely sophisticated form of memorization and interpolation." The author distinguishes between knowledge (retrieval) and intelligence (generalization), arguing these are not the same capability and training one system to do both at once, the same way, weakens both outcomes.
The proposed alternative is architectural: instead of one monolithic model, build a cluster of specialized subsystems. A knowledge cluster, optimized like a library, stores and retrieves facts faithfully, stays current, and covers broad ground—it does not need to be clever, just accurate and comprehensive. A rule cluster, by contrast, does not need to be enormous; a system that has genuinely learned "the rules of arithmetic" or "the grammar of a language" is small and tight, winning through correctness and generality rather than size. An orchestrator—the actual librarian—sits above both, knowing which specialist to consult for any given question, how to combine answers, and when both a fact and a rule are needed together. The author notes this idea—that intelligence emerges from many specialized, cooperating parts rather than one monolith—has been explored in cognitive science and computing for decades, but building it in modern AI systems has finally become practical.
The essay's most honest moment acknowledges a real complication: knowledge and intelligence do not separate as cleanly in reality as in analogy. Part of why current large models can reason at all seems to come from exposure to the breadth of their training data; they pick up patterns of logic and structure, not just facts. Strip away knowledge entirely and you might strip away material the rule-learning was built from. The author therefore reframes the claim as narrower and more useful: we have almost certainly over-invested in growing the library past meaningful returns while under-investing in the librarian—the part responsible for generalizing, reasoning, and routing rather than recalling. Even an imperfect move toward separation seems pointed in the right direction. The essay closes by suggesting a reframing of the default question in AI: instead of "how do we make the model bigger," ask "how do we make the division of labor smarter?" The next real leap in AI might not come from teaching the model more, but from admitting that "the model" was never really one thing to begin with.
The essay traces a core tension in AI development: the industry's decade-long strategy of scaling—training larger models on more data—has delivered impressive results, but appears to be hitting a ceiling. The author argues this is because scaling addresses only half the problem: growing the library of facts. Retrieval-augmented systems already prove the internet itself is a better knowledge store than any model's frozen parameters, yet current AI continues to burden itself with memorizing vast corpora. The real bottleneck, the piece contends, is the librarian function—the ability to learn and apply rules to novel situations rather than interpolate from training data.
The distinction between memorization and understanding crystallizes through the multiplication example: a child who memorizes times tables excels only on seen combinations, while one who understands multiplication's underlying principle can solve entirely novel problems. Large language models, despite their sophistication, often behave like the first child—extraordinarily adept at recombining patterns from training data but brittle outside that familiar territory. The failure mode described at the essay's opening—a model confidently assembling individually correct pieces into illogical wholes—exemplifies this gap.
The author's proposed architectural shift—from a single monolithic model to a cluster of specialized subsystems coordinated by an orchestrator—is not wholly novel in theory but has become practically feasible with modern AI. However, the essay's most honest concession is that knowledge and reasoning are currently tangled together in training; complete separation may be impossible. Still, the suggestion that the industry has likely over-weighted library-building relative to librarian-building implies a concrete direction for future research: smarter division of labor rather than raw scaling.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No discussion yet for this article
Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime