
Researchers successfully hand-coded neural network weights that memorize token sequences with the same linear scaling relationship as trained models—suggesting that efficient memorization follows from architecture alone. However, the hand-coded approach achieved lower absolute performance than gradient-trained models, indicating that training optimization adds a multiplicative boost beyond what the architecture structure alone provides.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Researchers hand-coded weights for single-layer MLPs that memorize labels for two-token input sequences, finding that their models' capacity to memorize facts with 90% accuracy scales roughly linearly with parameter count — matching the scaling behavior of trained models.
Why it matters
The finding suggests that efficient memorization scaling is achievable through direct weight construction rather than gradient-based training alone, offering insight into how neural network architecture determines learning capacity independent of the training method used.
What to watch
The hand-coded models' scaling prefactor (the constant multiplying the linear relationship) still falls short of trained models' by a factor whose specific value the article indicates but does not fully render in the available text.
The researchers constructed weights by hand for single-layer MLPs tasked with memorizing labels associated with two-token input sequences. Rather than using standard gradient-based training, they directly coded the weight values, then measured how many facts each model could memorize at 90% accuracy. Scaling the models up by increasing parameter count revealed a roughly linear relationship between model size and memorization capacity—the same linear scaling observed in models trained conventionally. This parallel scaling behavior is noteworthy because it shows that the efficiency of the scaling relationship is not unique to learned weights but emerges from the architecture itself. Yet the hand-coded approach did not fully match trained models: the scaling prefactor—the constant that multiplies the linear term—fell short of the trained models' prefactor by a specific factor (the article's rendering of this factor is incomplete in the available text). This gap indicates that while architecture determines the form of scaling (linear in this case), the training process optimizes weights in ways that improve the absolute capacity within that scaling regime. The finding suggests a decomposition of neural network efficiency: one part stems from structural properties of the architecture, while another part emerges from optimization during training.
The experiment addresses a foundational question in deep learning: whether efficient scaling emerges from network architecture alone or requires the optimization process that gradient-based training provides. By hand-coding weights rather than training them, the researchers isolate the contribution of architecture structure to memorization capacity. The result—that scaling behavior matches between hand-coded and trained models—suggests that the linear relationship between parameters and memorized facts is largely determined by the MLP structure itself, not by the particulars of the training algorithm. However, the gap in the scaling prefactor reveals that trained models extract additional efficiency from optimization, learning weight configurations that trained models can achieve but hand-coding cannot easily reproduce. This distinction between structural scaling and optimization-driven improvement clarifies what gradient descent adds beyond what is already implicit in the network architecture.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack