AIToday
Large Language ModelsarXiv cs.LGPublished: Apr 15, 2026, 13:00 JST1 min read

Researchers discover that normalization-free transformers achieve better gradient stability through subcritical signal propagation, offering insights into deep neural network initialization.

Researchers discover that normalization-free transformers achieve better gradient stability through subcritical signal propagation, offering insights into deep neural network initialization.

3 Key Points

  1. Study analyzes signal propagation in transformers using averaged partial Jacobian norm (APJN) to measure gradient amplification across layers

  2. Theory extends to bidirectional attention and permutation-symmetric token configurations by deriving recurrence relations for activation statistics

  3. Pre-LayerNorm architectures show power-law APJN growth, while tanh-like nonlinearities exhibit stretched-exponential growth indicating subcritical behavior

  4. Findings apply to Dynamic Tanh (DyT) and Dynamic erf (Derf) transformers, explaining why these normalization-free designs achieve better training stability

  5. Predictions match empirical measurements in deep vision transformers, bridging theoretical understanding and practical deep network behavior

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Sandisk's HBF claims 16x HBM capacityDIGITIMES Asia · 3m ago
  • World Labs unveils Atlas, a single AI model that generates, reconstructs, and simulates 3D worlds from just a few photosTHE DECODER · 4m ago
  • Saudi unveils $15B tech deals at LEAPFortune AI · 4m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleResearchers develop new behavioral profiling method to measure how AI agents balance task execution with safety refusals in real-world deployments