
Study reveals supervised fine-tuning (SFT) functions as a special case of policy gradient optimization with sparse rewards and unstable probability weighting, causing training instability
Group Fine-Tuning framework introduces Group Advantage Learning to create diverse response groups and normalized contrastive supervision, reducing reward sparsity issues
Dynamic Coefficient Rectification mechanism adaptively controls inverse-probability weights to stabilize the optimization process and prevent gradient explosion
GFT aims to unify knowledge injection with robust generalization, addressing single-path dependency and entropy collapse problems in current post-training methods
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Israeli startup DataAgent Ltd
SK Hynix presented a custom HBM concept at SEMICON Taiwan 2026, where compute functions are placed in the base…

Nvidia reported earnings that were both remarkable and boring, reflecting its focus on avoiding a consolidated…

Anthropic has agreed to a $35bn cloud-computing contract with Lambda, a Nvidia-backed cloud provider

The Supreme Court of Japan has included about ¥60 million in its fiscal 2027 budget request for AI-related exp…

The Consumer Affairs Agency said Tuesday it will use generative AI to analyze about 900,000 annual consultatio…
