
The approach lets multiple people independently fine-tune copies of a base model on different domains/languages with no communication, then combines them with a lightweight Mixture of Experts router trained in ~500 steps
Testing on Pythia models (410M to 6.9B parameters) shows consistent gains of 6.5-8% over the best individual specialist
A simple linear formula (R² = 0.856) can predict fusion effectiveness from specialist divergence before training begins: gain = 0.82 × divergence − 2.72
Code, paper, and project page are publicly available on GitHub and arXiv, with particularly promising cross-lingual fusion results
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Google DeepMind chief Koray Kavukcuoglu said being at the frontier of AI is the only thing that matters to the…

John Deere is testing an AI assistant called “JD” that answers farmers' questions on topics like equipment set…

Google has launched Google Pics, a new suite of creative design tools for Workspace users, built around Gemini…

OpenAI said today that it is integrating ChatGPT Health with Epic's electronic health record (EHR) system, whi…

Google is launching Google Pics, an AI-powered image creation and editing tool that will be part of Google Wor…

Google DeepMind launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite
