
What happened
Apple researchers tested all pairwise combinations of sixteen models on a round-trip task — one turns an arithmetic expression into a word problem, another recovers it. Swapping roles shifts accuracy by up to 60.4 points, and the best pair reaches 92.9%.
Why it matters
The channel is lossy and asymmetric, since pairing different models on each end beats using the same model twice, suggesting generation and extraction are separate skills rather than one general ability.
What to watch
At least 73.6% of failures start at generation, so gains hinge on improving the writing side. Watch whether the reported lift from roughly 3600 fine-tuning examples holds up in the disjoint-domain regime, where a gap to the frontier remains.
WHO IT HITSThis matters most to teams building multi-agent or chain-of-thought pipelines, where one model hands structured work to another in plain text; the finding suggests mixing models at each end can beat using one model throughout.
Summaries like this, in your inbox every morning.
The study frames natural language as a channel that models must squeeze structured information through, and asks how much tree-shaped content survives the trip. To test this, the authors built a round-trip protocol: a generator turns a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence acts as an exact oracle. Running every pairwise combination of sixteen models produces a communication matrix whose marginals separate how well a model writes from how well it reads.
The results point in three directions. The channel is lossy and asymmetric, since the identity of the writer and the reader matters more than the model family. The difficulty tracks the shape of the expression — operator count, depth, right-branching — rather than which lab built the model, and most failures begin on the generation side. And the channel is trainable: a modest set of fine-tuning examples sharing the evaluation's operators and tree shapes lifts every open-weight model above untrained Gemini-3.1-Pro, an upper bound under matched semantics, while a disjoint-domain regime with new operators and vocabulary also raises every open-weight model, confirming the gain is not an artifact of matched semantics.
For anyone wiring models together in multi-step or multi-agent setups, the practical reading is that who writes and who reads is a design choice worth testing rather than assuming away. The remaining gap to the frontier suggests the bottleneck is not fully closed by the fine-tuning described, and the authors identify tree-structured expression serialization as a primary limiting factor when models communicate hierarchical structure through natural language.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Google is paying about 100 digital publishers for content used in AI Overviews, AI Mode, and Gemini, with paym…

A Qiita walkthrough trained a five-label car-damage classifier on Gemini Enterprise Agent Platform AutoML usin…

Lauren Tan says she shipped about 2,000 pull requests a month to production on the SpaceX AI Grok Bot team

At its September 29, 2026 DevDay, OpenAI announced more than 20 items, including dots, an agent running on GPT…

A student made granite-code:8b and granite3.2:8b write a TORCS racing AI in 13 parts, checked by Python test s…

Alibaba's Qwen team open-sourced Qwen-Image-2.1 on September 20, 2026 — a 7B model generating 2048×2048 images…
