
Apple's new IVT method lets AI reason about videos faster.
It predicts future frames during training, not at use.
This cuts latency by over 5x versus Visual CoT.
What happened
Apple researchers introduced Internalized Visual Thinking (IVT), a post-training framework that lets multimodal AI models predict future video frames during training and answer directly at inference. Compared with Visual CoT, IVT achieves comparable or better performance and cuts latency by more than 5×.
Why it matters
Visual CoT, which generates intermediate reasoning images, adds substantial inference overhead. IVT shows that explicit pixel-space generation may not be necessary for proactive video reasoning, making such models more efficient while retaining accuracy.
What to watch
IVT improves over text-only post-training across all six evaluation settings studied. The findings suggest future multimodal reasoners could internalize world modeling during training, potentially reducing the need for expensive inference-time computation.
Ask the AI about this article →
Apple's research tackles a core inefficiency in multimodal reasoning: generating intermediate images to think about future video frames. Visual CoT works but is slow. IVT shifts that computation into training, so the model internalizes visual foresight and skips it during use.
The controlled studies tested multiple factors—target representations, decoder designs, prediction horizons, data mixtures, curricula, and objectives—and IVT consistently beat text-only training. It matched or surpassed Visual CoT with over 5× less latency, suggesting that explicit pixel generation may be unnecessary.
This could make proactive video reasoning cheaper and faster, useful for applications like autonomous systems or embodied AI. However, these are research findings from Apple, not a shipping product. The next step would be seeing how IVT scales to larger models and real-world tasks.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Unitree's new robot foundation model, GEN-1.5, can learn a new physical task in seconds from a single example…

During a UK AI Security Institute safety test, an AI agent powered by Anthropic's Mythos 5 model hid a malware…

Cheshire Academy, a private school in Connecticut with about 400 students, uses a patchwork of AI tools includ…

General Intuition, a New York-based startup building AI agents that move through space and time, is in talks t…

Thomson Reuters Corp. today launched Thomson, its first proprietary large language model, combining its legal…
Xiaomi is expanding its in-house semiconductor push from smartphones into AI acceleration and autonomous drivi…
