
Apple researchers have completed a comprehensive analysis of how to improve multimodal AI models (systems that understand both text and images) by aligning them to human preferences.
The study reveals that combining offline and online alignment methods can boost model performance, and introduces a new low-cost technique for generating training data that reduces hallucination—where models produce responses inconsistent with image content.
This work addresses a significant gap in understanding what makes multimodal alignment effective.
What happened
Apple researchers published a comprehensive study analyzing preference alignment techniques in multimodal large language models (MLLMs)—AI systems that process both text and images. The team categorized alignment algorithms into offline methods (such as Direct Preference Optimization) and online methods (such as online-DPO), and found that combining both approaches can improve model performance in certain scenarios.
Why it matters
Multimodal models often suffer from hallucination—generating responses inconsistent with image content or stating incorrect facts about visual information. Understanding which alignment methods work best helps developers reduce these errors and make MLLMs more reliable for tasks requiring accurate image understanding, which is increasingly important as these models see wider deployment.
What to watch
The researchers introduced a new technique called Bias-Driven Hallucination Sampling (BDHS) for creating preference training data that requires neither additional human annotation nor external models, while achieving competitive performance across multiple benchmarks—potentially lowering the cost and complexity of aligning future multimodal models.
Apple's Machine Learning research team has published a comprehensive study on preference alignment in multimodal large language models, authored by Elmira Amirloo, Jean-Philippe Fauconnier, Christoph Roesmann, Christian Kerl, Rinu Boney, Yusu Qian, Zirui Wang, Afshin Dehghan, Yinfei Yang, Zhe Gan, and Peter Grasch. The paper examines how to improve MLLMs—AI systems that understand both text and images—by aligning them to human preferences.
The core problem the researchers address is hallucination in multimodal contexts. Unlike text-only models, MLLMs can hallucinate in two ways: by stating incorrect facts, or by generating responses that contradict the visual information presented. The team's objective is to encourage these models to ground their outputs more firmly in actual image content. While recent work has introduced preference datasets and tested alignment methods like Direct Preference Optimization (DPO) and Proximal Policy Optimization (PPO), the lack of standardization across these studies has made it unclear which specific components—datasets, base models, or algorithms—drive the reported improvements.
The researchers take a systematic approach, independently analyzing each aspect of preference alignment. They categorize alignment algorithms into two broad groups: offline methods (exemplified by DPO) and online methods (exemplified by online-DPO). A key finding is that combining offline and online approaches can enhance model performance in certain scenarios, suggesting that no single method is universally optimal. The team also reviews existing multimodal preference datasets in detail, examining how construction choices affect downstream model quality.
The study's main contribution is the introduction of Bias-Driven Hallucination Sampling (BDHS), a new technique for generating multimodal preference data. Unlike prior approaches, BDHS requires neither additional human annotation nor external models, yet achieves performance competitive with previously published alignment work when tested across multiple benchmarks. This efficiency gain is significant because data preparation is often a major cost and resource bottleneck in alignment research, and removing those constraints may accelerate progress in making multimodal AI systems more reliable and aligned with human expectations.
The study addresses a notable gap in multimodal AI research: while preference alignment has become standard practice in training large language models, its application and effectiveness in multimodal systems—which handle both text and images—remained less well understood. The researchers systematically isolate and test individual components of alignment pipelines, recognizing that prior work varied widely in datasets, base models, and methods, making it difficult to identify which factors drove improvements.
By categorizing alignment approaches into offline and online methods, the team demonstrates that their combination can yield benefits beyond either alone in certain scenarios. This finding is practically significant because it suggests a flexible, modular approach to alignment rather than a single optimal technique. The introduction of BDHS is particularly noteworthy because it bypasses the traditional bottleneck of alignment research: the need for costly human annotation or reliance on separate external models to generate preference data. This lower-friction approach may accelerate the development and refinement of multimodal systems in industry.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic's Claude AI, working on the unsolved mathematics problem known as the Riemann hypothesis, initially…

DeepSeek released V4 Pro 0813, its latest Pro model, available through OpenRouter via API

Twitch announced today that users can now opt out of allowing Amazon to use their streams, VODs, clips, chats…

Anthropic added invisible watermarks to Claude's outputs to comply with the EU AI Act's requirement that AI-ge…

Eli Lilly and Company has committed to roll out Veeva Vault CRM across its global operations, adopting Veeva's…

Visa, the global payments network, is positioning itself to profit from AI-driven commerce by securing its rol…

The AI news that matters, in one minute each morning.
Sign up free